Residual and coefficient coding for video coding
By disabling Rice parameters and employing extended precision processing for transform skip residuals, the method addresses inefficiencies in video coding, achieving improved compression efficiency and reduced complexity in hardware decoders.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-04
AI Technical Summary
Existing video coding techniques face challenges in achieving better compression efficiency beyond the capabilities of standards like HEVC, particularly in handling transform skip residuals and coefficient coding, which can lead to increased complexity and inefficiencies in hardware decoders.
The method involves disabling the Rice parameter for transform skip residual coding, employing extended precision processing flags, and using persistent rice adaptation for binarization of transform coefficients, along with improved residual coding structures and quantization techniques to enhance video encoding efficiency.
This approach improves video encoding by reducing complexity and enhancing compression efficiency, particularly in hardware decoders, by optimizing transform skip residuals and coefficient coding processes.
Smart Images

Figure 2026035734000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to Provisional Application No. 63 / 181,110, filed April 28, 2021, the entire contents of which are incorporated herein by reference for all purposes.
[0002] This disclosure relates to video encoding and compression, and more particularly to improving and simplifying residual and coefficient coding for video coding. [Background technology]
[0003] Various video coding techniques may be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include versatile video coding (VVC), joint exploration test model (JEM), high-efficiency video coding (H.265 / HEVC), advanced video coding (H.264 / AVC), moving picture expert group (MPEG) coding, or similar. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, or similar) that exploit redundancy present in moving images or sequences. An important goal of video coding techniques is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]
[0004] Examples of this disclosure provide methods and apparatus for video encoding. [Means for solving the problem]
[0005] According to a first aspect of the present disclosure, there is provided a method for video encoding, the method may include, in response to determining, by a decoder, that transform skip is disabled, disabling presence of a Rice parameter for transform skip residual coding.
[0006] According to a second aspect of the present disclosure, a method for video encoding is provided. The method may include receiving, by a decoder, a sequence parameter set (SPS) alignment enable flag indicating whether an index ivlCurrRange is aligned before bypass decoding of syntax elements sb_coded_flag, abs_remainder, dec_abs_level, and coeff_sign_flagn based on a value of SPS alignment enable.
[0007] According to a third aspect of the present disclosure, there is provided a method for video encoding, the method may include receiving, by a decoder, an extended precision processing flag indicating whether extended dynamic range is employed for transform coefficients and during transform processing based on a value of the extended precision processing flag.
[0008] According to a fourth aspect of the present disclosure, there is provided a method for video encoding. The method may include receiving, by a decoder, a persistent rice adaptation enabled flag indicating whether Rice parameter derivation for binarization of abs_remaining and dec_abs_level is initialized at the beginning of each sub-block employing mode-dependent statistics accumulated from a previous sub-block based on a value of the persistent rice adaptation enabled flag.
[0009] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to be restrictive of the present disclosure.
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 2 is a block diagram of an encoder according to an example of the present disclosure. [Figure 2] FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3A] FIG. 10 is a diagram illustrating block divisions in a polymorphic tree structure according to an example of the present disclosure. [Figure 3B] FIG. 10 is a diagram illustrating block divisions in a polymorphic tree structure according to an example of the present disclosure. [Figure 3C] FIG. 10 is a diagram illustrating block divisions in a polymorphic tree structure according to an example of the present disclosure. [Figure 3D] FIG. 10 is a diagram illustrating block divisions in a polymorphic tree structure according to an example of the present disclosure. [Figure 3E] FIG. 10 is a diagram illustrating block divisions in a polymorphic tree structure according to an example of the present disclosure. [Figure 4] FIG. 10 illustrates a picture having 18×12 luma CTUs according to an example of the present disclosure. [Figure 5] FIG. 2 is a diagram of a picture having 18×12 luma CTUs according to an example of the present disclosure. [Figure 6A] FIG. 1 is a diagram of an example of disallowed ternary tree (TT) and binary tree (BT) partitioning in a VTM according to an example of the present disclosure. [Figure 6B] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 6C] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 6D] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 6E] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 6F] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 6G] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 6H] FIG. 10 is a diagram of an example of disallowed TT and BT partitioning in a VTM according to an example of the present disclosure. [Figure 7] FIG. 10 is a diagram of a residual coding structure for a transform block according to an example of the present disclosure. [Figure 8] FIG. 10 is a diagram of a residual coding structure for a transform skip block according to an example of the present disclosure. [Figure 9] FIG. 2 is a diagram of two scalar quantizers according to an example of the present disclosure. [Figure 10A] FIG. 10 is a diagram of state transitions according to an example of the present disclosure. [Figure 10B] FIG. 10 is a diagram of quantizer selection according to an example of the present disclosure. [Figure 11] FIG. 1 is a diagram of a template used to select a probabilistic model according to the present disclosure. [Figure 12] FIG. 2 is a diagram of an example of a block coded in palette mode according to the present disclosure. [Figure 13] FIG. 1 illustrates the use of a palette predictor to signal palette entries according to the present disclosure. [Figure 14A] FIG. 10 is a diagram of horizontal translation scanning according to the present disclosure. [Figure 14B] FIG. 10 is a diagram of vertical translation scanning according to the present disclosure. [Figure 15A] FIG. 10 illustrates a sub-block based index map scan for a palette in accordance with the present disclosure. [Figure 15B]FIG. 10 illustrates a sub-block based index map scan for a palette in accordance with the present disclosure. [Figure 16] FIG. 1 illustrates a method for encoding a video signal according to an example of the present disclosure. [Figure 17] FIG. 1 illustrates a method for encoding a video signal according to an example of the present disclosure. [Figure 18] FIG. 1 illustrates a computing environment coupled with a user interface according to an example of the present disclosure. [Figure 19] FIG. 1 illustrates a method for video encoding according to an example of the present disclosure. [Figure 20] FIG. 1 illustrates a method for video encoding according to an example of the present disclosure. [Figure 21] FIG. 1 illustrates a method for video encoding according to an example of the present disclosure. [Figure 22] FIG. 1 illustrates a method for video encoding according to an example of the present disclosure. [Figure 23] FIG. 1 is a block diagram illustrating an example system for encoding and decoding video blocks according to an example of this disclosure. [Figure 24] FIG. 2 is a block diagram illustrating an exemplary video encoder according to one example of this disclosure. [Figure 25] FIG. 2 is a block diagram illustrating an exemplary video decoder according to one example of this disclosure. [Figure 26] FIG. 1 illustrates a low-delay transform skip residual coding (TSRC) method according to an example of the present disclosure. [Figure 27] FIG. 1 illustrates a method for video encoding according to an example of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations described in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure. Rather, the implementations are merely examples of apparatus and methods consistent with aspects related to the present disclosure as set forth in the appended claims.
[0013] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It is also to be understood that the term "and / or," as used herein, is intended to mean and include any and all possible combinations of one or more of the associated listed items.
[0014] Terms such as "first," "second," and "third" may be used herein to describe various pieces of information, but it should be understood that these terms are not intended to limit the information. These terms are used only to distinguish one category of information from another. For example, first information may be referred to as second information, and similarly, second information may be referred to as first information, without departing from the scope of this disclosure. The term "if," as used herein, may be understood to mean "when," "in the event of," or "at the discretion of," depending on the context.
[0015] The first version of the HEVC standard was finalized in October 2013, providing approximately 50% bitrate reduction or equivalent perceptual quality compared to the previous generation video coding standard, H.264 / MPEG AVC. While the HEVC standard offers significant coding improvements over previous standards, it has been demonstrated that better coding efficiency than HEVC can be achieved by using additional coding tools. Based on this, both VCEG and MPEG have begun work on exploring new coding techniques for future video coding standardization. In October 2015, the ITU-T VECG and ISO / IEC MPEG formed the Joint Video Exploration Team (JVET) to begin significant research into advanced technologies that could enable significant improvements in coding efficiency. The JVET maintained a single reference software, called the Joint Exploration Model (JEM), by integrating several additional coding tools on top of the HEVC Test Model (HM).
[0016] In October 2017, ITU-T and ISO / IEC issued a joint call for proposals (CfP) for video compression with capabilities beyond HEVC. In April 2018, 23 responses were received and evaluated at the 10th JVET meeting, demonstrating approximately 40% improvement in compression efficiency over HEVC. Based on these evaluation results, JVET launched a new project to develop a new generation of video coding standard named Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.
[0017] Like HEVC, VVC is built on a block-based hybrid video coding framework.
[0018] Figure 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, Figure 1 shows a typical encoder 100. The encoder 100 includes a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.
[0019] In encoder 100, a video frame is partitioned into video blocks for processing. For a given video block, a prediction is formed based on either an inter-prediction technique or an intra-prediction technique.
[0020] A prediction residual, representing the difference between a current video block, which is part of video input 110, and its predictor, which is part of block predictor 140, is sent from summer 128 to transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantized coefficients are then provided to entropy coding 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 142 from intra / inter mode decision 116, such as video block partition information, motion vectors (MVs), reference picture indices, and intra prediction modes, is also provided via entropy coding 138 and stored in compressed bitstream 144. Compressed bitstream 144 comprises a video bitstream.
[0021] The encoder 100 also requires decoder-related circuitry to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed through inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with a block predictor 140 to generate unfiltered reconstructed pixels for the current video block.
[0022] Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples (called reference samples) of already coded neighboring blocks in the same video frame as the current video block.
[0023] Temporal prediction (also called "inter-prediction") uses reconstructed pixels from an already coded video picture to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, one reference picture index is sent that is used to identify which reference picture in reference picture storage the temporal prediction signal comes from.
[0024] Motion estimation 114 takes signals from video input 110 and picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 takes signals from video input 110, picture buffer 120, and the motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0025] After spatial prediction and / or temporal prediction are performed, intra / inter mode decision 116 within encoder 100 selects the best prediction mode, for example, based on a rate-distortion optimization method. Block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantized residual coefficients are inversely quantized by inverse quantization 134 and inversely transformed by inverse transform 136 to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Further in-loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied to the reconstructed CU before it is placed into reference picture storage in picture buffer 120 and used to encode future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 138, which further compresses and packs them to form the bitstream.
[0026] Figure 1 shows a block diagram of a typical block-based hybrid video coding system. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128 x 128 pixels. However, unlike HEVC, which partitions blocks solely based on a quadtree, VVC decomposes a single coding tree unit (CTU) into CUs based on a quadtree, binary tree, or ternary tree to accommodate various local characteristics. By definition, a coding tree block (CTB) is an N x N block of samples, where N is some value. Consequently, the division of components into CTBs is a partitioning. A CTU contains a CTB for luma samples, two corresponding CTBs for chroma samples for a picture with a three-sample array, or a CTB for samples for a monochrome picture or a picture coded using three separate color planes and the syntax structure used to code the samples. Furthermore, the concept of multiple partitioning unit types in HEVC has been eliminated; that is, in VVC, the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists. Instead, each CU is always used as the basic unit for both prediction and transformation without further partitioning. In the polymorphic tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using a binary tree structure and a ternary tree structure. As shown in Figures 3A, 3B, 3C, 3D, and 3E, there are five decomposition types: quadrant partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal triangular partitioning, and vertical triangular partitioning.
[0027] FIG. 3A shows a diagram illustrating block quadrants in a polymorphic tree structure according to the present disclosure.
[0028] FIG. 3B shows a diagram illustrating block vertical bisections in a polymorphic tree structure according to the present disclosure.
[0029] FIG. 3C shows a diagram illustrating horizontal bisection of blocks in a polymorphic tree structure according to the present disclosure.
[0030] FIG. 3D shows a diagram illustrating vertical third division of blocks in a polymorphic tree structure according to the present disclosure.
[0031] FIG. 3E shows a diagram illustrating horizontal trichotomization of blocks in a polymorphic tree structure according to the present disclosure.
[0032] In Figure 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples (called reference samples) of already coded neighboring blocks in the same video picture / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter prediction" or "motion-compensated prediction") predicts the current video block using reconstructed pixels from already coded video pictures. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal references. If multiple reference pictures are supported, an additional reference picture index is also sent, which is used to identify which reference picture in the reference picture store the temporal prediction signal comes from. After spatial prediction and / or temporal prediction, a mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The predictive block is then subtracted from the current video block, and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, which is then added back to the predictive block to form a reconstructed signal for the CU. Further in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF), may be applied to the reconstructed CU before it is placed in a reference picture store and used to encode future video blocks. To form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit, where they are further compressed and packed to form the bitstream.
[0033] Figure 2 shows a schematic block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of an exemplary decoder 200. The decoder 200 includes a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.
[0034] The decoder 200 is similar to the reconstruction-related section present in the encoder 100 of FIG. 1. In the decoder 200, an incoming input video bitstream 210 is first decoded through entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through inverse quantization 214 and inverse transform 216 to obtain a reconstructed prediction residual. A block predictor mechanism implemented in an intra / inter mode selector 220 is configured to perform either intra prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residual from the inverse transform 216 and the prediction output generated by the block predictor mechanism using an adder 218.
[0035] The reconstructed blocks may further pass through an in-loop filter 228 before being stored in a picture buffer 226, which serves as a reference picture store. The reconstructed video in the picture buffer 226 may be sent to drive a display device and may also be used to predict future video blocks. In situations where the in-loop filter 228 is on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0036] Figure 2 shows a schematic block diagram of a block-based video decoder. First, the video bitstream is entropy decoded in an entropy decoding unit. The coding mode and prediction information are sent to a spatial prediction unit (for intra-coding) or a temporal prediction unit (for inter-coding) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then summed. The reconstructed block may undergo further in-loop filtering before being stored in a reference picture store. The reconstructed video in the reference picture store is then sent to drive a display device and used to predict future video blocks.
[0037] In general, the basic intra prediction scheme applied in VVC is kept the same as that of HEVC, except that some modules are further extended and / or improved, e.g., intra sub-partitioning (ISP) coding mode, enhanced intra prediction with wide-angle intra direction, position-dependent intra prediction combination (PDPC), and 4-tap intra interpolation.
[0038] Partitioning Pictures, Tile Groups, Tiles, and CTUs in VVC In VVC, a tile is defined as a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile group is a group of an integer number of tiles of a picture that are exclusively contained in a single NAL unit. Essentially, the concept of a tile group is the same as that of a slice defined in HEVC. For example, a picture is divided into tile groups and tiles. A tile is a sequence of CTUs that covers a rectangular region of a picture. A tile group contains several tiles of a picture. Two modes of tile groups are supported: raster scan tile group mode and rectangular tile group mode. In raster scan tile group mode, a tile group contains a sequence of tiles within the tile raster scan of a picture. In rectangular tile group mode, a tile group contains several tiles of a picture that collectively form a rectangular region of the picture. The tiles in a rectangular tile group are in the order of the tile raster scan of the tile group.
[0039] Figure 4 shows an example of raster scan tile group partitioning of a picture, where the picture is divided into 12 tiles and three raster scan tile groups. Figure 4 includes tiles 410, 412, 414, 416, and 418. Each tile has 18 CTUs. More specifically, Figure 4 shows a picture with 18x12 luma CTUs partitioned into 12 tiles and three tile groups (for reference). The three tile groups are as follows: (1) the first tile group includes tiles 410 and 412; (2) the second tile group includes tiles 414, 416, 418, 420, and 422; and (3) the third tile group includes tiles 424, 426, 428, 430, and 432.
[0040] Figure 5 shows an example of rectangular tile group partitioning of a picture, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular tile groups. Figure 5 includes tiles 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, and 556. More specifically, Figure 5 shows a picture with 18x12 luminance CTUs partitioned into 24 tiles and 9 tile groups (for reference). Tile groups contain tiles, and tiles contain CTUs. The nine rectangular tile groups are: (1) two tiles 510 and 512; (2) two tiles 514 and 516; (3) two tiles 518 and 520; (4) four tiles 522, 524, 534, and 536; and (5) four tiles Le5 26, 528, 538, and 540, (6) four tiles 530, 532, 542, and 544, (7) two tiles 546 and 548, (8) two tiles 550 and 552, and (9) two tiles 554 and 556.
[0041] Large block size transform with high frequency zero setting in VVC. VTM4 allows transforms of large block sizes up to 64x64, which are primarily useful for high-resolution video, e.g., 1080p and 4K sequences. High-frequency transform coefficients of a transform block whose size (width or height, or both width and height) equals 64 are set to zero, resulting in only low-frequency coefficients being retained. For example, for an MxN transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the top 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without setting values to zero.
[0042] Virtual Pipeline Data Unit (VPDU) in VVC A virtual pipeline data unit (VPDU) is defined as a non-overlapping unit within a picture. In hardware decoders, consecutive VPDUs are processed simultaneously by multiple pipeline stages. The VPDU size is roughly proportional to the buffer size in most pipeline stages, so it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) and binary tree (BT) partitioning can result in an increase in VPDU size.
[0043] To keep the VPDU size as 64x64 luma samples, the following normative segmentation restrictions (with modifications to syntax signaling) apply to VTM5:
[0044] TT decomposition is not allowed for CUs with either width or height, or both width and height, equal to 128.
[0045] If CU is 128 × N, and N≦64 (i.e., width is equal to 128 and height is less than 128), then horizontal BT classification is not acceptable.
[0046] N × 128 CUs, where N≦64 (i.e., height equals 128 and width less than 128), then vertical BT classification is not acceptable.
[0047] 6A, 6B, 6C, 6D, 6E, 6F, 6G, and 6H show examples of impermissible TT and BT partitioning in a VTM.
[0048] Transform Coefficient Coding in VVC Transform coefficient coding in VVC is similar to HEVC in that they both use non-overlapping coefficient groups (also called CGs or sub-blocks). However, there are some differences between them. In HEVC, each CG of a coefficient has a fixed size of 4x4. In VVC Draft 6, the CG size became dependent on the TB size. As a result, various CG sizes are available in VVC: 1x16, 2x8, 8x2, 2x4, 4x2, and 16x1. CGs within a coding block, and the transform coefficients within a CG, are coded according to a predefined scan order.
[0049] To limit the maximum number of context-coded bins per pixel, the area of the TB and the type of video component (e.g., luma vs. chroma) are used to derive the maximum number of context-coded bins (CCBs) of the TB. The maximum number of context-coded bins is equal to TB_zosize * 1.75, where TB_zosize indicates the number of samples in the TB after coefficient zeroing. Note that coded_sub_block_flag, a flag indicating whether the CG contains non-zero coefficients, is not taken into account in the CCB count.
[0050] Coefficient zeroing is an operation performed on a transform block to force coefficients located in a particular region of the transform block to zero. For example, in current VVC, a 64x64 transform has an associated zeroing operation. As a result, all transform coefficients located outside the top-left 32x32 region within a 64x64 transform block are forced to zero. In fact, in current VVC, for any transform block with a size greater than 32 along a particular dimension, a coefficient zeroing operation is performed along that dimension to force coefficients located beyond the top-left 32x32 region to zero.
[0051] Transform coefficient coding in VVC begins with setting the variable remBinsPass1 to the maximum number of context-coded bins (MCCB) allowed. During the coding process, the variable is decremented by one each time a context-coded bin is signaled. While remBinsPass1 is 4 or greater, coefficients are first signaled through the sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag constructs, all of which use context-coded bins in the first pass. The remainder of the coefficient's level information is coded using the abs_remainder construct, using Golomb-Rice coding and bypass-coded bins in the second pass. If remBinsPass1 becomes less than 4 during the first pass, the current coefficient is not coded in the first pass. Instead, it is directly coded in the second pass using the dec_abs_level syntax element with Golomb-Rice coding and bypass coding bins. The Rice parameter derivation process for dec_abs_level[] is derived as specified in Table 3. After coding all levels described above, the signs (sign_flag) of all scan positions where sig_coeff_flag is equal to 1 are coded as bypass bins. This process is depicted in Figure 7. remBinsPass1 is reset for each TB. The transition from using context coding bins for sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag to using bypass coding bins for the remaining coefficients occurs at most once per TB. For a coefficient sub-block, if remBinsPass1 is less than 4 before coding the first coefficient, the entire coefficient sub-block is coded using bypass coding bins.
[0052] FIG. 7 shows a diagram of the residual coding structure of a transform block.
[0053] A unified (same) derivation of the Rice parameter (RicePara) is used to signal the syntax of abs_remainder and dec_abs_level. The only difference is that the base level baseLevel is set to 4 and 0 to code abs_remainder and dec_abs_level, respectively. The Rice parameter is determined based not only on the sum of the absolute levels of the five adjacent transform coefficients in the local template, but also on the corresponding base levels as follows:
[0054] RicePara=RiceParTable[max(min(31,sumAbs-5*baseLevel),0)]
[0055] The syntax and associated semantics for residual coding in the current VVC draft specification are shown in Tables 1 and 2, respectively. How to read Table 1 is provided in the appendix section of this disclosure and is also described in the VVC specification. [Table 1] JPEG2026035734000003.jpg249149 JPEG2026035734000004.jpg254160 JPEG2026035734000005.jpg250153 JPEG2026035734000006.jpg99165 [Table 2] JPEG2026035734000008.jpg250149 JPEG2026035734000009.jpg252155 JPEG2026035734000010.jpg114165 [Table 3] [Table 4]
[0056] Transform skip mode residual coding in VVC. Unlike HEVC, in which a single residual coding scheme is designed to code both transform coefficients and transform skip coefficients, in VVC, two separate residual coding schemes are used for the transform coefficients and transform skip coefficients (i.e., the residual), respectively.
[0057] In transform skip mode, the statistical properties of the residual signal are different from those of the transform coefficients, and no energy compaction around low-frequency components is observed. No signal transmission of the last x / y position, When all previous flags are equal to 0, coded_sub_block_flag is coded for all sub-blocks except the DC sub-block; sig_coeff_flag context modeling using two adjacent coefficients, par_level_flag to use only one context model, Add more than 5, 7, 9 flags, Modified Rice parameter derivation for residual binarization, The context modeling of the sign flag is determined based on the left and top neighboring coefficient values, and after sig_coeff_flag, the sign flag is parsed to keep all the context coding bins together. The (spatial) transform skip residual is modified to take into account various signal characteristics, including:
[0058] As shown in Figure 8, the syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag are coded in an interleaved manner for each residual sample in the first pass, followed by coding of the abs_level_gtX_flag bitplane in the second pass and abs_remainder in the third pass.
[0059] Path 1: sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag
[0060] Path 2: abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, abs_level_gt9_flag
[0061] Pass 3: abs_remainder
[0062] FIG. 8 shows a diagram of the residual coding structure of a transform skip block.
[0063] The syntax and associated semantics for transform skip mode residual coding in the current VVC draft specification are shown in Table 5 and Table 2, respectively. How to read Table 5 is provided in the appendix section of this disclosure and is also described in the VVC specification. [Table 5] JPEG2026035734000014.jpg252153 JPEG2026035734000015.jpg248159 JPEG2026035734000016.jpg49165
[0064] Quantization In the current VVC, the maximum QP value has been extended from 51 to 63, and the signaling of the initial QP has been changed accordingly. When a non-zero value of slice_qp_delta is coded, the initial value of SliceQpY can be modified at the slice segment layer. For transform skip blocks, when QP equals 4, the quantization step size is 1, so the minimum allowed quantization parameter (QP) is defined as 4.
[0065] Furthermore, the same HEVC scalar quantization is used with a new concept called dependent scalar quantization. Dependent scalar quantization refers to a technique in which the set of allowable reconstructed values of a transform coefficient depends on the value of the transform coefficient level preceding the current transform coefficient level in the reconstruction order. The main effect of this technique is that allowable reconstructed vectors are more densely packed in an N-dimensional vector space (N represents the number of transform coefficients in a transform block) compared to traditional independent scalar quantization as used in HEVC. That is, for a given average number of allowable reconstructed vectors per N-dimensional unit volume, the average distortion between the input vector and the nearest reconstructed vector is reduced. The dependent scalar quantization technique is realized by (a) defining two scalar quantizers with different reconstruction levels and (b) defining a process for switching between the two scalar quantizers.
[0066] The two scalar quantizers used are shown in Figure 9, denoted Q0 and Q1. The position of the available reconstruction levels is uniquely specified by the quantization step size Δ. The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the transform coefficient level preceding the current transform coefficient in the coding / reconstruction order.
[0067] Figure 9 shows a diagram of the two scalar quantizers used in the proposed dependent quantization technique.
[0068] As shown in Figures 10A and 10B, switching between two scalar quantizers (Q0 and Q1) is achieved via a state machine with four quantizer states (QState). QState can take on four different values: 0, 1, 2, and 3, which are uniquely determined by the parity of the transform coefficient level preceding the current transform coefficient in the encoding / reconstruction order. At the start of dequantization of a transform block, the state is set equal to 0. The transform coefficients are reconstructed in scan order (i.e., in the same order they are entropy decoded). After the current transform coefficient is reconstructed, the state is updated as shown in Figure 10, where k denotes the value of the transform coefficient level.
[0069] FIG. 10A shows a transition diagram illustrating the state transitions for the proposed dependent quantization.
[0070] FIG. 10B shows a table illustrating the quantizer selection for the proposed dependent quantization.
[0071] Signaling a default user-defined scaling matrix is also supported. The scaling matrices for DEFAULT mode are all flat and have elements equal to 16 for all TB sizes. IBC and intra coding modes currently share the same scaling matrix. Therefore, for a USER_DEFINED matrix, the MatrixType and MatrixType_DC numbers are updated as follows:
[0072] MatrixType: 30 = 2 (2 for Intra & IBC / Inter) x 3 (Y / Cb / Cr components) x 5 (Square TB size: 4 x 4 to 64 x 64 for Luminance, 2 x 2 to 32 x 32 for Saturation)
[0073] MatrixType_DC: 14 = 2 (2 for Intra & IBC / Inter × 1 for Y component) × 3 (TB size: 16 × 16, 32 × 32, 64 × 64) + 4 (2 for Intra & IBC / Inter × 2 for Cb / Cr component) × 2 (TB size: 16 × 16, 32 × 32)
[0074] DC values are coded separately for the following scaling matrices: 16x16, 32x32, and 64x64. For TBs smaller than 8x8, all elements in one scaling matrix are signaled. For TBs larger than 8x8, only 64 elements in one 8x8 scaling matrix are signaled as the base scaling matrix. To obtain square matrices larger than 8x8, the 8x8 base scaling matrix is upsampled (by element duplication) to the corresponding square size (i.e., 16x16, 32x32, 64x64). When zeroing out high-frequency coefficients of a 64-point transform is applied, the corresponding high frequencies in the scaling matrix are also zeroed out. That is, if the width or height of the TB is 32 or greater, only the left or upper half of the coefficients are retained, and the remaining coefficients are assigned zeros. Furthermore, because the bottom-right 4x4 element is never used, the number of elements signaled for a 64x64 scaling matrix is also reduced from 8x8 to three 4x4 submatrices.
[0075] Context modeling for transform coefficient coding. The choice of the probability model for the syntax elements related to the absolute values of the transform coefficient levels depends on the values of the absolute levels or partially reconstructed absolute levels in a local neighborhood. The template used is illustrated in Figure 11.
[0076] Figure 11 shows a diagram of the template used to select a probabilistic model. The black square designates the current scan position, and the square with an "x" represents the local neighborhood used.
[0077] The probability model selected depends on the sum of absolute levels (or partially reconstructed absolute levels) within the local neighborhood, and the number of absolute levels greater than 0 within the local neighborhood (given by the number of sig_coeff_flags equal to 1). Context modeling and binarization are based on the following measurements about the local neighborhood: numSig: number of non-zero levels in the local neighborhood, sumAbs1: the sum of the partially reconstructed absolute levels (absLevel1) after the first pass within the local neighborhood, sumAbs: the sum of the reconstructed absolute levels within the local neighborhood, Diagonal position (d): The sum of the horizontal and vertical coordinates of the current scan position within the transformation block Depends on.
[0078] Based on the values of numSig, sumAbs1, and d, a probability model is selected for encoding sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag. The Rice parameters for binarizing abs_remainder and dec_abs_level are selected based on the values of sumAbs and numSig.
[0079] In the current VVC, the reduced 32-point MTS (also known as RMTS32) is based on skipping high-frequency coefficients and is used to reduce the computational complexity of 32-point DST-7 / DCT-8. It also involves a change to coefficient coding, including all types of zeroing (i.e., RMTS32 and the existing zeroing for high-frequency components in DCT2). Specifically, binarization of the last non-zero coefficient position is coded based on the reduced TU size, and the context model selection for the last non-zero coefficient position is determined by the original TU size. Furthermore, 60 context models are used to code the sig_coeff_flag of the transform coefficients. The context model index selection is based on the sum of five previously partially reconstructed absolute levels, called locSumAbsPass1, and the dependent quantization state QState, as follows:
[0080] If cIdx is equal to 0, ctxInc is derived as follows: ctxInc=12*Max(0,QState-1)+Min((locSumAbsPass1+1)>>1,3)+(d<2?8:(d<5?4:0))
[0081] Otherwise (if cIdx is greater than 0), ctxInc is derived as follows: ctxInc=36+8*Max(0,QState-1)+Min((locSumAbsPass1+1)>>1,3)+(d<2?4:0)
[0082] Palette Mode The basic idea behind palette mode is that the samples in a CU are represented by a small set of representative color values. This set is called the palette. Color values that are excluded from the palette are those for the three color components. Value of It is also possible to indicate colors that are excluded from the palette by signaling them as escape colors that are signaled directly in the bitstream. This is illustrated in Figure 12.
[0083] An example of a block coded in palette mode is shown in Figure 12. Figure 12 includes a block 1210 coded in palette mode and a palette 1220.
[0084] In Figure 12, the palette size is 4. The first three samples use palette entries 2, 0, and 3, respectively, for reconstruction. black The color samples represent escape symbols. The CU-level flag palette_escape_val_present_flag indicates whether escape symbols are present in the CU. If escape symbols are present, the palette size is increased by 1 and the last index is used to represent the escape symbol. Thus, in Figure 12, index 4 is assigned to the escape symbol.
[0085] To decode a palette-coded block, the decoder needs the following information: Pallet table, Pallet Index It is necessary to have
[0086] If the palette index corresponds to an escape symbol, additional overhead is signaled to indicate the corresponding color value of the sample.
[0087] Additionally, the encoder needs to derive the appropriate palette to be used for that CU.
[0088] For palette derivation for lossy coding, a modified k-means clustering algorithm is used. The first sample of a block is added to the palette. Then, for each subsequent sample from the block, the sum of absolute differences (SAD) between the sample and each of the current palette colors is calculated. If the distortion of each component is less than the threshold for the palette entry corresponding to the smallest SAD, the sample is added to the cluster belonging to that palette entry. Otherwise, the sample is added as a new palette entry. When the number of samples mapped to a cluster exceeds the threshold, the centroid of that cluster is updated to become the palette entry for that cluster.
[0089] In the next step, the clusters are sorted in descending order of usage. The palette entry corresponding to each entry is then updated. Typically, the cluster centroid is used as the palette entry. However, given the cost of encoding palette entries, a rate-distortion analysis is performed to analyze whether any of the entries from the palette predictor, rather than the centroid, might be suitable for use as the updated palette entry. This process continues until all clusters have been processed or the maximum palette size is reached. Finally, if a cluster has only a single sample and no corresponding palette entry in the palette predictor, the sample is converted to an escape symbol. Furthermore, duplicate palette entries are removed, and the clusters are merged.
[0090] After palette derivation, each sample in the block is assigned the index of the nearest palette entry (in the SAD). The sample is then assigned to either "INDEX" or "COPY_ABOVE" mode. For each sample that can have "INDEX" or "COPY_ABOVE" mode, the cost of encoding the mode is then calculated. The mode with the lower cost is selected.
[0091] A palette predictor is maintained for encoding palette entries. The maximum size of the palette and the palette predictor are signaled in the SPS. The palette predictor is initialized at the beginning of each CTU row, each slice, and each tile.
[0092] For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. This is illustrated in Figure 13.
[0093] Figure 13 illustrates the use of a palette predictor to signal palette entries, including a previous palette 1310 and a current palette 1320.
[0094] The reuse flag is signaled using run-length coding of zeros. After this, the number of new palette entries is signaled using an exponential-Golomb code of degree 0. Finally, the component values of the new palette entries are signaled.
[0095] The palette index is coded using horizontal and vertical translation scan as shown in Figures 14A and 14B. The scan order is explicitly signaled in the bitstream using palette_transpose_flag.
[0096] FIG. 14A shows horizontal motion scanning.
[0097] FIG. 14B shows a vertical translation scan.
[0098] To encode the palette index, a line coefficient group (CG)-based palette mode is used, which divides the CU into multiple segments with 16 samples each based on the moving scan mode as shown in Figures 15A and 15B. The index run, palette index value, and quantized color for escape mode are encoded / analyzed sequentially for each CG.
[0099] FIG. 15A illustrates a sub-block based index map scan for a palette.
[0100] FIG. 15B illustrates a sub-block based index map scan for the palette.
[0101] Palette indices are coded using two main palette sample modes: "INDEX" and "COPY_ABOVE". As explained previously, escape symbols are assigned an index equal to the maximum palette size. In "COPY_ABOVE" mode, the palette index of the sample in the row above is duplicated. In "INDEX" mode, the palette index is explicitly signaled. The coding order of the palette run coding in each segment is as follows:
[0102] For each pixel, one context coding bin run_copy_flag=0 is signaled to indicate whether the pixel is in the same mode as the previous pixel, i.e., whether the previous scanned pixel and the current pixel are both of run type COPY_ABOVE, or whether the previous scanned pixel and the current pixel are both of run type INDEX and have the same index value; otherwise, run_copy_flag=1 is signaled.
[0103] If a pixel and the previous pixel are in different modes, one context coding bin copy_above_palette_indices_flag is signaled to indicate the execution type of the pixel, i.e., INDEX or COPY_ABOVE. INDEX mode is used by default, so the decoder does not need to analyze the execution type if the sample is in the first row (horizontal translation scan) or first column (vertical translation scan). Also, if the previously analyzed execution type is COPY_ABOVE, the decoder does not need to analyze the execution type. After palette execution coding of pixels in one segment, the index value (palette_idx_idc) and quantized escape color (palette_escape_val) for INDEX mode are bypass coded.
[0104] Improvements to residual and coefficient coding In VVC, when encoding transform coefficients, a unified (same) derivation of the Rice parameter (RicePara) is used to signal the syntax of abs_remainder and dec_abs_level. The only difference is that the base level baseLevel is set to 4 and 0 to encode abs_remainder and dec_abs_level, respectively. The Rice parameter is determined based not only on the sum of the absolute levels of the five adjacent transform coefficients in the local template, but also on the corresponding base levels as follows:
[0105] RicePara=RiceParTable[max(min(31,sumAbs-5*baseLevel),0)]
[0106] In other words, the binary codewords of the syntax elements abs_remainder and dec_abs_level are adaptively determined according to the level information of the neighboring coefficients. Because this codeword determination is performed sample by sample, additional logic is required to handle this codeword adaptation for coefficient coding.
[0107] Similarly, when encoding a residual block under transform skip mode, the binary codeword of the syntax element abs_remainder is adaptively determined according to the level information of the neighboring residual samples.
[0108] Furthermore, when coding syntax elements related to residual coding or transform coefficient coding, the choice of probability model depends on the level information of adjacent levels, which requires additional logic and additional context models.
[0109] In the current design, the binarization of escape samples is derived by invoking a cubic exponential-Golomb binarization process, the performance of which leaves room for further improvement.
[0110] In the current VVC, two different level mapping schemes are available, which are applied to the normal transform and the transform skip, respectively. Each level mapping scheme is associated with different conditions, mapping functions, and mapping positions. For blocks to which the normal transform is applied, one level mapping scheme is used after the number of context coding bins (CCBs) exceeds the limit. The mapping position, denoted as ZeroPos[n], and the mapping result, denoted as AbsLevel[xC][yC], are derived as specified in Table 2. For blocks to which the transform skip is applied, another level mapping scheme is used before the number of context coding bins (CCBs) exceeds the limit. The mapping position, denoted as predCoeff, and the mapping result, denoted as AbsLevel[xC][yC], are derived as specified in Table 5. Such a non-uniform design may not be optimal from a standardization perspective.
[0111] For profiles greater than 10 bits in HEVC, extended_precision_processing_flag equal to 1 specifies that extended dynamic range is used for coefficient analysis and inverse transform processing. In the current VVC, residual coding or transform skip coding of transform coefficients greater than 10 bits has been reported to cause significant performance degradation. There is room for further improvement in the performance.
[0112] Proposed method In this disclosure, several methods are proposed to address the issues mentioned in the section on improvements to residual and coefficient coding. It should be noted that the following methods may be applied alone or in combination.
[0113] According to a first aspect of the present disclosure, it is proposed to use a fixed set of binary codewords to encode a specific syntax element, such as abs_remainder, in residual coding. The binary codewords can be formed using various methods. Some exemplary methods are listed as follows:
[0114] First, to determine the codeword for abs_remainder, the same procedure as that used in current VVC is used, but a fixed Rice parameter (e.g., 1, 2, or 3) is always chosen.
[0115] Second, fixed-length binarization.
[0116] Third, there is truncated Rice binarization.
[0117] Fourth, a truncated binary (TB) binarization process.
[0118] Fifth, the kth exponential-Golomb binarization process (EGk).
[0119] Sixth, there is the restricted kth-order exponential-Golomb binarization.
[0120] According to a second aspect of the present disclosure, it is proposed to use a fixed set of code words to encode certain syntax elements, such as abs_remainder and dec_abs_level, in transform coefficient coding. The binary code words can be formed using various methods. Some exemplary methods are listed as follows:
[0121] First, to determine the code words for abs_remainder and dec_abs_level, the same procedure as used in current VVC is used, but with a fixed Rice parameter, e.g., 1, 2, or 3. As used in current VVC, the value of baseLevel can still be different for abs_remainder and dec_abs_level (e.g., to encode abs_remainder and dec_abs_level, baseLevel is set to 4 and 0, respectively).
[0122] Second, to determine the code words for abs_remainder and dec_abs_level, the same procedure as currently used in VVC is used, but with a fixed Rice parameter, e.g., 1, 2, or 3. The values for baseLevel for abs_remainder and dec_abs_level are chosen to be the same, e.g., both use 0 or both use 4.
[0123] Third, fixed-length binarization.
[0124] Fourth, truncated Rice binarization.
[0125] Fifth, the truncated binary (TB) binarization process.
[0126] Sixth, the kth exponential-Golomb binarization process (EGk).
[0127] Seventh, restricted k-th order exponential-Golomb binarization.
[0128] According to a third aspect of the present disclosure, it is proposed to use a single context for coding syntax elements related to residual coding or coefficient coding (e.g., abs_level_gtx_flag), and context selection based on neighboring decoded level information may be eliminated.
[0129] According to a fourth aspect of the present disclosure, it is proposed to use a variable set of binary codewords to code a specific syntax element, such as abs_remainder, in residual coding, and the selection of the set of binary codewords is determined according to specific coded information of a current block, such as a quantization parameter (QP) associated with TB / CB and / or slice, a prediction mode of a CU (e.g., IBC mode or intra or inter), and / or a slice type (e.g., I slice, P slice, or B slice). Various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed as follows:
[0130] First, to determine the codeword for bs_remainder, the same procedure as that used in current VVC is used, but with different Rice parameters.
[0131] Second, the kth order Exponential-Golomb Binarization Process (EGk).
[0132] Third, there is the restricted kth-order exponential-Golomb binarization. [Table 6]
[0133] The same method as described in the fourth embodiment can be used to convert coefficientIt is also applicable to symbolization. According to the fifth aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode specific syntax elements, such as abs_remainder and dec_abs_level, in transform coefficient coding. The selection of the set of binary codewords is determined according to specific encoded information of the current block, such as quantization parameters (QP) associated with TB / CB and / or slices, the prediction mode of the CU (e.g., IBC mode or intra or inter), and / or slice type (e.g., I slice, P slice, or B slice). Also in this case, various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed as follows.
[0134] First, the same procedure as that used in the current VVC is used to determine the codeword of bs_remainder, but different Rice parameters are used.
[0135] Second, it is the k-th exponential Golomb binary process (EGk).
[0136] Third, it is limited k-th exponential Golomb binary.
[0137] In these above methods, different Rice parameters may be used to derive different sets of binary codewords. For a given block of residual samples, the Rice parameter used is determined according to the CU QP shown as QPCU, rather than adjacent level information. As shown in Table 6, a specific example is illustrated, where TH1 to TH4 are predetermined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predetermined Rice parameters. It is worth noting that the same logic can be implemented in different ways in practice. For example, a specific equation or look-up table may also be used to derive the same Rice parameter from the QP value of the current CU as shown in Table 6.
[0138] According to a fifth aspect of the present disclosure, a set of parameters and / or thresholds related to codeword determination for a syntax element of transform coefficient coding and / or transform skip residual coding is signaled in a bitstream, and the determined codeword is used as a binarized codeword when encoding the syntax element through an entropy coder, e.g., arithmetic coding.
[0139] It should be noted that the set of parameters and / or thresholds may be a complete set or a subset of all parameters and thresholds related to codeword decisions of the syntax element. The set of parameters and / or thresholds may be signaled at various levels in the video bitstream. For example, they may be signaled at the sequence level (e.g., sequence parameter set) at the coding tree unit (CTU) level, the picture level (e.g., picture parameter set and / or picture header), the slice level (e.g., slice header), or at the coding unit (CU) level.
[0140] In one example, the Rice parameters used to determine the codewords for encoding the abs_remainder syntax in transform skip residual coding are signaled in a slice header, a picture header, a PPS, and / or an SPS. The signaled Rice parameters are used to determine the codewords for encoding the abs_remainder syntax when the CU is coded as a transform skip mode and the CU is associated with the slice header, picture header, PPS, and / or SPS described above.
[0141] According to a sixth aspect of the present disclosure, the set of parameters and / or thresholds related to code word determination shown in the first and second aspects is used for the syntax elements of transform coefficient coding and / or transform skip residual coding. Also, different sets may be used depending on whether the current block contains luma residual / coefficients or chroma residual / coefficients. The determined code word is used as a binarized code word when encoding the syntax elements through an entropy coder, for example, arithmetic coding.
[0142] In one example, the abs_remainder codeword associated with the transform residual coding used in current VVC is used for both luma and chroma blocks, but different fixed Rician parameters are used by the luma and chroma blocks, respectively (e.g., K1 for luma blocks and K2 for chroma blocks, where K1 and K2 are integers).
[0143] According to a seventh aspect of the present disclosure, a set of parameters and / or thresholds related to codeword determination for a syntax element of transform coefficient coding and / or transform skip residual coding is signaled in a bitstream. Different sets may be signaled for luma blocks and chroma blocks. The determined codewords are used as binarized codewords when encoding the syntax element through an entropy coder, e.g., arithmetic coding.
[0144] The same method as described in the above embodiment is also applicable to escape value encoding in palette mode, eg, palette_escape_val.
[0145] According to an eighth aspect of the present disclosure, different k-th order Exponential-Golomb binarizations may be used to derive different sets of binary code words for encoding escape values in palette mode. In one example, for a given block of escape samples, the Exponential-Golomb parameter used, i.e., value k, is determined by the QP CUIn deriving the value of the parameter k based on a given QP value of the block, the same example as shown in Table 6 can be used. In that example, four different thresholds (TH1 to TH4) are listed, and the QPs of these thresholds and CU It is worth noting that five different k values (K0 to K4) can be derived based on k = 1 / k, but the number of thresholds is for illustrative purposes only. In practice, a different number of thresholds may be used to partition the entire QP value range into a different number of QP value segments, and for each QP value segment, a different k value may be used to derive a corresponding binary code word for encoding an escape value for a block coded in palette mode. It is also worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may be used to derive the same Rice parameter.
[0146] According to a ninth aspect of the present disclosure, a set of parameters and / or thresholds related to code word determination for a syntax element of an escape sample are signaled in a bitstream, and the determined code word is used as a binarized code word when encoding the syntax element through an entropy coder, e.g., arithmetic coding.
[0147] It should be noted that the set of parameters and / or thresholds may be a complete set or a subset of all parameters and thresholds related to codeword decisions of the syntax element. The set of parameters and / or thresholds may be signaled at various levels in the video bitstream. For example, they may be signaled at the sequence level (e.g., sequence parameter set) at the coding tree unit (CTU) level, the picture level (e.g., picture parameter set and / or picture header), the slice level (e.g., slice header), or at the coding unit (CU) level.
[0148] In one example according to this embodiment, k-th order Exponential-Golomb binarization is used to determine a codeword for encoding the construct palette_escape_val in palette mode, and the value k is signaled to the decoder in the bitstream. The value k may be signaled at various levels, for example, the value k may be signaled in a slice header, a picture header, a PPS, and / or an SPS, etc. The signaled Exponential-Golomb parameter is used to determine a codeword for encoding the construct palette_escape_val when the CU is coded in palette mode and the CU is associated with the slice header, picture header, PPS, and / or SPS, etc., as described above.
[0149] Level mapping harmonization for transform skip mode and normal transform mode According to a tenth aspect of the present disclosure, the same condition for applying level mapping is used for both transform skip mode and normal transform mode. In one example, it is proposed to apply level mapping after the number of context coding bins (CCBs) exceeds a limit for both transform skip mode and normal transform mode. In another example, it is proposed to apply level mapping before the number of context coding bins (CCBs) exceeds a limit for both transform skip mode and normal transform mode.
[0150] According to an eleventh aspect of the present disclosure, the same method for deriving mapping positions in level mapping is used for both transform skip mode and normal transform mode. In one example, it is proposed that the method for deriving mapping positions in level mapping used under transform skip mode is also applied to normal transform mode. In another example, it is proposed that the method for deriving mapping positions in level mapping used under normal transform mode is also applied to transform skip mode.
[0151] According to a twelfth aspect of the present disclosure, the same level mapping method is applied to both the transform skip mode and the normal transform mode. In one example, it is proposed to apply the level mapping function used under the transform skip mode to the normal transform mode as well. In another example, it is proposed to apply the level mapping function used under the normal transform mode to the transform skip mode as well.
[0152] Simplifying the derivation of Rice parameters in residual coding According to a thirteenth aspect of the present disclosure, it is proposed to use simple logic, such as shift or division operations, instead of a lookup table to derive the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding. According to the present disclosure, the lookup table specified in Table 4 may be eliminated. In one example, the Rice parameter cRiceParam is derived as cRiceParam=(locSumAbs>>n), where n is a positive number, e.g., 3. It is worth noting that in practice, other different logic, such as division by a value equal to the nth power of 2, may be used to achieve the same result. An example of a corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font. [Table 7]
[0153] According to a fourteenth aspect of the present disclosure, it is proposed to use fewer adjacent positions for deriving the Rice parameters when encoding the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding. In one example, it is proposed to use only two adjacent positions for deriving the Rice parameters when encoding the syntax elements abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 8]
[0154] In another example, it is proposed to use only one adjacent position for the derivation of the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 9]
[0155] According to a fifteenth aspect of the present disclosure, for the derivation of the Rice parameters when encoding the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding, it is proposed to use a different logic for adjusting the value of locSumAbs based on the value of baseLevel. In one example, additional scale and offset operations are applied in the form of "(locSumAbs-BaseLevel*5)*alpha+beta". The corresponding decoding process based on the VVC draft when alpha takes the value 1.5 and beta takes the value 1 is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 10]
[0156] According to a sixteenth aspect of the present disclosure, it is proposed to remove the clipping operation for the Rice parameter derivation in the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding. According to the present disclosure, an example of the decoding process of the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 11]
[0157] According to the current disclosure, an example of the decoding process of the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 12]
[0158] According to a seventeenth aspect of the present disclosure, for the derivation of the Rice parameters when encoding the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding, it is proposed to change the initial value of locSumAbs from 0 to a non-zero integer. In one example, an initial value of 1 is assigned to locSumAbs, and the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 13]
[0159] According to an eighteenth aspect of the present disclosure, it is proposed to use the maximum value of adjacent position level values instead of the sum of adjacent position level values to derive the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 14]
[0160] According to a nineteenth aspect of the present disclosure, when encoding the syntax elements abs_remainder / dec_abs_level using Golomb-Rice coding, it is proposed to derive the Rice parameter based on the relative amplitude of each AbsLevel value at adjacent positions and the reference level value. In one example, the Rice parameter is derived based on how many of the AbsLevel values at adjacent positions are greater than the reference level. An example of a corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font. [Table 15]
[0161] In another example, the Rice parameter is derived based on the sum of the (AbsLevel-BaseLevel) values of neighboring positions where the AbsLevel value is greater than the base level. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 16]
[0162] According to the current disclosure, an example of the decoding process of the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 17]
[0163] Simplifying Level Mapping Position Derivation in Residual Coding According to a twentieth aspect of the present disclosure, it is proposed to remove QState from the derivation of ZeroPos[n], so that ZeroPos[n] is derived only from cRiceParam. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 18]
[0164] According to a 21st aspect of the present disclosure, it is proposed to derive ZeroPos[n] based on the value of locSumAbs. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 19]
[0165] According to a 22nd aspect of the present disclosure, it is proposed to derive ZeroPos[n] based on the AbsLevel values of adjacent positions. In one example, ZeroPos[n] is derived based on the maximum value of AbsLevel[xC+1][yC] and AbsLevel[xC][yC+1]. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 20]
[0166] According to a 23rd aspect of the present disclosure, it is proposed to derive both cRiceParam and ZeroPos[n] based on the maximum of all AbsLevel values at adjacent positions. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font: [Table 21]
[0167] The same method as described in the above embodiment is also applicable to deriving predCoeff in transform skip mode residual coding. In one example, the variable predCoeff is derived as follows:
[0168] predCoeff=Max(absLeftCoeff,absAboveCoeff)+1
[0169] Residual coding of transform coefficients In this disclosure, methods are provided to simplify and / or further improve existing designs of residual coding to address the problems pointed out in the "Improvements to Residual and Coefficient Coding" section. Generally, the main features of the techniques proposed in this disclosure are summarized as follows:
[0170] First, we adjust the Rice parameter derivation used under conventional residual coding based on the current design.
[0171] Second, we modify the binarization method used under normal residual coding.
[0172] Third, we modify the Rice parameter derivation used under normal residual coding.
[0173] Rice parameter derivation in residual coding based on the current design According to a 24th aspect of the present disclosure, it is proposed to use a variable method of Rice parameter derivation to encode a specific syntax element, such as abs_remainder / dec_abs_level, in residual encoding, and according to a quantization parameter or encoding bit depth associated with specific encoded information of the current block, such as TB / CB and / or slice / profile, and / or according to a new flag associated with the TB / CB / slice / picture / sequence level, such as extended_precision_processing_flag, the selection result is determined. Various methods may be used to derive the Rice parameter, and some exemplary methods are listed below.
[0174] First, cRiceParam = (cRiceParam << a) + (cRiceParam >> b) + c, where a, b, and c are positive numbers, for example, {a, b, c} = {1, 1, 0}. It is worth noting that in practice, other different logics, such as multiplication operations with values equal to 2 to the power of n, may be used to achieve the same result.
[0175] Second, cRiceParam = (cRiceParam << a) + b, where a and b are positive numbers, for example, {a, b} = {1, 1}. It is worth noting that in practice, other different logics, such as multiplication operations with values equal to 2 to the power of n, may be used to achieve the same result.
[0176] Third, cRiceParam = (cRiceParam * a) + b, where a and b are positive numbers, for example, {a, b} = {1.5, 0}. It is worth noting that in practice, other different logics, such as multiplication operations with values equal to 2 to the power of n, may be used to achieve the same result.
[0177] An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deletions indicated in italic font. The changes to the VVC draft are shown in bold italic font in Table 22. It is worth noting that in practice the same logic can be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters from the BitDepth value of the current CU / sequence. [Table 22] JPEG2026035734000034.jpg81165
[0178] In another example, when BitDepth is equal to or greater than a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the rice parameter cRiceParam is set as follows: cRiceParam=(cRiceParam<<a)+(cRiceParam> >b)+c, where a, b, and c are positive numbers, e.g., 1. The corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font. The changes relative to the VVC draft are shown in bold italic font in Table 23. It is worth noting that in practice the same logic can be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter from the BitDepth value of the current CU / sequence. [Table 23] JPEG2026035734000036.jpg72164
[0179] A binary method for residual coding of profiles greater than 10 bits. According to a 25th aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode a specific syntax element, e.g., abs_remainder / dec_abs_level, in residual coding, and the selection result is determined according to the quantization parameter or coding bit depth associated with specific coded information of the current block, e.g., TB / CB and / or slice / profile, and / or according to a new flag associated with the TB / CB / slice / picture / sequence level, e.g., extended_precision_processing_flag. Various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed below.
[0180] First, to determine the codeword of abs_remainder, the same procedure as that currently used in VVC is used, but a fixed Rice parameter (e.g., 2, 3, 4, 5, 6, 7, or 8) is always selected. The fixed value may vary under different conditions according to the quantization parameter or coding bit depth associated with specific coded information of the current block, e.g., TB / CB and / or slice / profile, and / or according to a syntax element associated with the TB / CB / slice / picture / sequence level, e.g., rice_parameter_value. As shown in Table 24, a specific example is illustrated, where TH1 to TH4 are predetermined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predetermined Rice parameters. It is worth noting that the same logic can be implemented in different ways in practice. For example, a specific equation or look-up table may also be used to derive the same Rice parameter from the BitDepth value of the current CU / sequence as shown in Table 24.
[0181] Second, it is fixed-length binarization.
[0182] Third, it is truncated Rice binarization.
[0183] Fourth, the truncated binary (TB) binarization process.
[0184] Fifth, the kth exponential-Golomb binarization process (EGk).
[0185] Sixth, there is the restricted kth-order exponential-Golomb binarization. [Table 24]
[0186] In one example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, the rice parameter cRiceParam is fixed as n, where n is a positive number (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value may be different in different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font. The changes to the VVC draft are shown in bold italic font in Table 25. [Table 25]
[0187] In another example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, it is proposed to use only one fixed value for the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, with changes shown in bold italic font and deleted content in italic font. The changes relative to the VVC draft are shown in bold italic font in Table 26. [Table 26]
[0188] In yet another example, when BitDepth is greater than or equal to a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is fixed as n, where n is a positive number, e.g., 4, 5, 6, 7, or 8. The fixed value may be different in different conditions. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16). The changes to the VVC draft are shown in bold italic font in Table 27, with the modified content in bold italic and the deleted content in italic. [Table 27]
[0189] In yet another example, it is proposed to use only one fixed value for the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level when BitDepth is greater than a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16). The corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), changes are shown in bold italic font, and deleted content is shown in italic font. The changes relative to the VVC draft are shown in bold italic font in Table 28. [Table 28]
[0190] Rice parameter derivation in residual coding According to a 26th aspect of the present disclosure, it is proposed to use variable methods of Rice parameter derivation to code specific syntax elements, such as abs_remainder / dec_abs_level, in residual coding, and the selection result is determined according to specific coded information of the current block, such as the quantization parameter or coding bit depth associated with TB / CB and / or slice / profile, and / or according to a new flag, such as extended_precision_processing_flag, associated with TB / CB / slice / picture / sequence level. Various methods may be used to derive the Rice parameters, and some exemplary methods are listed below.
[0191] First, it is proposed to use a counter to derive the Rice parameter. The counter is determined according to the value of the coded coefficient and the specific coded information of the current block, such as the component ID. A specific example is riceParameter=counter / a, where a is a positive number, such as 4, which maintains two counters (resolved by luma / chroma). These counters are reset to 0 at the beginning of each slice. Upon coding, the counter is updated as follows if this is the first coefficient coded within a sub-TU: if(coeffValue>=(3< <rice))counter++ if(((coeffValue<<1)<(1<<riceParameter))&&(counter> 0))counter--
[0192] Second, it is proposed to add a shift operation in the derivation of the Rice parameters in VVC. The shift is determined according to the value of the coded coefficient. An example of the corresponding decoding process based on the VVC draft is shown below, where the shift is determined according to the counter of Method 1, and the changes are shown in bold italic font and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 29. [Table 29]
[0193] First, it is proposed to add a shift operation in the derivation of the Rice parameters in VVC. The shift is determined according to the specific coded information of the current block, such as the TB / CB and / or the coding bit depth associated with the slice profile (e.g., 14-bit profile or 16-bit profile). An example of the corresponding decoding process based on the VVC draft is shown below, where the shift is determined according to the counters in Method 1, with changes indicated in bold italic font and deleted content indicated in italic font. The changes to the VVC draft are shown in bold italic font in Table 30. [Table 30]
[0194] Transform skip residual coding According to a 27th aspect of the present disclosure, it is proposed to use a variable set of binary codewords to code a specific syntax element, such as abs_remainder, in transform skip residual coding, and the selection result is determined according to specific coded information of the current block, such as a quantization parameter or coding bit depth associated with TB / CB and / or slice / profile, and / or according to a new flag associated with TB / CB / slice / picture / sequence level, such as extended_precision_processing_flag. Various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed below.
[0195] First, to determine the sign word of abs_remainder, the same procedure as currently used in VVC is used, but a fixed Rice parameter (e.g., 2, 3, 4, 5, 6, 7, or 8) is always selected. The fixed value may vary under different conditions according to specific encoded information of the current block, such as the quantization parameter associated with the TB / CB and / or slice / profile, frame type (e.g., I, P, or B), component ID (e.g., luminance or chrominance), color format (e.g., 420, 422, or 444), or encoded bit depth, and / or according to syntax elements associated with the TB / CB / slice / picture / sequence level, such as rice_parameter_value. As shown in Table 7, a specific example is illustrated, where TH1 to TH4 are predetermined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predetermined Rice parameters. It is worth noting that the same logic can be implemented in different ways in practice. For example, a specific equation or look-up table may also be used to derive the same Rice parameter from the BitDepth value of the current CU / sequence as shown in Table 7.
[0196] Second, it is fixed-length binarization.
[0197] Third, it is truncated Rice binarization.
[0198] Fourth, it is a truncated binary (TB) binarization process.
[0199] Fifth, it is a k-th exponential Golomb binarization process (EGk).
[0200] Sixth, it is limited k-th exponential Golomb binarization.
[0201] An example of a corresponding decoding process based on the VVC draft is shown below, with changes to the VVC draft shown in bold italic font in Table 31 and deletions shown in italic font. It is worth noting that in practice the same logic may be implemented in different ways. For example, specific equations or lookup tables may also be used to derive the same Rice parameters. [Table 31]
[0202] In another example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, it is proposed to use only one fixed value for the Rice parameter when encoding the abs_remainder syntax element. The corresponding decoding process based on the VVC draft is shown below, with changes shown in bold italic font and deleted content shown in italic font. The changes relative to the VVC draft are shown in bold italic font in Table 32. [Table 32]
[0203] In yet another example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, the rice parameter cRiceParam is fixed to n, where n is a positive number (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value may be different under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, with changes indicated in bold italic font and deleted content indicated in italic font. The changes to the VVC draft are shown in bold italic font in Table 33. [Table 33]
[0204] In yet another example, when BitDepth is greater than or equal to a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is fixed to n, where n is a positive number, e.g., 4, 5, 6, 7, or 8. The fixed value may be different under different conditions. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), and changes are shown in bold italic font and deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 34. [Table 34]
[0205] In yet another example, one control flag is signaled in the slice header to indicate whether signaling of the Rice parameter for transform skip blocks is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled for each transform skip slice to indicate the Rice parameter for that slice. When the control flag is signaled as disabled (e.g., set equal to “0”), no further syntax elements are signaled at lower levels to indicate the Rice parameter for transform skip slices, and a default Rice parameter (e.g., 1) is used for all transform skip slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2), changes are shown in bold italic font, and deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 35. It is worth noting that sh_ts_residual_coding_rice_index can be coded in various ways and / or may have a maximum value. For example, to encode / decode the same syntax element, one may also use u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right).
[0206] Slice Header Syntax [Table 35]
[0207] sh_ts_residual_coding_rice_flag equal to 1 specifies that sh_ts_residual_coding_rice_index may be present in the current slice. sh_ts_residual_coding_rice_flag equal to 0 specifies that sh_ts_residual_coding_rice_index is not present in the current slice. When sh_ts_residual_coding_rice_flag is not present, the value of sh_ts_residual_coding_rice_flag is inferred to be equal to 0. sh_ts_residual_coding_rice_index specifies the rice parameter used in the residual_ts_coding() syntax construct. [Table 36]
[0208] In yet another example, one control flag is signaled in the sequence parameter set (or in the sequence parameter set range extension syntax) to indicate whether signaling of Rice parameters for transform skip blocks is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled for each transform skip slice to indicate the Rice parameter for that slice. When the control flag is signaled as disabled (e.g., set equal to “0”), no further syntax elements are signaled at lower levels to indicate Rice parameters for transform skip slices, and a default Rice parameter (e.g., 1) is used for all transform skip slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). Changes to the VVC draft are shown in bold italics in Table 37, and deleted content is italicized. It is worth noting that sh_ts_residual_coding_rice_idx can be coded in various ways and / or may have a maximum value. For example, to encode / decode the same syntax element, one may also use u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right).
[0209] Sequence Parameter Set RBSP Syntax [Table 37]
[0210] sps_ts_residual_coding_rice_present_in_sh_flag equal to 1 specifies that sh_ts_residual_coding_rice_idx may be present in SH syntax structures that reference an SPS. sps_ts_residual_coding_rice_present_in_sh_flag equal to 0 specifies that sh_ts_residual_coding_rice_idx is not present in SH syntax structures that reference an SPS. When sps_ts_residual_coding_rice_present_in_sh_flag is not present, the value of sps_ts_residual_coding_rice_present_in_sh_flag is inferred to be equal to 0.
[0211] Slice Header Syntax [Table 38]
[0212] sh_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax construct. [Table 39]
[0213] In one or more examples of the present disclosure, as shown in Figure 27, it is proposed to disable the presence of Rice parameters for transform skip residual coding if transform skip is disabled in step 2702. In one specific example, to achieve such a design objective, it is proposed to use sps_transform_skip_enabled_flag to condition the presence of sps_ts_residual_coding_rice_present_in_sh_flag, as shown in step 2704. For example, when the flag sps_transform_skip_enabled_flag is equal to 0 (i.e., when transform skip is disabled in the current picture), sps_ts_residual_coding_rice_present_in_sh_flag is inferred to be 0 rather than signaled. When the flag sps_transform_skip_enabled_flag is equal to 1, sps_ts_residual_coding_rice_present_in_sh_flag is further signaled. Changes to the current VVC Working Draft are shown in italic font below. [Table 40]
[0214] In yet another example, when the transform skip flag (sps_transform_skip_enabled_flag) is signaled as enabled, one control flag is further signaled in the sequence parameter set (or in the sequence parameter set range extension syntax) to indicate whether signaling of Rice parameters for transform skip blocks is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled per transform skip slice to indicate the Rice parameter for that slice. When the control flag is signaled as disabled (e.g., set equal to "0"), no further syntax elements are signaled at lower levels to indicate the Rice parameters for the transform skip slices, and the default Rice parameter (e.g., 1) is used for all transform skip slices. An example of a corresponding decoding process based on the VVC draft is shown below. Changes to the VVC draft are shown in italic font.
[0215] Sequence Parameter Set RBSP Syntax [Table 41]
[0216] sps_ts_residual_coding_rice_present_in_sh_flag equal to 1 specifies that sh_ts_residual_coding_rice_idx may be present in SH syntax structures that reference an SPS. sps_ts_residual_coding_rice_present_in_sh_flag equal to 0 specifies that sh_ts_residual_coding_rice_idx_minus1 is not present in SH syntax structures that reference an SPS. When sps_ts_residual_coding_rice_present_in_sh_flag is not present, the value of sps_ts_residual_coding_rice_present_in_sh_flag is inferred to be equal to 0.
[0217] Slice Header Syntax [Table 42]
[0218] sh_ts_residual_coding_rice_idx_minus1+1 specifies the rice parameter used in the residual_ts_coding() syntax construct. When sh_ts_residual_coding_rice_idx_minus1 is not present, the value of sh_ts_residual_coding_rice_idx_minus1 is inferred to be equal to 0.
[0219] 9.3.3.11 abs_remainder[] thresholding process
[0220] The inputs to this process are the syntax element abs_remainder[n], the color component cIdx, the current sub-block index i, the luma position (x0,y0) specifying the top-left sample of the current luma transform block relative to the top-left luma sample of the picture, the current coefficient scan position (xC,yC), the binary logarithm of the transform block width log2TbWidth, and a request to binarize the binary logarithm of the transform block height log2TbHeight.
[0221] The output of this process is a binarization of the syntax elements.
[0222] The variables lastAbsRemainder and lastRiceParam are derived as follows:
[0223] - If this process is called for the first time for the current subblock index i, then lastAbsRemainder and lastRiceParam are both set equal to 0.
[0224] - Otherwise (if this is not the first time this process is called for the current subblock index i), lastAbsRemainder and lastRiceParam are set equal to the values of abs_remainder[n] and cRiceParam, respectively, derived during the last call of the binarization process for the syntax element abs_remainder[n] specified in this section.
[0225] The rice parameter cRiceParam is derived as follows:
[0226] - If transform_skip_flag[x0][y0][cIdx] is equal to 1 and sh_ts_residual_coding_disabled_flag is equal to 0, the rice parameter cRiceParam is set to be equal to sh_ts_residual_coding_rice_idx_minus1+1.
[0227] - Otherwise, the Rice parameter cRiceParam is derived by invoking the Rice parameter derivation process in abs_remainder[] as specified in section 9.3.3.2 with the variable baseLevel set equal to 4, the color component index cIdx, the luminance position (x0,y0), the current coefficient scan position (xC,yC), the binary logarithm of the transform block width log2TbWidth, and the binary logarithm of the transform block height log2TbHeight as inputs.
[0228] In yet another example, one syntax element is signaled for each transform skip slice to indicate the Rice parameter for that slice. An example of the corresponding decoding process based on the VVC draft is shown below. The changes to the VVC draft are shown in bold italic font in Table 40. It is worth noting that sh_ts_residual_coding_rice_idx can be coded in various ways and / or may have a maximum value. For example, u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right), may also be used to code / decode the same syntax element.
[0229] Slice Header Syntax [Table 43]
[0230] sh_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax construct. If sh_ts_residual_coding_rice_idx is not present, the value of sh_ts_residual_coding_rice_idx is inferred to be equal to 0. [Table 44]
[0231] In yet another example, one control flag is signaled in the picture parameter set range extension syntax to indicate whether signaling of Rice parameters for transform skip blocks is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled to indicate the Rice parameters for that picture. When the control flag is signaled as disabled (e.g., set equal to "0"), no further syntax elements are signaled at lower levels to indicate Rice parameters for transform skip slices, and a default Rice parameter (e.g., 1) is used for all transform skip slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes to the VVC draft are shown in bold italic font in Table 42. It is worth noting that pps_ts_residual_coding_rice_idx can be coded in various ways and / or may have a maximum value. For example, to encode / decode the same syntax element, one may also use u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right).
[0232] Picture Parameter Set Range Extension Syntax [Table 45]
[0233] pps_ts_residual_coding_rice_flag equal to 1 indicates that the current picture contains pps_ts_residual_coding_rice_ idxpps_ts_residual_coding_rice_flag equal to 0 specifies that pps_ts_residual_coding_rice_idx is not present in the current picture. When pps_ts_residual_coding_rice_flag is not present, the value of pps_ts_residual_coding_rice_flag is inferred.
[0234] pps_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax structure. [Table 46]
[0235] In yet another example, it is proposed to use only variable Rice parameters for encoding the syntax element abs_remainder. The value of the applied Rice parameter may be determined according to specific coded information of the current block, such as block size, quantization parameter, bit depth, transform type, etc. In a specific embodiment, it is proposed to adjust the Rice parameter based on the coded bit depth and the quantization parameter applied to one CU. The corresponding decoding process based on the VVC draft is shown below, with changes to the VVC draft indicated in bold italic font in Table 44, and deleted content indicated in italic font. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter. [Table 47] JPEG2026035734000061.jpg250153 JPEG2026035734000062.jpg20164
[0236] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 33 or 34). Changes to the VVC draft are shown in bold italic font in Table 45, with deleted content in italics. It is worth noting that in practice the same logic may be implemented in different ways. For example, specific equations or lookup tables may also be used to derive the same Rice parameters. [Table 48]
[0237] In yet another example, the corresponding decoding process based on the VVC draft is shown below, A and T.H. B is a predetermined threshold (e.g., TH A =8,TH B = 33 or 34). Changes to the VVC draft are shown in bold italics in Table 46, with deletions in italics. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 49]
[0238] In yet another example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, it is proposed to use only variable Rice parameters for encoding the syntax element of abs_remainder. The variable value may be determined according to specific coded information of the current block, e.g., block size, quantization parameter, bit depth, transform type, etc. In a specific embodiment, it is proposed to adjust the Rice parameter based on the coded bit depth and the quantization parameter applied to one CU. The corresponding decoding process based on the VVC draft is shown below. The changes relative to the VVC draft are shown in bold italic font in Table 47. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter. [Table 50] JPEG2026035734000066.jpg57164
[0239] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 18, 19). The changes to the VVC draft are shown in bold italic font in Table 48. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 51]
[0240] In yet another example, the corresponding decoding process based on the VVC draft is shown below, A and T.H. B is a predetermined threshold (e.g., TH A =8,TH B= 18 or 19). Changes to the VVC draft are shown in bold italic font in Table 49. It is worth noting that in practice the same logic can be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 52]
[0241] FIG. 16 illustrates a method for video encoding. The method may be applied to, for example, an encoder. In step 1610, the encoder may receive a video input. The video input may be, for example, a live stream. In step 1612, the encoder may obtain a quantization parameter based on the video input. The quantization parameter may be calculated, for example, by a quantization unit within the encoder. In step 1614, the encoder may derive a Rice parameter based on at least one predetermined threshold, a coding bit depth, and the quantization parameter. The Rice parameter may be used, for example, to signal the syntax of abs_remainder and dec_abs_level. In step 1616, the encoder may entropy encode a video bitstream based on the Rice parameter. The video bitstream may be entropy encoded, for example, to generate a compressed video bitstream.
[0242] In yet another example, when BitDepth is greater than 10, it is proposed to use only fixed values (e.g., 2, 3, 4, 5, 6, 7, or 8) for the Rice parameter when encoding the abs_remainder syntax element. The fixed value may vary under different conditions according to the specific coded information of the current block, such as the quantization parameter. The corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 18, 19). The changes relative to the VVC draft are shown in bold italic font in Table 50. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter. [Table 53]
[0243] In yet another example, the corresponding decoding process based on the VVC draft is shown below, A and T.H. B is a predetermined threshold (e.g., TH A =8,TH B = 18 or 19). Changes to the VVC draft are shown in bold italic font in Table 51. It is worth noting that in practice the same logic can be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 54]
[0244] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 33 or 34). The changes to the VVC draft are shown in bold italic font in Table 52. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 55]
[0245] In yet another example, the corresponding decoding process based on the VVC draft is shown below, A and T.H. B is a predetermined threshold (e.g., TH A =8,TH B = 33 or 34). Changes to the VVC draft are shown in bold italic font in Table 53. It is worth noting that in practice the same logic can be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 56]
[0246] It is worth noting that in the above illustration, the equation used to calculate a specific Rice parameter is used only as an example to illustrate the proposed concept. For those skilled in the art of modern video coding techniques, other mapping functions (or equivalent mapping equations) are already applicable to the proposed concept (i.e., determining the Rice parameter for the transform skip mode based on the coded bits and the applied quantization parameter). It should also be noted that the current VVC design allows the value of the applied quantization parameter to vary at the coding block group level. Therefore, the proposed Rice parameter adjustment scheme can provide flexible adaptation of the Rice parameter for the transform skip mode at the coding block group level.
[0247] Signal transmission information for normal residual coding and transform-skip residual coding According to a 28th aspect of the present disclosure, it is proposed to signal certain syntax elements in transform skip residual coding, such as the Rice parameter of the binary codeword for coding abs_remainder, the shift parameter and offset parameter for deriving the Rice parameter used for abs_remainder / dec_abs_level in normal residual coding, and to decide whether to signal them according to certain coded information of the current block, such as the quantization parameter or coding bit depth associated with the TB / CB and / or slice / profile, and / or according to a new flag associated with the TB / CB / slice / picture / sequence level, such as sps_residual_coding_info_present_in_sh_flag.
[0248] In one example, one control flag is signaled in a slice header to indicate whether signaling of Rice parameters for transform skip blocks and signaling of shift parameters and / or offset parameters for derivation of Rice parameters in transform blocks is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled for each transform skip slice to indicate the Rice parameters for that slice, and two syntax elements are further signaled for each transform slice to indicate the shift parameters and / or offset parameters for derivation of Rice parameters for that slice. When the control flag is signaled as disabled (e.g., set equal to "0"), no further syntax elements are signaled at lower levels to indicate Rice parameters for transform skip slices, and default Rice parameters (e.g., 1) are used for all transform skip slices; and no further syntax elements are signaled at lower levels to indicate shift parameters and offset parameters for derivation of Rice parameters for transform slices, and default shift parameters and / or offset parameters (e.g., 0) are used for all transform slices. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes to the VVC draft are shown in bold italic font in Table 54. It is worth noting that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_index can be coded in various ways and / or may have maximum values. For example, to code / decode the same syntax element, u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right), may also be used.
[0249] 17 illustrates a method for video decoding. The method may be applied to, for example, an encoder. In step 1710, the encoder may receive a video input. In step 1712, the encoder may signal a Rice parameter of a binary codeword for a coding syntax element. The coding syntax element may include abs_remainder in transform skip residual coding. In step 1714, the encoder may entropy encode a video bitstream based on the Rice parameter and the video input.
[0250] Slice Header Syntax [Table 57]
[0251] sh_residual_coding_rice_flag equal to 1 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_residual_coding_rice_index may be present in the current slice. sh_residual_coding_rice_flag equal to 0 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_residual_coding_rice_index are not present in the current slice.
[0252] sh_residual_coding_rice_shift specifies the shift parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_shift is not present, the value of sh_residual_coding_rice_shift is inferred to be equal to 0.
[0253] sh_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_offset is not present, the value of sh_residual_coding_rice_offset is inferred to be equal to 0.
[0254] sh_ts_residual_coding_rice_index specifies the rice parameter used in the residual_ts_coding() syntax construct. When sh_ts_residual_coding_rice_index is not present, the value of sh_ts_residual_coding_rice_index is inferred to be equal to 0. [Table 58] [Table 59]
[0255] In another example, one control flag is signaled in a sequence parameter set (or in a sequence parameter set range extension syntax) to indicate whether signaling of Rice parameters for transform skip blocks and signaling of shift parameters and / or offset parameters for the derivation of Rice parameters in transform blocks is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled for each transform skip slice to indicate the Rice parameters for that slice, and two syntax elements are further signaled for each transform slice to indicate the shift parameters and / or offset parameters for the derivation of Rice parameters for that slice. When the control flag is signaled as disabled (e.g., set equal to "0"), no further syntax elements are signaled at lower levels to indicate the Rice parameters of the transform skip slices, and a default Rice parameter (e.g., 1) is used for all transform skip slices; no further syntax elements are signaled at lower levels to indicate the shift and / or offset parameters for the Rice parameter derivation of the transform slices, and a default shift and / or offset parameter (e.g., 0) is used for all transform slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes relative to the VVC draft are shown in bold italic font in Table 57. It is worth noting that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx can be coded in various ways and / or may have maximum values. For example, to encode / decode the same syntax element, one may also use u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right).
[0256] Sequence Parameter Set RBSP Syntax [Table 60]
[0257] sps_residual_coding_info_present_in_sh_flag equal to 1 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx may be present in SH syntax structures that reference an SPS. sps_residual_coding_info_present_in_sh_flag equal to 0 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx are not present in SH syntax structures that reference an SPS. When sps_residual_coding_info_present_in_sh_flag is not present, the value of sps_residual_coding_info_present_in_sh_flag is inferred to be equal to 0.
[0258] Slice Header Syntax [Table 61]
[0259] sh_residual_coding_rice_shift specifies the shift parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_shift is not present, the value of sh_residual_coding_rice_shift is inferred to be equal to 0.
[0260] sh_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_offset is not present, the value of sh_residual_coding_rice_offset is inferred to be equal to 0.
[0261] sh_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax construct. When sh_ts_residual_coding_rice_index is not present, the value of sh_ts_residual_coding_rice_index is inferred to be equal to 0. [Table 62] [Table 63]
[0262] In yet another example, one syntax element is signaled for each transform skip slice to indicate the Rice parameter for that slice, and two syntax elements are signaled for each transform slice to indicate the shift and / or offset parameters for the derivation of the Rice parameter for that slice. An example of a corresponding decoding process based on the VVC draft is shown below. The changes to the VVC draft are shown in bold italic font in Table 61. It is worth noting that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx can be coded in various ways and / or have maximum values. For example, u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right), may also be used to code / decode the same syntax element.
[0263] Slice Header Syntax [Table 64]
[0264] sh_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax construct. When sh_ts_residual_coding_rice_idx is not present, the value of sh_ts_residual_coding_rice_idx is inferred to be equal to 0.
[0265] sh_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_offset is not present, the value of sh_residual_coding_rice_offset is inferred to be equal to 0.
[0266] sh_residual_coding_rice_shift is abs_remainder[] and dec_abs_level[] Rice parameter derivation process Used for shift· Specifies the parameter. When sh_residual_coding_rice_shift is not present, the value of sh_residual_coding_rice_shift is inferred to be equal to 0. [Table 65] [Table 66]
[0267] In yet another example, one control flag is signaled in the picture parameter set range extension syntax to indicate whether signaling of Rice parameters for transform skip blocks and signaling of shift and / or offset parameters for the derivation of Rice parameters in transform blocks are enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled to indicate Rice parameters for transform skip residual coding of the picture, and two syntax elements are further signaled to indicate shift and / or offset parameters for the derivation of Rice parameters for the picture per normal residual coding. When the control flag is signaled as disabled (e.g., set equal to "0"), no further syntax elements at lower levels are signaled to indicate the Rice parameters for transform-skip residual coding, and the default Rice parameters (e.g., 1) are used for all transform-skip residual coding; no further syntax elements at lower levels are signaled to indicate the shift and / or offset parameters for the Rice parameter derivation for regular residual coding, and the default shift and / or offset parameters (e.g., 0) are used for all regular residual coding. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes relative to the VVC draft are shown in bold italic font in Table 64. It is worth noting that pps_residual_coding_rice_shift, pps_residual_coding_rice_offset, and pps_ts_residual_coding_rice_idx can be coded in various ways and / or may have maximum values. For example, to encode / decode the same syntax element, one may also use u(n), an unsigned integer using n bits, or f(n), a fixed-pattern bit string using n bits written left bit first (left to right).
[0268] Picture Parameter Set Range Extension Syntax [Table 67]
[0269] pps_residual_coding_info_flag equal to 1 indicates that pps_residual_coding_rice_shift, pps_residual_coding_rice_offset, and pps_ts_residual_coding_rice_ idx may be present. pps_residual_coding_info_flag equal to 0 specifies that pps_residual_coding_rice_shift, pps_residual_coding_rice_offset, and pps_ts_residual_coding_rice_idx are not present in the current picture. When pps_residual_coding_info_flag is not present, the value of pps_residual_coding_info_flag is inferred.
[0270] pps_residual_coding_rice_shift specifies the shift parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When pps_residual_coding_rice_shift is not present, the value of pps_residual_coding_rice_shift is inferred to be equal to 0.
[0271] pps_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When pps_residual_coding_rice_offset is not present, the value of pps_residual_coding_rice_offset is inferred to be equal to 0.
[0272] pps_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax structure. idx When does not exist, pps_ts_residual_coding_rice_ idx The value of is inferred to be equal to 0. [Table 68] [Table 69]
[0273] According to a 29th aspect of the present disclosure, it is proposed to use different Rice parameters for coding certain syntax elements, e.g., abs_remainder, in transform skip residual coding, shift and offset parameters for deriving the Rice parameters used for abs_remainder / dec_abs_level in normal residual coding, and to decide which one to use according to certain coded information of the current block, e.g., quantization parameters or coding bit depth associated with TB / CB and / or slice / profile, and / or according to a new flag associated with TB / CB / slice / picture / sequence level, e.g., sps_residual_coding_info_present_in_sh_flag.
[0274] In one example, a control flag is signaled in the slice header to indicate whether the derivation process of the Rice parameter for the transform skip block and the derivation process of the shift parameter and / or offset parameter for the Rice parameter in the transform block are enabled or disabled. When the control flag is signaled as enabled, the Rice parameter may vary under different conditions according to the specific coded information of the current block, such as the quantization parameter and bit depth. Also, the shift parameter and / or offset parameter for the derivation of the Rice parameter in normal residual coding may vary under different conditions according to the specific coded information of the current block, such as the quantization parameter and bit depth. When the control flag is signaled as disabled (e.g., set equal to "0"), the default Rice parameter (e.g., 1) is used for all transform skip slices, and the default shift parameter and / or offset parameter (e.g., 0) is used for all transform slices. An example of a corresponding decoding process based on the VVC draft is shown below, and the TH A and T.H. B is a predetermined threshold (e.g., TH A =8,TH B = 18 or 19). Changes to the VVC draft are shown in bold italic font in Table 67. It is worth noting that in practice the same logic can be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters.
[0275] Slice Header Syntax [Table 70]
[0276] sh_residual_coding_rice_flag equal to 1 specifies that the bit-depth dependent Rice parameter derivation process is used for the current slice. sh_residual_coding_rice_flag equal to 0 specifies that the bit-depth dependent Rice parameter derivation process is not used for the current slice. [Table 71] [Table 72]
[0277] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 18, 19). The changes to the VVC draft are shown in bold italic font in Table 70. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameters. [Table 73] JPEG2026035734000090.jpg82164
[0278] According to another aspect of the present disclosure, it is proposed to add constraints that the values of these encoding tools above flag to provide the same general constraint control as others in the general constraint information.
[0279] For example, sps_ts_residual_coding_rice_present_in_sh_flag equal to 1 specifies that sh_ts_residual_coding_rice_idx may be present in SH syntax structures that reference an SPS. sps_ts_residual_coding_rice_present_in_sh_flag equal to 0 specifies that sh_ts_residual_coding_rice_idx is not present in SH syntax structures that reference an SPS. According to the present disclosure, it is proposed to add the syntax element gci_no_ts_residual_coding_rice_constraint_flag in the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The additions are highlighted in italics. [Table 74] [Table 75]
[0280] In another example, pps_ts_residual_coding_rice_flag equal to 1 indicates that there is no pps_ts_residual_coding_rice_ idx pps_ts_residual_coding_rice_flag equal to 0 specifies that pps_ts_residual_coding_rice_idx is not present in the current picture. According to the present disclosure, it is proposed to add the syntax element gci_no_ts_residual_coding_rice_constraint_flag in the general constraint information syntax to provide the same general constraint control as the other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The additions are highlighted in italics. [Table 76]
Table 77
[0281] In yet another example, that the sps_rice_adaptation_enabled_flag is equal to 1 indicates that the Rice parameters for the binarization of abs_remaining[] and dec_abs_level can be derived by an expression.
[0282] The expression may include RiceParam = RiceParam + shiftVal, and shiftVal = (localSumAbs < Tx[0])? Rx[0] : ((localSumAbs < Tx[1])? Rx[1] : ((localSumAbs < Tx[2])? Rx[2] : ((localSumAbs < Tx[3])? Rx[3] : Rx[4]))), where the lists Tx[] and Rx[] are specified as Tx[] = {32, 128, 512, 2048}>>(1523) Rx[] = {0, 2, 4, 6, 8}.
[0283] According to the present disclosure, in order to provide the same general constraint control as other flags, it is proposed to add a syntax element gci_no_rice_adaptation_constraint_flag into the general constraint information syntax. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are highlighted in italics.
Table 78
Table 79
[0284] Since the proposed Rice parameter adaptation scheme is only used for transform skip residual coding (TSRC), the proposed method can be effective when TSRC is enabled. Correspondingly, one or more embodiments of the present disclosure propose to add one bitstream constraint that requires the value of gci_no_rice_adaptation_constraint_flag to be 1 when the transform skip mode is disabled from the general constraint information level, e.g., when the value of gci_no_transform_skip_constraint_flag is set to 1.
[0285] In yet another example, sps_range_extension_flag equal to 1 specifies the presence of the sps_range_extension() syntax structure in the SPS RBSP syntax structure. sps_range_extension_flag equal to 0 specifies the absence of this syntax structure. According to the present disclosure, it is proposed to add the syntax element gci_no_range_extension_constraint_flag in the general constraint information syntax to provide the same general constraint control as the other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The additions are highlighted in italics. [Table 80] [Table 81]
[0286] 19 illustrates a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 1902, the decoder may receive a sequence parameter set (SPS) range extension flag indicating whether a syntax structure sps_range_extension is present in a raw byte sequence payload (RBSP) syntax structure of a slice head (SH) based on a value of the SPS range extension flag.
[0287] In step 1904, the decoder may determine that sps_range_extension is present in the SH RBSP syntax structure in response to determining that the value of the SPS range extension flag is equal to one.
[0288] In step 1906, the decoder may determine that sps_range_extension is not present in the SH RBSP syntax structure in response to determining that the value of the range extension flag is equal to 0.
[0289] In yet another example, sps_cabac_bypass_alignment_enabled_flag equal to 1 specifies that the value of ivlCurrRange may be aligned before bypass decoding of the syntax elements sb_coded_flag[][], abs_remainder[], dec_abs_level[n], and coeff_sign_flag[]. sps_cabac_bypass_alignment_enabled_flag equal to 0 specifies that the value of ivlCurrRange is not aligned before bypass decoding. According to the present disclosure, it is proposed to add the syntax element gci_no_cabac_bypass_alignment_constraint_flag within the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes relative to the VVC draft are highlighted. The additions are highlighted in italics. [Table 82] [Table 83]
[0290] 20 shows a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 2002, the decoder may receive a sequence parameter set (SPS) alignment enable flag indicating whether the index ivlCurrRange is aligned before bypass decoding of the syntax elements sb_coded_flag, abs_remainder, dec_abs_level, and coeff_sign_flagn based on the value of SPS alignment enable.
[0291] In step 2004, the decoder may determine that the ivlCurrRange is aligned before bypass decoding in response to determining that the value of the SPS alignment valid flag is equal to one.
[0292] In step 2006, the decoder may determine that the ivlCurrRange is not aligned prior to bypass decoding in response to determining that the value of the SPS alignment valid flag is equal to 0.
[0293] In yet another example, extended_precision_processing_flag equal to 1 specifies that extended dynamic range may be used for transform coefficients and transform processing. Extended_precision_processing_flag equal to 0 specifies that extended dynamic range is not used. According to the present disclosure, it is proposed to add a syntax element gci_no_extended_precision_processing_constraint_flag in the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The additions are highlighted in italics. [Table 84] [Table 85]
[0294] 21 illustrates a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 2102, the decoder may receive an extended precision processing flag indicating whether extended dynamic range is employed for transform coefficients and during the transform process based on a value of the extended precision processing flag.
[0295] In step 2104, the decoder may determine that an extended dynamic range is to be employed for the transform coefficients and during the transform process in response to determining that the value of the extended precision processing flag is equal to one.
[0296] In step 2106, the decoder may determine that extended dynamic range is not employed for the transform coefficients or during the transform process in response to determining that the value of the extended precision processing flag is equal to 0.
[0297] In yet another example, persistent_rice_adaptation_enabled_flag equal to 1 specifies that the Rice parameter derivation for binarization of abs_remaining[] and dec_abs_level at the beginning of each sub-block may be initialized using mode-dependent statistics accumulated from the previous sub-block. persistent_rice_adaptation_enabled_flag equal to 0 specifies that the previous sub-block state is not used in the Rice parameter derivation. According to the present disclosure, it is proposed to add the syntax element gci_no_persistent_rice_adaptation_constraint_flag in the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The additions are highlighted in italics. [Table 86] [Table 87]
[0298] 22 illustrates a method for video encoding according to an example of this disclosure. The method may be applied to, for example, a decoder. In step 2202, the decoder may receive a persistent Rice adaptation enable flag that indicates whether Rice parameter derivation for binarization of abs_remaining and dec_abs_level is initialized at the beginning of each sub-block employing mode-dependent statistics accumulated from the previous sub-block based on the value of the persistent Rice adaptation enable flag.
[0299] In step 2204, the decoder may determine, in response to determining that the value of the persistent Rice adaptation enable flag is equal to one, that the Rice parameter derivation for binarization is initialized at the beginning of each sub-block employing mode-dependent statistics accumulated from the previous sub-block.
[0300] In step 2206, the decoder may determine that the previous sub-block state is not to be employed in the Rice parameter derivation in response to determining that the value of the persistent Rice adaptation valid flag is equal to 0.
[0301] The above methods may be implemented using an apparatus including one or more circuits, including an application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), controller, microcontroller, microprocessor, or other electronic component. The apparatus may use the circuits in combination with other hardware or software components to perform the above-described methods. Each module, sub-module, unit, or sub-unit disclosed above may be at least partially implemented using one or more circuits.
[0302] Determination of rice parameters On the encoder side, TSRC encoding may require multiple encoding passes to derive the best Rice parameter. This multi-pass encoding may not be suitable for practical hardware encoder design. To solve this problem, a low-delay TSRC encoding method is also proposed. According to a 30th aspect of the present disclosure, it is proposed to derive the Rice parameter according to specific encoded information of the current slice, such as the quantization parameter and / or coding bit depth associated with the slice / picture / sequence, and / or the hash ratio associated with the level of the slice / picture / sequence. Various methods may be used to derive the Rice parameter, and some exemplary methods are listed below. It should be noted that the following methods may be applied alone or in combination.
[0303] 1. The Rice parameters mentioned in the above embodiments may further depend on the video resolution, which includes both the temporal resolution (e.g., frame rate) and spatial resolution (e.g., picture width and height) of the video.
[0304] 2. The Rice parameter may vary at the sequence level, picture level, slice level, and / or any predefined region. In one specific example, different Rice values are used for pictures with different temporal layer IDs (related to nuh_temporal_id_plus1 specified in the VVC specification). Alternatively, the Rice parameter may include a value determined based on the QP value used at the sequence level, picture level, slice level, and / or any predefined region. For example, the Rice parameter = Clip3(1, 8, (TH-QP) / 6), where TH is a predetermined threshold (e.g., 18, 19).
[0305] 3. The Rice parameter may be set to a default value, e.g., 1, according to the change in coded information between the current slice and the previous slice. In one specific example, when the temporal layer ID has changed compared to the previous picture, the default Rice value is used for the picture. Alternatively, when ΔQ is greater than TH, the default Rice value is used for the picture, where ΔQ is calculated as abs(QPcurrent-QPprevious), and TH is a predetermined threshold value. The Rice parameter (e.g., 0, 5) is used. For example, when the hash ratio from the intra-block duplication mode of the current slice is greater than TH, the Rice parameter = 1, where TH is a predetermined threshold value, e.g., Max(41*(number of CTUs), 4200).
[0306] 4. The Rice parameter for each slice is based on the value of abs_remainder coded in the preceding slice according to the coding order. In a specific example, after one slice is coded, the number of bins for binarization of abs_remainder using a different Rice parameter is calculated, and then that number is used to determine the Rice parameter for the subsequent slice. For example, the Rice parameter that achieves the smallest number of bins in the preceding slice is selected for the current slice. As another example, if the current slice and its preceding slice use the same QP, the Rice parameter that achieves the smallest number of bins in the preceding slice is selected for the current slice; otherwise, the number of bins generated in the preceding slice using the default Rice parameter (i.e., 1) is scaled by TH before being compared with other Rice parameters, and the Rice parameter that results in the smallest number of bins is selected for the current slice, where TH is a predetermined threshold, for example, 0.9.
[0307] 5. The Rice parameter for each slice is based on the value of abs_remainder coded in the previous slice according to the coding order. The Rice parameter may be adjusted according to changes in coded information between the current slice and the previous slice. In a specific example, the Rice parameter that achieves the minimum number of bins in the previous slice is selected for the current slice. Also, when ΔQ is greater than TH, the Rice value may be adjusted, where ΔQ is calculated as abs(QPcurrent-QPprevious), and TH is a predetermined threshold value (e.g., 0, 5). The adjustment may be adding a predetermined offset (e.g., +1, -1) or scaling by a predetermined value.
[0308] 26 shows a flow diagram of a low-delay transform skip residual coding (TSRC) method according to an example of the present disclosure. The method may be applied to, for example, an encoder. In step 2602, the encoder may derive Rice parameters based on coded information of a current slice of video. The coded information may include one or more of the following parameters: a quantization parameter or coding bit depth associated with the slice, picture, or sequence of video, or a hash ratio associated with the slice, picture, or sequence of video.
[0309] It should be noted that the above encoder method may also be applied on the decoder side: in one specific example, the Rice parameters do not need to be signaled to the decoder, and the encoder / decoder derives the Rice parameters using the same method.
[0310] 18 illustrates a computing environment 1810 coupled with a user interface 1860. The computing environment 1810 may be part of a data processing server. The computing environment 1810 includes a processor 1820, a memory 1840, and an I / O interface 1850.
[0311] The processor 1820 typically controls the overall operation of the computing environment 1810, such as operations related to display, data acquisition, data communication, and image processing. The processor 1820 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Additionally, the processor 1820 may include one or more modules that facilitate interaction between the processor 1820 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0312] Memory 1840 is configured to store various types of data to support the operation of computing environment 1810. Memory 1840 may include predefined software 1842. Examples of such data include instructions for any applications or methods operating on computing environment 1810, video data sets, image data, etc. Memory 1840 may be implemented using any type of volatile or non-volatile memory device, or combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0313] The I / O interface 1850 provides an interface between the processor 1820 and peripheral interface modules such as a keyboard, click wheel, buttons, and the like. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1850 may be coupled to an encoder and a decoder.
[0314] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs, such as those contained in memory 1840, executable by processor 1820 in computing environment 1810 for performing the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.
[0315] A non-transitory computer-readable storage medium stores a plurality of programs executed by a computing device having one or more processors therein, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the above-described method for motion estimation.
[0316] In some embodiments, the computing environment 1810 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0317] 23 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel, according to some implementations of the present disclosure. As shown in FIG. 23, system 10 includes a source device 12 that generates and encodes video data to be subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, or the like. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0318] In some implementations, destination device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or communication device capable of moving encoded video data from source device 12 to destination device 14. In one example, link 16 may include a communication medium that enables source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.
[0319] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, digital versatile disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, storage device 32 may correspond to a file server or another intermediate storage device capable of holding the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network-attached storage (NAS) devices, or local disk drives. Destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0320] 23 , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a video camera for a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, implementations described herein may be applicable to video encoding generally and may be applicable to wireless and / or wired applications.
[0321] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.
[0322] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem and receive encoded video data via link 16. The encoded video data communicated over link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use in decoding the video data by video decoder 30. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0323] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0324] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and is applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data in accordance with any of these current or future standards.
[0325] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If implemented partially in software, the electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.
[0326] 24 is a block diagram illustrating an example video encoder 20 according to some implementations described in this application. Video encoder 20 may perform intra-predictive and inter-predictive coding of video blocks within video frames. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. It should be noted that the term "frame" may be used synonymously with the term "image" or "picture" in the field of video coding.
[0327] As shown in FIG. 24 , video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, adder 50, transform processing unit 52, quantization unit 54, and entropy coding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partition unit 45, intra-prediction processing unit 46, and intra-block replication (BC) unit 48. In some implementations, video encoder 20 also includes inverse quantization unit 58, inverse transform processing unit 60, and adder 62 for video block reconstruction. An in-loop filter 63, such as a deblocking filter, may be disposed between adder 62 and DPB 64 to filter block boundaries and remove blocky artifacts from the reconstructed video. In addition to the deblocking filter, another in-loop filter, such as a sample adaptive offset (SAO) filter and / or an adaptive in-loop filter (ALF), may also be used to filter the output of adder 62. In some examples, the in-loop filter may be omitted, and the decoded video block may be provided directly to DPB 64 by summer 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.
[0328] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from video source 18 shown in FIG. 23. DPB 64 is a buffer that stores reference video data (e.g., reference frames or reference pictures) for use in encoding video data by video encoder 20 (e.g., in an intra-predictive coding mode or an inter-predictive coding mode). Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various examples, video data memory 40 may be on-chip with other components of video encoder 20 or off-chip relative to those components.
[0329] As shown in FIG. 24, after receiving video data, partition unit 45 in prediction processing unit 41 partitions the video data into video blocks. This partitioning also includes partitioning the video frame into slices, tiles (e.g., sets of video blocks), or other larger coding units (CUs) according to a predefined decomposition structure, such as a quad-tree (QT) structure associated with the video data. A video frame may be or be considered as a two-dimensional array or matrix of samples having sample values. The samples in the array are also sometimes referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. A video frame may be divided into multiple video blocks, for example, by using QT partitioning. A video block may also be or be considered as a two-dimensional array or matrix of samples having sample values, although with smaller dimensions than a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further partitioned into one or more block partitions or sub-blocks (which may again form blocks), for example, by iteratively using QT partitioning, binary-tree (BT) partitioning, or triple-tree (TT) partitioning, or any combination thereof. It should be noted that the term "block" or "video block" as used herein may be a portion of a frame or picture, specifically a rectangular (square or non-square) portion. For example, with respect to HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or a corresponding block, e.g., a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB), and / or sub-block.
[0330] Prediction processing unit 41 may select one of multiple possible predictive coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for the current video block based on the error result (e.g., coding rate and distortion level). Prediction processing unit 41 may provide the resulting intra-predictively coded block or inter-predictively coded block to summer 50 to generate a residual block, and may also provide the resulting intra-predictively coded block or inter-predictively coded block to summer 62 to reconstruct a coded block for subsequent use as part of a reference frame. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to entropy coding unit 56.
[0331] To select an appropriate intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple coding passes, e.g., to select an appropriate coding mode for each block of video data.
[0332] In some implementations, motion estimation unit 42 determines the inter-prediction mode for the current video frame by generating motion vectors that indicate the displacement of video blocks in the current video frame relative to predictive blocks in a reference video frame according to a predetermined pattern in the sequence of video frames. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of video blocks in the current video frame or picture relative to predictive blocks in a reference frame relative to the current block being coded in the current frame. The predetermined pattern designates the video frames in the sequence as P frames or B frames. Intra BC unit 48 may determine vectors, e.g., block vectors, for intra BC coding in a manner similar to the determination of motion vectors by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine block vectors.
[0333] The predictive block of a video block may be or correspond to a block or reference block of a reference frame that is deemed to closely match, in terms of pixel difference, the video block to be encoded, which may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metric. In some implementations, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter pixel positions, eighth pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion search for whole pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy.
[0334] Motion estimation unit 42 calculates a motion vector for a video block in an inter-predictively coded frame by comparing the position of the video block with the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), which each identify one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy coding unit 56.
[0335] The motion compensation performed by motion compensation unit 44 may include fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the current video block, motion compensation unit 44 may locate the predictive block to which the motion vector points in one of the reference frame lists, obtain the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting pixel values of the predictive block provided by motion compensation unit 44 from pixel values of the current video block being coded. The pixel difference values forming the residual video block may include a luma difference component, a chroma difference component, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use in decoding the video blocks of the video frame by video decoder 30. The syntax elements may include, for example, a syntax element defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.
[0336] In some implementations, the intra BC unit 48 may generate a vector and fetch a predictive block in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is within the same frame as the current block being coded, and the vector is referred to as a block vector rather than a motion vector. Specifically, the intra BC unit 48 may determine an intra prediction mode to use to code the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, e.g., during separate coding passes, and test their performance through rate-distortion analysis. The intra BC unit 48 may then select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using the rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to create the coded block, as well as the bitrate (i.e., number of bits) used to create the coded block. Intra BC unit 48 may calculate ratios from the distortions and rates of the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for that block.
[0337] In other examples, intra BC unit 48 may use all or part of motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction according to implementations described herein. In either case, for intra block duplication, the predictive block may be a block deemed to closely match the block to be coded, as may be determined by SAD, SSD, or other difference metric in terms of pixel difference, and identification of the predictive block may include calculating values for sub-integer pixel positions.
[0338] Whether the predictive block is a block from the same frame according to intra prediction or a block from a different frame according to inter prediction, video encoder 20 may form a residual video block by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. The pixel difference values forming the residual video block may include both luma and chroma component differences.
[0339] Intra-prediction processing unit 46 may intra-predict the current video block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44, or as an alternative to intra-block replication prediction performed by intra BC unit 48, as described above. Specifically, intra-prediction processing unit 46 may determine an intra-prediction mode to use to encode the current block. To do so, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 (or, in some examples, a mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy coding unit 56. Entropy coding unit 56 may encode the information indicating the selected intra-prediction mode into the bitstream.
[0340] After prediction processing unit 41 determines a predictive block for the current video block by inter-prediction or intra-prediction, adder 50 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more TUs, is provided to transform processing unit 52. Transform processing unit 52 converts the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0341] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54, which quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.
[0342] Following quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) encoding, or another entropy encoding methodology or technique. The encoded bitstream may then be transmitted to video decoder 30 as shown in FIG. 23 or archived in storage device 32 as shown in FIG. 23 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements of the current video frame being encoded.
[0343] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate reference blocks for prediction of other video blocks. As described above, motion compensation unit 44 may generate motion-compensated prediction blocks from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0344] Summer 62 may add the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to create a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.
[0345] 25 is a block diagram illustrating an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described above for the video encoder 20 in connection with FIG. 24. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.
[0346] In some examples, units of video decoder 30 may be assigned tasks to perform implementations of the present application. Also, in some examples, implementations of the present disclosure may be divided among one or more of the units of video decoder 30. For example, intra BC unit 85 may perform implementations of the present application alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 may not include intra BC unit 85, and the functions of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.
[0347] Video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores coded video data from the coded video bitstream. DPB 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra-predictive or inter-predictive coding modes). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. 25, for purposes of illustration, video data memory 79 and DPB 92 are depicted as two separate components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30 or off-chip relative to those components.
[0348] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.
[0349] When a video frame is coded as an intra-predictively coded (I) frame or for intra-coded predictive blocks in other types of frames, intra-prediction unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video frame based on a signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.
[0350] When a video frame is coded as an inter-predictive (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 may generate one or more predictive blocks of a video block of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the predictive blocks may be created from a reference frame in one of the reference frame lists. Video decoder 30 may construct reference frame lists List 0 and List 1 using a default construction technique based on the reference frames stored in DPB 92.
[0351] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 creates a predictive block of the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The predictive block may be within the same reconstructed region of the picture as the current video block as defined by video encoder 20.
[0352] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by analyzing the motion vectors and other syntax elements, and then use the prediction information to create a predictive block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-prediction or inter-prediction) used to encode the video blocks of the video frame, the inter-prediction frame type (e.g., B or P), construction information for one or more of the frame's reference frame lists, the motion vectors of each inter-predictively coded video block of the frame, the inter-prediction status of each inter-predictively coded video block of the frame, and other information for decoding video blocks in the current video frame.
[0353] Similarly, intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block was predicted using an intra BC mode, construction information regarding which video blocks of the frame are within the reconstruction region and which video blocks should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.
[0354] Motion compensation unit 82 may also perform interpolation using an interpolation filter used by video encoder 20 during encoding of the video block to calculate sub-integer pixel interpolated values of the reference block. In this case, motion compensation unit 82 may determine the interpolation filter used by video encoder 20 from the received syntax element and use the interpolation filter to create the predictive block.
[0355] Inverse quantization unit 86 uses the same quantization parameters calculated by video encoder 20 for each video block in a video frame to determine the degree of quantization to inverse quantize the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct residual blocks in the pixel domain.
[0356] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs a decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. To further process the decoded video block, an in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF filter, may be disposed between adder 90 and DPB 92. In some examples, in-loop filter 91 may be omitted, and the decoded video block may be provided directly to DPB 92 by adder 90. The decoded video block in a given frame is then stored in DPB 92, which stores reference frames used for subsequent motion compensation of the next video block. DPB 92, or a memory device separate from DPB 92, may also store the decoded video for later presentation on a display device, such as display device 34 of FIG. 23 .
[0357] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the disclosure. Many modifications, variations and alternative implementations will become apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.
[0358] The examples have been chosen and described to explain the principles of the disclosure and to enable those skilled in the art to understand the disclosure in various embodiments and to make best use of the underlying principles and various implementations with various modifications suited to the particular applications contemplated. Therefore, it should be understood that the scope of the disclosure is not limited to the particular examples of implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the disclosure.
Claims
1. 1. A method for video encoding, comprising: Disabling the presence of Rice parameters for transform skip residual coding in response to determining, by the decoder, that transform skip is disabled. A method comprising:
2. Adopting the transform skip flag sps_transform_skip_enabled_flag to condition the presence of sps_ts_residual_coding_rice_present_in_sh_flag The method for video encoding of claim 2 further comprising:
3. in response to determining that the flag sps_transform_skip_enabled_flag is equal to 0 and transform skip is disabled in the current picture, inferring that the sps_ts_residual_coding_rice_present_in_sh_flag is 0 and not signaled. The method for video encoding of claim 2 further comprising:
4. signaling the sps_ts_residual_coding_rice_present_in_sh_flag in response to determining that the flag sps_transform_skip_enabled_flag is equal to 1 and transform skip is not disabled in the current picture. The method for video encoding of claim 2 further comprising:
5. signaling a control flag in the sequence parameter set or in the sequence parameter set range extension syntax to indicate whether signaling of Rice parameters for the transform skip block is enabled or disabled in response to determining that the transform skip flag sps_transform_skip_enabled_flag is signaled as enabled. The method for video encoding of claim 2 further comprising:
6. signaling one syntax element for each transform skip slice to indicate a Rice parameter for each corresponding transform skip slice in response to determining that the control flag is signaled as valid. The method for video encoding of claim 5 further comprising:
7. 6. The method for video encoding of claim 5, further comprising, in response to determining that the control flag is signaled as invalid and set equal to 0, adopting default Rice parameters for all transform skip slices, and no further syntax elements are signaled to indicate the Rice parameters for the transform skip slices.
8. determining, in response to determining that sps_ts_residual_coding_rice_present_in_sh_flag is equal to 1, that an index sh_ts_residual_coding_rice_idx is present in a slice header (SH) syntax structure that references a sequence parameter set (SPS); The method for video encoding of claim 2 further comprising:
9. determining that the index sh_ts_residual_coding_rice_present_in_sh_flag is equal to 0, and The method for video encoding of claim 2 further comprising:
10. inferring the value of sps_ts_residual_coding_rice_present_in_sh_flag to be 0 in response to determining that sps_ts_residual_coding_rice_present_in_sh_flag is not present; The method for video encoding of claim 2 further comprising:
11. 1. An apparatus for video encoding, comprising: one or more processors; a memory configured to store instructions executable by the one or more processors; 11. An apparatus comprising:
12. 11. A non-transitory computer-readable storage medium for video encoding storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any of claims 1 to 10.