Encoding of Residuals and Coefficients for Video Encoding
Patent Information
- Application Number
- JP2023547334
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-04
- Filing Date
- 2022-01-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-25
Smart Images

Figure 0007698050000101 
Figure 0007698050000102 
Figure 0007698050000103
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to Provisional Application No. 63 / 145,964, filed on February 4, 2021, the entire content of which is incorporated herein by reference for all purposes.
[0002] This disclosure relates to the encoding and compression of video. More particularly, this disclosure relates to improvements and simplifications in the encoding of residuals and coefficients for video encoding.
Background Art
[0003] To compress video data, various video encoding techniques may be used. Video encoding is performed according to one or more video encoding standards. For example, video encoding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High - Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Expert Group (MPEG) coding, or the like. Video encoding generally utilizes prediction methods (e.g., inter - prediction, intra - prediction, or the like) that exploit redundancy present in the moving images or sequences. An important goal of video encoding techniques is to compress video data into a form that uses a lower bit - rate while avoiding or minimizing a degradation in video quality.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Examples of this disclosure provide methods and apparatuses for video encoding.
Means for Solving the Problem
[0005] According to a first aspect of the present disclosure, a method for video encoding is provided. The method may include receiving, by a decoder, a sequence parameter set (SPS) range extension flag indicating whether a syntax structure sps_range_extension exists within a raw byte sequence payload (RBSP) syntax structure of a slice head (SH) based on a value of the SPS range extension flag.
[0006] According to a second aspect of the present disclosure, a method for video encoding is provided. The method may include receiving, by a decoder, an SPS alignment valid flag indicating whether an index ivlCurrRange is aligned before bypass decoding of syntax elements sb_coded_flag, abs_remainder, dec_abs_level, and coeff_sign_flagn based on a value of the SPS alignment valid. Flag of the value.
[0007] According to a third aspect of the present disclosure, a method for video encoding is provided. The method may include receiving, by a decoder, an extended precision processing flag indicating whether an extended dynamic range is adopted for transform coefficients and during transform processing based on a value of the extended precision processing flag.
[0008] According to a fourth aspect of the present disclosure, a method for video encoding is provided. The method may include receiving, by a decoder, a persistent rice adaption enabled flag indicating whether, at the start of each sub-block that employs mode-dependent statistics accumulated from a previous sub-block based on a value of the persistent rice adaption enabled flag, initialization of rice parameter derivation for binarization of abs_remaining and dec_abs_level is performed.
[0009] It should be understood that the above summary and the following detailed description are merely exemplary and explanatory and are not intended to limit the present disclosure.
[0010] The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate examples in accordance with the present disclosure and together with the description serve to explain the principles of the present disclosure.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 6E
Figure 6F
Figure 6G
Figure 6H
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 15A
Figure 15B
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
DETAILED DESCRIPTION OF THE INVENTION
[0012] Next, refer in detail to the exemplary embodiments, examples of which are shown in the accompanying drawings. The following description refers to the accompanying drawings, and unless otherwise specified, the same numbers in different drawings represent the same or similar elements. The implementation forms described in the following description of the exemplary embodiments do not represent all implementation forms in accordance with the present disclosure. Rather, the implementation forms are merely examples of apparatuses and methods in accordance with aspects related to the present disclosure described in the appended claims.
[0013] The terms used in the present disclosure are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. The singular forms "a", "an", and "the" are intended to include the plural forms as well when used in the present disclosure and the appended claims, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein is intended to mean and include any and all possible combinations of one or more of the associated listed items.
[0014] The terms "first", "second", "third", etc. may be used in this specification to describe various information, but it should be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information may be referred to as the second information, and similarly, the second information may be referred to as the first information. The term "case" when used in this specification may be understood to mean "when", "upon", or "depending on the judgment" depending on the context.
[0015] In October 2013, the first version of the HEVC standard was finalized, which provides about 50% bit-rate reduction or equivalent perceptual quality compared to the previous-generation video coding standard H.264 / MPEG AVC. The HEVC standard achieves a significant improvement in coding compared to previous standards, but it has been proven that better coding efficiency than HEVC can be achieved by using additional coding tools. Based on this, both VCEG and MPEG started exploring new coding technologies for future video coding standardization. In October 2015, the Joint Video Exploration Team (JVET) was formed by ITU-T VECG and ISO / IEC MPEG, and important research on advanced technologies that can enable a significant improvement in coding efficiency was started. By JVET, a reference software called the Joint Exploration Model (JEM) was maintained by integrating some additional coding tools on top of the HEVC Test Model (HM).
[0016] In October 2017, ITU-T and ISO / IEC issued a call for proposals (CfP) regarding video compression with capabilities beyond HEVC. In April 2018, 23 responses were received and evaluated at the 10th JVET meeting, which demonstrated that the compression efficiency was improved by about 40% compared to HEVC. Based on such evaluation results, JVET launched a new project to develop a new generation of video coding standard named Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate the reference implementation of the VVC standard.
[0017] Similar to HEVC, VVC is built based on a block-based hybrid video coding framework.
[0018] FIG. 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, FIG. 1 shows a typical encoder 100. The encoder 100 includes a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.
[0019] In encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter prediction method or an intra prediction method.
[0020] A prediction residual representing the difference between the current video block, which is part of the video input 110, and its predictor, which is part of the block predictor 140, is sent from adder 128 to transform 130. Then, for entropy reduction, the transform coefficients are sent from transform 130 to quantization 132. The quantized coefficients are then supplied to entropy coding 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 142 from the intra / inter mode decision 116, such as video block partitioning information, motion vectors (MVs), reference picture indexes, and intra prediction modes, is also supplied through entropy coding 138 and stored in the compressed bitstream 144. The compressed bitstream 144 includes the video bitstream.
[0021] In encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed through inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with the block predictor 140 to generate the unfiltered reconstructed pixels of the current video block.
[0022] Spatial prediction (or "intra prediction") predicts the current video block using samples (referred to as reference samples) from already encoded adjacent blocks within the same video frame as the current video block.
[0023] Temporal prediction (also called "inter prediction") predicts the current video block using reconstructed pixels from already encoded video pictures. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, one reference picture index is sent additionally to identify from which reference picture within the reference picture storage the temporal prediction signal is coming.
[0024] Motion estimation 114 fetches signals from the video input 110 and the picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 fetches signals from the video input 110, the picture buffer 120, and the motion estimation signal from motion estimation 114 and outputs a motion compensation signal to the intra / inter mode decision 116.
[0025] After spatial prediction and / or temporal prediction is performed, the intra / inter mode decision 116 within the encoder 100 selects the best prediction mode, for example, based on a rate distortion optimization method. Next, the block predictor 140 is subtracted from the current video block, and the resulting prediction residual is decorrelated using the transform 130 and quantization 132. The resulting quantized residual coefficients are inverse quantized by the inverse quantization 134 and inverse transformed by the inverse transform 136 to form a reconstructed residual, and then the reconstructed residual is added back to the prediction block to form the reconstructed signal of the CU. Before the reconstructed CU is placed in the reference picture storage of the picture buffer 120 and used to encode future video blocks, further in-loop filtering 122, such as a block deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), may be applied to the reconstructed CU. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and the quantized residual coefficients are all sent to the entropy coding unit 138, which further compresses and packs them to form a bitstream.
[0026] FIG. 1 shows a block diagram of a general block-based hybrid video coding system. The input video signal is processed block by block (referred to as a coding unit (CU)). In VTM-1.0, a CU may be up to 128×128 pixels. However, unlike HEVC that partitions blocks based only on quadtree, in VVC, one coding tree unit (CTU) is decomposed into CUs to adapt to various local characteristics based on quadtree / binary tree / trinary tree. By definition, a coding tree block (CTB) is a sample of an N×N block with a certain value N, and as a result, the division of components into CTBs is a partition. A CTU is encoded using a CTB of luminance samples, two corresponding CTBs of chrominance samples of a picture having three sample arrays, or a CTB of samples of a picture encoded using a syntax structure used to encode a monochrome picture or three separate color planes and samples. Further, the concept of multiple partition unit types in HEVC is eliminated, that is, in VVC, the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists, and instead, each CU is always used as a basic unit for both prediction and transform without further partitioning. In the polymorphic tree structure, first, one CTU is partitioned by a quadtree structure. Then, each quadtree leaf node can be further partitioned by a binary tree structure and a trinary tree structure. As shown in FIGS. 3A, 3B, 3C, 3D, and 3E, there are five decomposition types: quaternary partitioning, vertical binary partitioning, horizontal binary partitioning, horizontal trinary partitioning, and vertical trinary partitioning.
[0027] FIG. 3A shows a diagram illustrating block quaternary partitioning in the polymorphic tree structure according to the present disclosure.
[0028] FIG. 3B shows a diagram illustrating block vertical binary partitioning in the polymorphic tree structure according to the present disclosure.
[0029] FIG. 3C shows a diagram illustrating block horizontal binary partitioning in the polymorphic tree structure according to the present disclosure.
[0030] FIG. 3D shows a diagram illustrating a block vertical three-way division in the polymorphic tree structure according to the present disclosure.
[0031] FIG. 3E shows a diagram illustrating a block horizontal three-way division in the polymorphic tree structure according to the present disclosure.
[0032] In FIG. 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") predicts the current video block using samples (referred to as reference samples) of already encoded adjacent blocks within the same video picture / slice. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also called "inter prediction" or "motion compensated prediction") predicts the current video block using reconstructed pixels from already encoded video pictures. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, when multiple reference pictures are supported, an additional reference picture index is sent to identify from which reference picture in the reference picture store the temporal prediction signal is coming. After spatial prediction and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. Then, the prediction block is subtracted from the current video block, and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, and then the reconstructed residual is added back to the prediction block to form the reconstructed signal of the CU. Before the reconstructed CU is put into the reference picture store for use in encoding future video blocks, further in-loop filtering such as block deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) may be applied to the reconstructed CU. To form the output video bitstream, the encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit and further compressed and packed to form the bitstream.
[0033] Figure 2 shows a schematic block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of a typical decoder 200. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.
[0034] The decoder 200 is similar to the reconstruction-related section present in the encoder 100 of FIG. 1. In the decoder 200, first, the incoming input video bitstream 210 is decoded through entropy decoding 212 to derive quantized coefficient levels and prediction-related information. Next, the quantized coefficient levels are processed through inverse quantization 214 and inverse transform 216 to obtain the reconstructed prediction residuals. The block predictor mechanism implemented in the intra / inter mode selector 220 is configured to perform either intra prediction 222 or motion compensation 224 based on the decoded prediction information. By summing the reconstructed prediction residuals from the inverse transform 216 and the prediction output generated by the block predictor mechanism using the adder 218, a set of non-filtered reconstructed pixels is obtained.
[0035] The reconstructed blocks may further pass through the in-loop filter 228 before being stored in the picture buffer 226 that functions as a reference picture store. The reconstructed video in the picture buffer 226 is sent to drive a display device and may also be used to predict future video blocks. In situations where the in-loop filter 228 is on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0036] Figure 2 shows a schematic block diagram of a block-based video decoder. First, in the entropy decoding unit, the video bitstream is entropy decoded. The coding mode and prediction information are sent to the spatial prediction unit (in the case of intra coding) or the temporal prediction unit (in the case of inter coding) to form a prediction block. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual block. Then, the prediction block and the residual block are added. The reconstructed block may further pass through in-loop filtering before being stored in the reference picture buffer. Next, the reconstructed video in the reference picture buffer is sent to drive the display device and is used to predict future video blocks.
[0037] Generally, the basic intra prediction method applied in VVC remains the same as that of HEVC, except that several modules are further extended and / or improved. For example, the intra sub-partition (ISP) coding mode, extended intra prediction by wide-angle intra directions, position-dependent intra prediction combination (PDPC), and 4-tap intra interpolation.
[0038] Partitioning of Pictures, Tile Groups, Tiles, and CTUs in VVC In VVC, a tile is defined as a rectangular region of CTUs within a specific tile column and a specific tile row within a picture. A tile group is a group of an integer number of tiles of a picture that is exclusively included in a single NAL unit. Basically, the concept of a tile group is the same as that of a slice defined in HEVC. For example, a picture is divided into tile groups and tiles. A tile is a sequence of CTUs that covers a rectangular region of a picture. A tile group contains several tiles of a picture. Two modes of tile groups, namely the raster scan tile group mode and the rectangular tile group mode, are supported. In the raster scan tile group mode, a tile group contains a sequence of tiles in the tile raster scan of a picture. In the rectangular tile group mode, a tile group contains several tiles of a picture that collectively form a rectangular region of a picture. Tiles within a rectangular tile group are in the order of the tile raster scan of the tile group.
[0039] Figure 4 shows an example of raster scan tile group segmentation of a picture, where the picture is divided into 12 tiles and 3 raster scan tile groups. Figure 4 includes tiles 410, 412, 414, 416, and 418. Each tile has 18 CTUs. More specifically, Figure 4 shows a picture with 18×12 luma CTUs segmented into 12 tiles and 3 tile groups (for reference). The 3 tile groups are as follows: (1) the first tile group includes tiles 410 and 412, (2) the second tile group includes tiles 414, 416, 418, 420, and 422, and (3) the third tile group includes tiles 424, 426, 428, 430, and 432.
[0040] FIG. 5 shows an example of rectangular tile group partitioning of a picture, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows), as well as 9 rectangular tile groups. FIG. 5 includes tiles 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, and 556. More specifically, FIG. 5 shows a picture having an 18×12 luminance CTU partitioned into 24 tiles and 9 tile groups (for reference). A tile group contains tiles, and a tile contains a CTU. The 9 rectangular tile groups include (1) 2 tiles 510 and 512, (2) 2 tiles 514 and 516, (3) 2 tiles 518 and 520, (4) 4 tiles 522, 524, 534, and 536, (5) 4 tile groups 526, 528, 538, and 540, (6) 4 tiles 530, 532, 542, and 544, (7) 2 tiles 546 and 548, (8) 2 tiles 550 and 552, (9) 2 tiles 554 and 556.
[0041] Conversion of Large Block Sizes with High-Frequency Zero Setting in VVC In VTM4, conversion of large block sizes up to 64×64 is possible, and this conversion is mainly useful for high-resolution videos, such as 1080p and 4K sequences. The high-frequency conversion coefficients of a conversion block whose size (width or height, or both width and height) is equal to 64 are set to zero, and as a result, only the low-frequency coefficients are retained. For example, in the case of an M×N conversion block with block width M and block height N, when M is equal to 64, only the left 32 columns of the conversion coefficients are retained. Similarly, when N is equal to 64, only the top 32 rows of the conversion coefficients are retained. When the conversion skip mode is used for a large block, the entire block is used without setting the values to zero.
[0042] Virtual Pipeline Data Unit (VPDU) in VVC The virtual pipeline data unit (VPDU) is defined as a non-overlapping unit within a picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. The VPDU size is approximately proportional to the buffer size in most pipeline stages, and thus it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) partition and the binary tree (BT) partition may cause an increase in the VPDU size.
[0043] To keep the VPDU size as 64×64 luminance samples, the following canonical partition restrictions (with syntax signaling modifications) are applied to VTM5.
[0044] TT decomposition is not allowed for a CU where either the width or the height, or both the width and the height, are equal to 128.
[0045] For a 128×N CU where N≤64 (i.e., the width is equal to 128 and the height is less than 128), horizontal BT is not allowed.
[0046] For an N×128 CU where N≤64 (i.e., the height is equal to 128 and the width is less than 128), vertical BT is not allowed.
[0047] Figures 6A, 6B, 6C, 6D, 6E, 6F, 6G, and 6H show examples of non-allowed TT partitioning and BT partitioning in VTM.
[0048] Transform coefficient coding in VVC The conversion coefficient coding in VVC is similar to HEVC in that it uses both non-overlapping coefficient groups (also called CGs or sub-blocks). However, there are also some differences between them. In HEVC, each CG of coefficients has a fixed size of 4×4. In VVC Draft 6, the CG size has become dependent on the TB size. As a result, various CG sizes (1×16, 2×8, 8×2, 2×4, 4×2, and 16×1) are available in VVC. The CGs inside the coding block and the conversion coefficients within the CGs are coded according to a predefined scanning order.
[0049] To limit the maximum number of context-coded bins per pixel, the area of the TB and the type of video component (e.g., luminance component vs. chrominance component) are used to derive the maximum number of context-coded bins (CCBs: context-coded bins) of the TB. The maximum number of context-coded bins is equal to TB_zosize * 1.75. Here, TB_zosize indicates the number of samples in the TB after coefficient zeroing. Note that the coded_sub_block_flag, which is a flag indicating whether the CG contains non-zero coefficients, is not considered in the CCB count.
[0050] Coefficient zeroing is an operation performed on the conversion block to force the coefficients located in a specific region of the conversion block to 0. For example, in the current VVC, the 64×64 conversion has associated zeroing operations. As a result, all the conversion coefficients located outside the upper-left 32×32 region within the 64×64 conversion block are forced to 0. In fact, in the current VVC, for any conversion block having a size exceeding 32 along a particular dimension, coefficient zeroing operations are performed along that dimension to force the coefficients located beyond the upper-left 32×32 region to 0.
[0051] In the conversion coefficient coding in VVC, first, the variable remBinsPass1 is set to the maximum number of context-coded bins (MCCB: maximum number of context-coded bins) allowed. In the coding process, the variable decreases by 1 each time a context-coded bin is signaled. While remBinsPass1 is 4 or more, the coefficient is first signaled through the syntax of sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag, all of which use context-coded bins in the first pass. The remaining part of the coefficient level information is coded in the second pass using Golomb-Rice codes and bypass-coded bins with the syntax element of abs_remainder. When remBinsPass1 becomes less than 4 while coding the first pass, the current coefficient is not coded in the first pass and is directly coded in the second pass using Golomb-Rice codes and bypass-coded bins with the syntax element of dec_abs_level. The Rice parameter derivation process of dec_abs_level[] is derived as specified in Table 3. After coding all levels described above, finally, the sign (sign_flag) of all scan positions where sig_coeff_flag is equal to 1 is coded as a bypass bin. Such a process is depicted in Figure 7. remBinsPass1 is reset for each TB. The transition from the use of context-coded bins of sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag to the use of bypass-coded bins for the remaining coefficients occurs only once per TB at most. For a coefficient sub-block, if remBinsPass1 is less than 4 before coding the first coefficient, the entire coefficient sub-block is coded using bypass-coded bins.
[0052] Figure 7 shows a diagram of the residual coding structure of the conversion block.
[0053] To signal the syntax of abs_remainder and dec_abs_level, the derivation of a unified (same) Rice parameter (RicePara) is used. The only difference is that the reference levels baseLevel are set to 4 and 0 respectively for encoding abs_remainder and dec_abs_level. The Rice parameter is determined based not only on the sum of the absolute levels of five adjacent transform coefficients within the local template, but also on the corresponding reference levels as follows.
[0054] RicePara=RiceParTable[max(min(31,sumAbs-5*baseLevel),0)]
[0055] The syntax of residual coding and the related semantics in the current VVC draft specification are shown in Table 1 and Table 2 respectively. The way to read Table 1 is shown in the attached section of this disclosure and is also described in the VVC specification.
Table 1
Table 2
Table 3
Table 4
[0056] Residual Coding in the Conversion Skip Mode in VVC Unlike HEVC which is designed such that a single residual coding method encodes both transform coefficients and transform skip coefficients, in VVC, two separate residual coding methods are used for transform coefficients and transform skip coefficients (i.e., residuals), respectively.
[0057] In the transform skip mode, the statistical characteristics of the residual signal are different from those of the transform coefficients, and no energy compression around the low-frequency components is observed. Residual coding no signaling of the last x / y position, when all previous flags are equal to 0, coded_sub_block_flag is coded for all sub-blocks except the DC sub-block, sig_coeff_flag context modeling using two adjacent coefficients, par_level_flag using only one context model, adding flags exceeding 5, 7, 9, modified Rice parameter derivation for residual binarization, context modeling of the sign flag is determined based on the left and upper adjacent coefficient values, and after sig_coeff_flag, the sign flag is parsed to hold all context-coded bits together is modified considering various signal characteristics of the (spatial) transform skip residual.
[0058] As shown in FIG. 8, in the first pass, the syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag are coded in an interleaved manner for each residual sample, followed by the coding of the abs_level_gtX_flag bitplane in the second pass and the coding of abs_remainder in the third pass.
[0059] Pass 1: sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag
[0060] Path 2: abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, abs_level_gt9_flag
[0061] Path 3: abs_remainder
[0062] FIG. 8 shows a diagram of the residual coding structure of the conversion skip block.
[0063] The syntax and related semantics of the residual coding of the conversion skip mode in the current VVC draft specification are shown in Table 5 and Table 2, respectively. The reading method of Table 5 is shown in the attached section of this disclosure and is also described in the VVC specification.
Table 5
[0064] Quantization In the current VVC, the maximum QP value is extended from 51 to 63, and accordingly, the signaling of the initial QP is changed. When a non-zero value of slice_qp_delta is coded, the initial value of SliceQpY can be modified in the slice segment layer. For the conversion skip block, when QP is equal to 4, the size of the quantization step is 1, so the minimum allowable quantization parameter (QP) is defined as 4.
[0065] Furthermore, the same HEVC scalar quantization is used together with a new concept called dependent scalar quantization. Dependent scalar quantization refers to a technique in which the set of allowable reconstruction values of the transform coefficients depends on the values of the transform coefficient levels prior to the current transform coefficient level in the reconstruction order. The main effect of this technique is that, compared to the conventional independent scalar quantization used in HEVC, the allowable reconstruction vectors are packed more densely in the N-dimensional vector space (N represents the number of transform coefficients in the transform block). That is, for a given average number of allowable reconstruction vectors per unit volume in the N-dimension, the average distortion between the input vector and the closest reconstruction vector is reduced. The technique of dependent scalar quantization is realized by (a) defining two scalar quantizers using different reconstruction levels, and (b) defining a process for switching between the two scalar quantizers.
[0066] The two scalar quantizers used are shown in FIG. 9 in the states indicated by Q0 and Q1. The positions of the available reconstruction levels are uniquely specified by the quantization step size Δ. The scalar quantizer used (Q0 or Q1) is not explicitly signaled in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the transform coefficient levels preceding the current transform coefficient in the encoding / reconstruction order.
[0067] FIG. 9 shows a diagram of the two scalar quantizers used in the proposed dependent quantization technique.
[0068] As shown in FIGS. 10A and 10B, the switching between two scalar quantizers (Q0 and Q1) is realized via a state machine having four quantizer states (QState). QState can take four different values of 0, 1, 2, and 3. This is uniquely determined by the parity of the transform coefficient levels preceding the current transform coefficient in the encoding / reconstruction order. At the start of the inverse quantization of the transform block, the state is set to be equal to 0. The transform coefficients are reconstructed in scan order (i.e., in the same order as they are entropy decoded). After the current transform coefficient is reconstructed, the state is updated as shown in FIG. 10, where k indicates the value of the transform coefficient level.
[0069] FIG. 10A shows a transition diagram showing state transitions for the proposed dependent quantization.
[0070] FIG. 10B shows a table showing quantizer selection for the proposed dependent quantization.
[0071] Transmitting a default user-defined scaling matrix is also supported. The scaling matrices in the DEFAULT mode are all flat, and the elements are equal to 16 for all TB sizes. The IBC and intra coding modes currently share the same scaling matrix. Therefore, for the USER_DEFINED matrix, the numbers of MatrixType and MatrixType_DC are updated as follows.
[0072] MatrixType:30 = 2 (2 for intra & IBC / inter) × 3 (Y / Cb / Cr components) × 5 (square TB sizes: 4×4 to 64×64 for luminance, 2×2 to 32×32 for chrominance).
[0073] MatrixType_DC:14 = 2 (in the case of Intra & IBC / Inter for the 2×Y components, 1 in the case of 1) × 3 (TB sizes: 16×16, 32×32, 64×64) + 4 (in the case of Intra & IBC / Inter for the 2×Cb / Cr components, 2 in the case of 2) × 2 (TB sizes: 16×16, 32×32).
[0074] The DC values are encoded separately for the following scaling matrices: 16×16, 32×32, and 64×64. For TBs smaller than 8×8 in size, all elements within one scaling matrix are signaled. For TBs with a size of 8×8 or larger, only 64 elements within one 8×8 scaling matrix are signaled as the basic scaling matrix. To obtain a square matrix of a size larger than 8×8, the 8×8 basic scaling matrix is upsampled to the corresponding square size (i.e., 16×16, 32×32, 64×64) (by replication of elements). When zero - setting of the high - frequency coefficients of the 64 - point transform is applied, the corresponding high - frequencies of the scaling matrix are also zero - set. That is, when the width or height of the TB is 32 or more, only the left or upper half of the coefficients is retained, and 0 is assigned to the remaining coefficients. Further, since the lower - right 4×4 elements are never used, the number of elements signaled for the 64×64 scaling matrix is also reduced from 8×8 to three 4×4 sub - matrices.
[0075] Context Modeling for Transform Coefficient Encoding The selection of the probability model for the syntax elements related to the absolute value of the transform coefficient level depends on the absolute level in the local neighborhood or the value of the partially reconstructed absolute level. The templates used are illustrated in Figure 11.
[0076] Figure 11 shows a diagram of the templates used to select the probability model. The black squares specify the current scan position, and the squares with "x" represent the local neighborhood used.
[0077] The selected probability model depends on the sum of the absolute levels (or partially reconstructed absolute levels) within the local neighborhood, and the number of absolute levels greater than 0 within the local neighborhood (given by the number of sig_coeff_flags equal to 1). Context modeling and binarization are based on the following measurements for the local neighborhood, namely, numSig: the number of non-zero levels within the local neighborhood, sumAbs1: the sum of the partially reconstructed absolute levels (absLevel1) after the first pass within the local neighborhood, sumAbs: the sum of the reconstructed absolute levels within the local neighborhood, The diagonal position (d) depends on the sum of the horizontal and vertical coordinates of the current scan position within the transform block.
[0078] Based on the values of numSig, sumAbs1, and d, a probability model is selected for encoding sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag. The Rice parameters for binarizing abs_remainder and dec_abs_level are selected based on the values of sumAbs and numSig.
[0079] In the current VVC, the reduced 32-point MTS (also called RMTS32) is based on skipping high-frequency coefficients and is used to reduce the computational complexity of 32-point DST-7 / DCT-8. Also, this involves a change in coefficient coding that includes all types of zero settings (i.e., RMTS32 and the existing zero settings for high-frequency components in DCT2). Specifically, the binarization of the last non-zero coefficient position coding is coded based on the reduced TU size, and the context model selection for the last non-zero coefficient position coding is determined by the original TU size. Furthermore, 60 context models are used to code the sig_coeff_flag of the transform coefficients. The selection of the context model index is based on the sum of five absolute levels that were previously partially reconstructed, called locSumAbsPass1, and the state of the dependent quantization, QState, as follows.
[0080] When cIdx is equal to 0, ctxInc is derived as follows. ctxInc = 12 * Max(0, QState - 1) + Min((locSumAbsPass1 + 1) >> 1, 3) + (d < 2? 8 : (d < 5? 4 : 0))
[0081] Otherwise (when cIdx is greater than 0), ctxInc is derived as follows. ctxInc = 36 + 8 * Max(0, QState - 1) + Min((locSumAbsPass1 + 1) >> 1, 3) + (d < 2? 4 : 0)
[0082] Palette mode The basic idea behind the palette mode is that the samples within a CU are represented by a small set of representative color values. This set is called a palette. It is also possible to indicate the colors excluded from the palette by signaling the color values excluded from the palette as escape colors whose by values of the three color components are directly signaled in the bitstream. This is illustrated in Figure 12.
[0083] Figure 12 shows an example of a block encoded in palette mode. Figure 12 includes a block 1210 encoded in palette mode and a palette 1220.
[0084] In Figure 12, the palette size is 4. The first three samples use palette entries 2, 0, and 3 respectively for reconstruction. The blue sample represents an escape symbol. The CU-level flag palette_escape_val_present_flag indicates whether an escape symbol exists within the CU. If an escape symbol exists, the palette size is increased by 1 only, and the last index is used to indicate the escape symbol. Thus, in Figure 12, index 4 is assigned to the escape symbol.
[0085] To decode a palette-encoded block, the decoder needs the following information, namely the palette table, the palette index and has.
[0086] If the palette index corresponds to an escape symbol, additional overhead is signaled to indicate the corresponding color value of the sample.
[0087] Furthermore, on the encoder side, it is necessary to derive the appropriate palette used for that CU.
[0088] In the case of deriving a palette for irreversible symbolization, a modified k - means clustering algorithm is used. The first sample of the block is added to the palette. Then, for each subsequent sample from the block, the sum of absolute differences (SAD) between the sample and each of the current palette colors is calculated. If the distortion of each component is less than the threshold of the palette entry corresponding to the minimum SAD, the sample is added to the cluster belonging to that palette entry. Otherwise, the sample is added as a new palette entry. When the number of samples mapped to a cluster exceeds the threshold, the centroid of that cluster is updated to become the palette entry for that cluster.
[0089] In the next step, the clusters are sorted in descending order of usage. Then, the palette entries corresponding to each entry are updated. Usually, the cluster centroid is used as the entry for the palette. However, considering the cost of encoding the palette entry, rate - distortion analysis is performed to analyze whether any entry from the palette predictor rather than the centroid might be more suitable to be used as the updated palette entry. This process continues until all clusters are processed or the maximum palette size is reached. Finally, if a cluster has only a single sample and the corresponding palette entry is not in the palette predictor, the sample is converted to an escape symbol. Additionally, duplicate palette entries are removed and their clusters are merged.
[0090] After palette derivation, each sample in the block is assigned the index of the nearest palette entry (within the SAD). Then the sample is assigned to either the "INDEX" or "COPY_ABOVE" mode. For each sample for which the "INDEX" or "COPY_ABOVE" mode is possible, the cost of encoding the mode is then calculated. The mode with the lower cost is selected.
[0091] For the encoding of palette entries, a palette predictor is maintained. The maximum size of the palette and the palette predictor are signaled in the SPS. The palette predictor is initialized at the start of each CTU row, each slice, and each tile.
[0092] For each entry of the palette predictor, a reuse flag indicating whether it is part of the current palette is signaled. This is illustrated in FIG. 13.
[0093] FIG. 13 shows the use of the palette predictor for signaling palette entries. FIG. 13 includes a previous palette 1310 and a current palette 1320.
[0094] The reuse flag is sent using zero run-length coding. After this, the number of new palette entries is signaled using exponential Golomb coding of order 0. Finally, the component values of the new palette entries are signaled.
[0095] Palette indices are encoded using horizontal and vertical motion scans as shown in FIGS. 14A and 14B. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag.
[0096] FIG. 14A shows a horizontal motion scan.
[0097] FIG. 14B shows a vertical motion scan.
[0098] To encode the palette indices, a line coefficient group (CG)-based palette mode is used in which the CU is divided into a plurality of segments having 16 samples based on the motion scan mode as shown in FIGS. 15A and 15B, and the index run, the palette index value, and the quantized color for the escape mode are sequentially encoded / parsed for each CG.
[0099] FIG. 15A shows a sub-block based index map scan for a palette.
[0100] FIG. 15B shows a sub-block based index map scan for a palette.
[0101] Palette indices are encoded using two main palette sample modes, "INDEX" and "COPY_ABOVE". As previously explained, an index equal to the maximum palette size is assigned to the escape symbol. In the "COPY_ABOVE" mode, the palette index of the sample in the above row is replicated. In the "INDEX" mode, the palette index is explicitly signaled. The encoding order of the palette run encoding in each segment is as follows.
[0102] For each pixel, one context encoding bin run_copy_flag = 0 is signaled, indicating whether the pixel is in the same mode as the previous pixel, i.e., whether both the previous scanned pixel and the current pixel are of execution type COPY_ABOVE, or whether both the previous scanned pixel and the current pixel are of execution type INDEX and have the same index value. Otherwise, run_copy_flag = 1 is signaled.
[0103] When the pixel and the previous pixel are in different modes, a context - encoded bit copy_above_palette_indices_flag indicating the execution type of the pixel, i.e., INDEX or COPY_ABOVE, is signaled. Since the INDEX mode is used by default, when the sample is in the first row (horizontal scan) or the first column (vertical scan), the decoder does not need to analyze the execution type. Also, when the previously analyzed execution type is COPY_ABOVE, the decoder does not need to analyze the execution type. After the palette execution encoding of the pixels within one segment, the index value (palette_idx_idc) and the quantized escape color (palette_escape_val) in the INDEX mode are bypass - encoded.
[0104] Improvements to the coding of residuals and coefficients In VVC, when encoding transform coefficients, a unified (same) Rice parameter (RicePara) derivation is used to signal the syntax of abs_remainder and dec_abs_level. The only difference is that the reference levels baseLevel are set to 4 and 0 respectively for encoding abs_remainder and dec_abs_level. The Rice parameter is determined based not only on the sum of the absolute levels of five adjacent transform coefficients within the local template but also on the corresponding reference levels as follows.
[0105] RicePara = RiceParTable[max(min(31,sumAbs - 5*baseLevel),0)]
[0106] In other words, the binary - coded words of the syntax elements abs_remainder and dec_abs_level are adaptively determined according to the level information of adjacent coefficients. Since this determination of the coded word is performed for each sample, additional logic is required to handle this coded - word adaptation for coefficient coding.
[0107] Similarly, when encoding a residual block in transform skip mode, the binary codeword of the syntax element abs_remainder is adaptively determined according to the level information of adjacent residual samples.
[0108] Furthermore, when encoding syntax elements related to residual coding or transform coefficient coding, the selection of the probability model depends on the level information of adjacent levels, which requires additional logic and additional context models.
[0109] In the current design, the binarization of escape samples is derived by invoking a tertiary exponential Golomb binarization process. There is still room for improvement in its performance.
[0110] In the current VVC, two different level mapping methods are available, which are applied to normal transform and transform skip respectively. Each level mapping method is associated with different conditions, mapping functions, and mapping positions. For blocks to which normal transform is applied, after the number of context coding bins (CCBs) exceeds the limit, a certain level mapping method is used. The mapping position indicated as ZeroPos[n] and the mapping result indicated as AbsLevel[xC][yC] are derived as specified in Table 2. For blocks to which transform skip is applied, before the number of context coding bins (CCBs) exceeds the limit, another level mapping method is used. The mapping position indicated as predCoeff and the mapping result indicated as AbsLevel[xC][yC] are derived as specified in Table 5. Such a non-unified design may not be optimal from the perspective of standardization.
[0111] For profiles beyond 10 bits in HEVC, the fact that the extended_precision_processing_flag is equal to 1 specifies that an extended dynamic range is used for coefficient analysis and inverse transform processing. In the current VVC, residual coding or transform skip coding of transform coefficients beyond 10 bits has been reported as a cause of significant performance degradation. There is room for further improvement in its performance.
[0112] Proposed method In the present disclosure, several methods are proposed to address the issues described in the section on improvements to residual and coefficient coding. It should be noted that the following methods may be applied alone or in combination.
[0113] According to a first aspect of the present disclosure, it is proposed to use a fixed set of binary codewords to code a specific syntax element, for example abs_remainder, in residual coding. The binary codewords can be formed using various methods. Some exemplary methods are listed as follows.
[0114] First, the same procedure as that used in the current VVC is used to determine the codeword of abs_remainder, but a fixed Rice parameter (for example, 1, 2, or 3) is always selected.
[0115] Second, it is fixed-length binarization.
[0116] Third, it is truncated Rice binarization.
[0117] Fourth, it is a truncated Binary (TB) binarization process.
[0118] Fifth, it is a k-th exponential Golomb binarization process (EGk).
[0119] Sixth, it is limited k-th exponential Golomb binarization.
[0120] According to a second aspect of the present disclosure, it is proposed to use a fixed set of codewords to encode certain syntax elements, such as abs_remainder and dec_abs_level, in transform coefficient coding. The binary codewords can be formed using various methods. Some exemplary methods are listed as follows.
[0121] First, to determine the codewords for abs_remainder and dec_abs_level, the same procedure as used in the current VVC is used, but a fixed Rice parameter, such as 1, 2, or 3, is used. As in the current VVC, the value of baseLevel can still be different for abs_remainder and dec_abs_level (for example, to encode abs_remainder and dec_abs_level, baseLevel is set to 4 and 0 respectively).
[0122] Second, to determine the codewords for abs_remainder and dec_abs_level, the same procedure as used in the current VVC is used, but a fixed Rice parameter, such as 1, 2, or 3, is used. The values of baseLevel for abs_remainder and dec_abs_level are selected to be the same, for example, both use 0 or both use 4.
[0123] Third, it is fixed-length binarization.
[0124] Fourth, it is truncated Rice binarization.
[0125] Fifth, it is a truncated binary (TB) binarization process.
[0126] Sixth, it is a k-th exponential Golomb binarization process (EGk).
[0127] Seventhly, it is limited k-th exponential Golomb binarization.
[0128] According to a third aspect of the present disclosure, it is proposed to use a single context for encoding syntax elements related to residual coding or coefficient coding (e.g., abs_level_gtx_flag), and context selection based on adjacent decoded level information can be removed.
[0129] According to a fourth aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode a specific syntax element, e.g., abs_remainder, in residual coding, and the selection of the set of binary codewords is determined according to specific encoded information of the current block, e.g., quantization parameter (QP) associated with TB / CB and / or slice, prediction mode of the CU (e.g., IBC mode or intra or inter), and / or slice type (e.g., I slice, P slice, or B slice). Various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed as follows.
[0130] First, the same procedure as that currently used in VVC is used to determine the codeword of bs_remainder, but different Rice parameters are used.
[0131] Second, it is a k-th exponential Golomb binarization process (EGk).
[0132] Third, it is limited k-th exponential Golomb binarization. [Table 6]
[0133] The same method as that described in the fourth aspect is used for transformation. CoefficientIt is also applicable to symbolization. According to the fifth aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode specific syntax elements, such as abs_remainder and dec_abs_level, in transform coefficient coding. The selection of the set of binary codewords is determined according to specific encoded information of the current block, such as quantization parameters (QP) associated with TB / CB and / or slices, the prediction mode of the CU (e.g., IBC mode or intra or inter), and / or slice type (e.g., I slice, P slice, or B slice). Also in this case, various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed as follows.
[0134] First, the same procedure as that currently used in the VVC is used to determine the codeword of bs_remainder, but different Rice parameters are used.
[0135] Second, it is the k-th exponential Golomb binarization process (EGk).
[0136] Third, it is limited k-th exponential Golomb binarization.
[0137] In these above methods, different Rice parameters may be used to derive different sets of binary codewords. For a given block of residual samples, the Rice parameter used is determined according to the CU QP shown as QPCU, rather than adjacent level information. As shown in Table 6, a specific example is illustrated, where TH1 to TH4 are predetermined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predetermined Rice parameters. It is worth noting that the same logic can actually be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter from the QP value of the current CU as shown in Table 6.
[0138] According to a fifth aspect of the present disclosure, a set of parameters and / or thresholds related to codeword determination for syntax elements of transform coefficient coding and / or transform skip residual coding is signaled within a bitstream. The determined codeword is used as a binary codeword when encoding the syntax element through an entropy encoder, such as arithmetic coding.
[0139] It should be noted that the set of parameters and / or thresholds can be a complete set or a subset of all parameters and thresholds related to codeword determination of the syntax elements. The set of parameters and / or thresholds can be signaled at various levels in a video bitstream. For example, they can be signaled at the sequence level (e.g., sequence parameter set), picture level (e.g., picture parameter set, and / or picture header), slice level (e.g., slice header), or coding unit (CU) level in a coding tree unit (CTU).
[0140] In one example, the Rice parameter used to determine the codeword for encoding the abs_remainder syntax in transform skip residual coding is signaled in the slice header, picture header, PPS, and / or SPS. The signaled Rice parameter is used to determine the codeword for encoding the syntax abs_remainder when the CU is encoded in transform skip mode and the CU is associated with the above-mentioned slice header, picture header, PPS, and / or SPS.
[0141] According to a sixth aspect of the present disclosure, a set of parameters and / or thresholds related to the codeword determination shown in the first and second aspects is used for the syntax elements of transform coefficient coding and / or transform skip residual coding. Also, different sets may be used depending on whether the current block contains a luminance residual / coefficient or a chrominance residual / coefficient. The determined codeword is used as a binarized codeword when encoding the syntax element through an entropy encoder, for example, arithmetic coding.
[0142] In one example, the codeword of abs_remainder associated with the transform residual coding used in the current VVC is used for both luminance blocks and chrominance blocks, but different fixed Rice parameters are used by the luminance block and the chrominance block, respectively. (For example, K1 for the luminance block and K2 for the chrominance block, where K1 and K2 are integers)
[0143] According to a seventh aspect of the present disclosure, a set of parameters and / or thresholds related to codeword determination for the syntax elements of transform coefficient coding and / or transform skip residual coding is signaled in the bitstream. Also, different sets may be signaled for the luminance block and the chrominance block. The determined codeword is used as a binarized codeword when encoding the syntax element through an entropy encoder, for example, arithmetic coding.
[0144] The same method as described in the above aspect is also applicable to the escape value coding in the palette mode, for example, palette_escape_val.
[0145] According to an eighth aspect of the present disclosure, different k-th order exponential Golomb binarizations may be used to derive different sets of binary codewords for encoding escape values in the palette mode. In one example, for a given block of escape samples, the exponential Golomb parameter used, that is, the value k, is the QP CUIt is determined according to the QP value of the block shown as. When deriving the value of parameter k based on a given QP value of a block, the same example as that shown in Table 6 can be used. In that example, four different thresholds (from TH1 to TH4) are listed, and based on these thresholds and QP CU up to five different k values (from K0 to K4) can be derived, but it is worth mentioning that the number of thresholds is only for the purpose of explanation. In practice, different numbers of thresholds may be used to divide the entire QP value range into different numbers of QP value segments, and for each QP value segment, different k values may be used to derive the corresponding binary codeword for encoding the escape value of the block encoded in palette mode. It is also worth noting that the same logic can actually be implemented in different ways. For example, for example, a specific equation or look-up table may be used to derive the same Rice parameter.
[0146] According to a ninth aspect of the present disclosure, a set of parameters and / or thresholds related to codeword determination for the syntax element of the escape sample is signaled in the bitstream. The determined codeword is used as a binary codeword when encoding the syntax element through an entropy encoder, for example, arithmetic coding.
[0147] It should be noted that the set of parameters and / or thresholds can be a complete set or a subset of all parameters and thresholds related to codeword determination of the syntax element. The set of parameters and / or thresholds can be signaled at various levels in the video bitstream. For example, they can be signaled at the sequence level (e.g., sequence parameter set), picture level (e.g., picture parameter set, and / or picture header), slice level (e.g., slice header) at the coding tree unit (CTU) level, or at the coding unit (CU) level.
[0148] In one example according to this aspect, in order to determine a codeword for encoding the palate_escape_val syntax in palette mode, k-th order exponential Golomb binarization is used, and the value k is signaled to the decoder in the bitstream. The value k may be signaled at various levels. For example, the value k may be signaled in a slice header, a picture header, a PPS, and / or an SPS, etc. The signaled exponential Golomb parameter is used to determine a codeword for encoding the syntax palette_escape_val when the CU is encoded in palette mode and the CU is associated with the above-mentioned slice header, picture header, PPS, and / or SPS, etc.
[0149] Harmonization of level mapping for transform skip mode and normal transform mode According to a tenth aspect of the present disclosure, the same conditions for applying level mapping are used for both the transform skip mode and the normal transform mode. In one example, it is proposed to apply level mapping after the number of context coding bins (CCBs) exceeds a limit for both the transform skip mode and the normal transform mode. In another example, it is proposed to apply level mapping before the number of context coding bins (CCBs) exceeds a limit for both the transform skip mode and the normal transform mode.
[0150] According to an eleventh aspect of the present disclosure, the same method for deriving the mapping position in level mapping is used for both the transform skip mode and the normal transform mode. In one example, it is proposed to apply the method for deriving the mapping position in the level mapping used under the transform skip mode to the normal transform mode as well. In another example, it is proposed to apply the method for deriving the mapping position in the level mapping used under the normal transform mode to the transform skip mode as well.
[0151] According to a twelfth aspect of the present disclosure, the same level mapping method is applied to both the conversion skip mode and the normal conversion mode. In one example, it is proposed to apply the level mapping function used under the conversion skip mode also to the normal conversion mode. In another example, it is proposed to apply the level mapping function used under the normal conversion mode also to the conversion skip mode.
[0152] Simplification of Rice parameter derivation in residual coding According to a thirteenth aspect of the present disclosure, for the derivation of the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level using the Golomb-Rice code, it is proposed to use simple logic such as shift operations or division operations instead of a look-up table. According to the present disclosure, the look-up table as specified in Table 4 may be removed. In one example, the Rice parameter cRiceParam is derived as cRiceParam=(locSumAbs>>n), where n is a positive number, for example 3. It is worth noting that in practice, other different logics, for example, division operations by a value equal to 2 to the power of n, may be used to achieve the same result. An example of the corresponding decoding process based on the VVC draft is shown below, with the changed parts indicated in bold italic font and the deleted content indicated in italic font.
Table 7
[0153] According to a fourteenth aspect of the present disclosure, for the derivation of the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level using the Golomb-Rice code, it is proposed to use fewer adjacent positions. In one example, it is proposed to use only two adjacent positions for the derivation of the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, with the changed parts indicated in bold italic font and the deleted content indicated in italic font.
Table 8
[0154] In another example, it is proposed to use only one adjacent position for the derivation of the Rice parameter when encoding the syntax element abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content in italic font.
Table 9
[0155] According to the 15th aspect of the present disclosure, it is proposed to use different logics for adjusting the value of locSumAbs based on the value of baseLevel for the derivation of the Rice parameter when encoding the syntax element abs_remainder / dec_abs_level using the Golomb-Rice code. In one example, additional scale operations and offset operations are applied in the form of "(locSumAbs - BaseLevel * 5) * alpha + beta". The corresponding decoding process based on the VVC draft when alpha takes the value of 1.5 and beta takes the value of 1 is shown below, with the changes indicated in bold italic font and the deleted content in italic font.
Table 10
[0156] According to the 16th aspect of the present disclosure, it is proposed to remove the clip operation for the derivation of the Rice parameter in the syntax element abs_remainder / dec_abs_level using the Golomb-Rice code. According to the current disclosure, an example of the decoding process of the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content in italic font.
Table 11
[0157] According to the present disclosure, an example of the decoding process of the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font. [Table 12]
[0158] According to the 17th aspect of the present disclosure, in order to derive the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level using the Golomb-Rice code, it is proposed to change the initial value of locSumAbs from 0 to a non-zero integer. In one example, an initial value of 1 is assigned to locSumAbs, and the corresponding decoding process based on the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font. [Table 13]
[0159] According to the 18th aspect of the present disclosure, in order to derive the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level using the Golomb-Rice code, it is proposed to use the maximum value of the adjacent position level values instead of the sum of the adjacent position level values. An example of the corresponding decoding process based on the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font. [Table 14]
[0160] According to a nineteenth aspect of the present disclosure, when encoding the syntax elements abs_remainder / dec_abs_level using a Golomb-Rice code, it is proposed to derive a Rice parameter based on the relative amplitudes of each AbsLevel value at adjacent positions and a reference level value. In one example, the Rice parameter is derived based on how many of the AbsLevel values at adjacent positions are greater than the reference level. An example of a corresponding decoding process based on the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content indicated in italic font.
Table 15
[0161] In another example, the Rice parameter is derived based on the sum of the (AbsLevel - BaseLevel) values at adjacent positions where the AbsLevel value is greater than the reference level. An example of a corresponding decoding process based on the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content indicated in italic font.
Table 16
[0162] According to the present disclosure, an example of the decoding process of the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content indicated in italic font.
Table 17
[0163] Simplification of level mapping position derivation in residual coding According to the 20th aspect of the present disclosure, it is proposed to remove QState from the derivation of ZeroPos[n] so that ZeroPos[n] is derived only from cRiceParam. An example of the corresponding decoding process based on the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font. [Table 18]
[0164] According to the 21st aspect of the present disclosure, it is proposed to derive ZeroPos[n] based on the value of locSumAbs. An example of the corresponding decoding process based on the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font. [Table 19]
[0165] According to the 22nd aspect of the present disclosure, it is proposed to derive ZeroPos[n] based on the values of AbsLevel at adjacent positions. In one example, ZeroPos[n] is derived based on the maximum value of AbsLevel[xC + 1][yC] and AbsLevel[xC][yC + 1]. An example of the corresponding decoding process based on the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font. [Table 20]
[0166] According to the 23rd aspect of the present disclosure, it is proposed to derive both cRiceParam and ZeroPos[n] based on the maximum value of all AbsLevel values at adjacent positions. An example of the corresponding decoding process based on the VVC draft is shown below, where the changed parts are indicated in bold italic font and the deleted content is indicated in italic font.
Table 21
[0167] The same method as described in the above aspect is also applicable to the derivation of predCoeff in the residual coding of the transform skip mode. In one example, the variable predCoeff is derived as follows.
[0168] predCoeff = Max(absLeftCoeff, absAboveCoeff) + 1
[0169] Residual coding of transform coefficients In the present disclosure, in order to address the problems pointed out in the "Improvements to Residual and Coefficient Coding" section, methods are provided for simplifying and / or further improving the existing design of residual coding. Generally, the main features of the techniques proposed in the present disclosure are summarized as follows.
[0170] First, adjust the Rice parameter derivation used under normal residual coding based on the current design.
[0171] Second, change the binary method used under normal residual coding.
[0172] Third, change the Rice parameter derivation used under normal residual coding.
[0173] Rice parameter derivation in residual coding based on the current design According to a 24th aspect of the present disclosure, it is proposed to use a variable method of Rice parameter derivation to encode a specific syntax element, such as abs_remainder / dec_abs_level, in residual encoding, and according to a quantization parameter or encoding bit depth associated with specific encoded information of the current block, such as TB / CB and / or slice / profile, and / or according to a new flag associated with the TB / CB / slice / picture / sequence level, such as an extended_precision_processing_flag, a selection result is determined. Various methods may be used to derive the Rice parameter, and some exemplary methods are listed below.
[0174] First, cRiceParam = (cRiceParam << a) + (cRiceParam >> b) + c, where a, b, and c are positive numbers, for example, {a, b, c} = {1, 1, 0}. It is worth noting that in practice, other different logics, such as a multiplication operation with a value equal to a power of 2, may be used to achieve the same result.
[0175] Second, cRiceParam = (cRiceParam << a) + b, where a and b are positive numbers, for example, {a, b} = {1, 1}. It is worth noting that in practice, other different logics, such as a multiplication operation with a value equal to a power of 2, may be used to achieve the same result.
[0176] Third, cRiceParam = (cRiceParam * a) + b, where a and b are positive numbers, for example, {a, b} = {1.5, 0}. It is worth noting that in practice, other different logics, such as a multiplication operation with a value equal to a power of 2, may be used to achieve the same result.
[0177] An example of the corresponding decoding process based on the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content in italic font. The changes to the VVC draft are shown in bold italic font in Table 22. It is worth noting that the same logic can actually be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter from the current CU / sequence BitDepth value.
Table 22
[0178] In another example, when the BitDepth is greater than or equal to a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is derived as cRiceParam = (cRiceParam << a) + (cRiceParam >> b) + c, where a, b, and c are positive numbers, e.g., 1. The corresponding decoding process based on the VVC draft is shown below, with the changes indicated in bold italic font and the deleted content in italic font. The changes to the VVC draft are shown in bold italic font in Table 23. It is worth noting that the same logic can actually be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter from the current CU / sequence BitDepth value.
Table 23
[0179] Binary method in residual coding for profiles exceeding 10 bits According to the 25th aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode a specific syntax element, such as abs_remainder / dec_abs_level, in residual coding, and according to a quantization parameter or coding bit depth associated with specific coded information of the current block, such as TB / CB and / or slice / profile, and / or according to a new flag associated with TB / CB / slice / picture / sequence level, such as extended_precision_processing_flag, the selection result is determined. Various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed below.
[0180] First, to determine the codeword of abs_remainder, the same procedure as that currently used in VVC is used, but a fixed Rice parameter (e.g., 2, 3, 4, 5, 6, 7, or 8) is always selected. The fixed value may vary under different conditions according to a quantization parameter or coding bit depth associated with specific coded information of the current block, such as TB / CB and / or slice / profile, and / or according to a syntax element associated with TB / CB / slice / picture / sequence level, such as rice_parameter_value. As shown in Table 24, a specific example is illustrated, where TH1 to TH4 are predetermined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predetermined Rice parameters. It is worth noting that the same logic can be implemented in different ways in practice. For example, a specific equation or look-up table may also be used to derive the same Rice parameter from the BitDepth value of the current CU / sequence as shown in Table 24.
[0181] Second, it is fixed-length binarization.
[0182] Third, it is truncated Rice binarization.
[0183] Fourthly, it is a truncation binary (TB) binarization process.
[0184] Fifthly, it is a k-th exponential Golomb binarization process (EGk).
[0185] Sixthly, it is a limited k-th exponential Golomb binarization. [Table 24]
[0186] In one example, when a new flag, for example extended_precision_processing_flag, is equal to 1, the Rice parameter cRiceParam is fixed as n, where n is a positive number (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value may vary under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, with the changed parts indicated in bold italic font and the deleted content indicated in italic font. The changed parts with respect to the VVC draft are indicated in bold italic font in Table 25. [Table 25]
[0187] In another example, when a new flag, for example extended_precision_processing_flag, is equal to 1, it is proposed to use only one fixed value for the Rice parameter when encoding the syntax elements abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, with the changed parts indicated in bold italic font and the deleted content indicated in italic font. The changed parts with respect to the VVC draft are indicated in bold italic font in Table 26. [Table 26]
[0188] In yet another example, when the BitDepth is greater than or equal to a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is fixed as n, where n is a positive number, e.g., 4, 5, 6, 7, or 8. The fixed value may vary under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16). The changes to the VVC draft are shown in bold italic font in Table 27, with the changes in bold italic and the deleted content in italic.
Table 27
[0189] In yet another example, when the BitDepth is greater than a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), it is proposed to use only one fixed value for the Rice parameter when encoding the syntax element abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the changes are shown in bold italic font, and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 28.
Table 28
[0190] Rice Parameter Derivation in Residual Encoding According to a 26th aspect of the present disclosure, it is proposed to use a variable method of Rice parameter derivation to encode certain syntax elements in residual encoding, such as abs_remainder / dec_abs_level, and according to certain encoded information of the current block, such as quantization parameters or encoded bit depths associated with TB / CB and / or slice / profile, and / or according to a new flag associated with TB / CB / slice / picture / sequence level, such as extended_precision_processing_flag, the selection result is determined. Various methods may be used to derive the Rice parameter, and some exemplary methods are listed below.
[0191] First, it is proposed to use a counter to derive the Rice parameter. The counter is determined according to the value of the encoded coefficient and certain encoded information of the current block, such as the component ID. A specific example is riceParameter = counter / a, where a is a positive number, such as 4, which maintains two counters (decomposed by luminance / chroma). These counters are reset to 0 at the start of each slice. When encoded, the counter is updated as follows if this is the first coefficient encoded within the sub-TU. if(coeffValue>=(3<<rice))counter++ if(((coeffValue<<1)<(1<<riceParameter))&&(counter>0))counter--
[0192] Second, it is proposed to add a shift operation in the derivation of the Rice parameter in VVC. The shift is determined according to the value of the encoded coefficient. An example of the corresponding decoding process based on the VVC draft is shown below. The shift is determined according to the counter of method 1, and the changed parts are shown in bold italic font, and the deleted content is shown in italic font. The changed parts to the VVC draft are shown in bold italic font in Table 29.
Table 29
[0193] First, it is proposed to add a shift operation in the derivation of the Rice parameter in VVC. The shift is determined according to the encoded bit depth associated with specific encoded information of the current block, such as TB / CB and / or slice profile (e.g., 14-bit profile or 16-bit profile). An example of the corresponding decoding process based on the VVC draft is shown below, where the shift is determined according to the counter of Method 1, the changed parts are shown in bold italic font, and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 30.
Table 30
[0194] Residual Encoding of Transform Skip According to the 27th aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode specific syntax elements, such as abs_remainder, in the residual encoding of transform skip, and the selection result is determined according to the quantization parameter or the encoded bit depth associated with specific encoded information of the current block, such as TB / CB and / or slice / profile, and / or according to a new flag associated with the TB / CB / slice / picture / sequence level, such as the extended_precision_processing_flag. Various methods may be used to derive the variable set of binary codewords, and some exemplary methods are listed below.
[0195] First, to determine the sign word of abs_remainder, the same procedure as that currently used in VVC is used, but a fixed Rice parameter (e.g., 2, 3, 4, 5, 6, 7, or 8) is always selected. The fixed value may vary under different conditions according to specific encoded information of the current block, such as the quantization parameter associated with the TB / CB and / or slice / profile, frame type (e.g., I, P, or B), component ID (e.g., luminance or chrominance), color format (e.g., 420, 422, or 444), or encoded bit depth, and / or according to syntax elements associated with the TB / CB / slice / picture / sequence level, such as rice_parameter_value. As shown in Table 7, a specific example is illustrated, where TH1 to TH4 are predetermined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predetermined Rice parameters. It is worth noting that the same logic can actually be implemented in different ways. For example, a specific equation or look-up table may also be used to derive the same Rice parameter from the BitDepth value of the current CU / sequence as shown in Table 7.
[0196] Second, it is fixed-length binarization.
[0197] Third, it is truncated Rice binarization.
[0198] Fourth, it is a truncated binary (TB) binarization process.
[0199] Fifth, it is a k-th exponential Golomb binarization process (EGk).
[0200] Sixth, it is limited k-th exponential Golomb binarization.
[0201] An example of the corresponding decoding process based on the VVC draft is shown below. The changes to the VVC draft are shown in bold italic font in Table 31, and the deleted content is shown in italic font. It is worth noting that in practice, the same logic can be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter.
Table 31
[0202] In another example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, it is proposed to use only one fixed value for the Rice parameter when encoding the abs_remainder syntax element. The corresponding decoding process based on the VVC draft is shown below. The changes are shown in bold italic font, and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 32.
Table 32
[0203] In yet another example, when a new flag, e.g., extended_precision_processing_flag, is equal to 1, the Rice parameter cRiceParam is fixed to n, where n is a positive number (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value may vary under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below. The changes are shown in bold italic font, and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 33.
Table 33
[0204] In yet another example, when the BitDepth is greater than or equal to a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is fixed to n, where n is a positive number, e.g., 4, 5, 6, 7, or 8. The fixed values may vary under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the changes are shown in bold italic font, and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 34.
Table 34
[0205] In yet another example, one control flag is signaled within the slice header to indicate whether signaling of the Rice parameters for the conversion skip block is valid or not. When the control flag is signaled as valid, one more syntax element is further signaled for each conversion skip slice to indicate the Rice parameter of that slice. When the control flag is signaled as invalid (e.g., when set to be equal to "0"), no further syntax elements are signaled at a lower level to indicate the Rice parameter of the conversion skip slice, and a default Rice parameter (e.g., 1) is used for all conversion skip slices. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2), the changes are shown in bold italic font, and the deleted content is shown in italic font. The changes to the VVC draft are shown in bold italic font in Table 35. It is worth noting that sh_ts_residual_coding_rice_index can be coded in various ways and / or can have a maximum value. For example, for encoding / decoding the same syntax element, u(n), an unsigned integer using n bits, or f(n), a fixed pattern bit sequence using n bits with the left bit described first (from left to right), may also be used.
[0206] Slice header syntax [Table 35]
[0207] That sh_ts_residual_coding_rice_flag is equal to 1 specifies that sh_ts_residual_coding_rice_index may exist within the current slice. That sh_ts_residual_coding_rice_flag is equal to 0 specifies that sh_ts_residual_coding_rice_index does not exist within the current slice. When sh_ts_residual_coding_rice_flag does not exist, the value of sh_ts_residual_coding_rice_flag is inferred to be equal to 0. sh_ts_residual_coding_rice_index specifies the Rice parameter used in the residual_ts_coding() syntax structure.
Table 36
[0208] In yet another example, one control flag is signaled within the sequence parameter set (or within the sequence parameter set range extension syntax) to indicate whether signaling of the Rice parameters for the conversion skip block is effective or not. When the control flag is signaled as effective, one syntax element is further signaled for each conversion skip slice to indicate the Rice parameter of that slice. When the control flag is signaled as ineffective (e.g., when set to be equal to "0"), no further syntax elements are signaled at a lower level to indicate the Rice parameters of the conversion skip slices, and a default Rice parameter (e.g., 1) is used for all conversion skip slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes to the VVC draft are shown in bold italics in Table 37, and the deleted content is shown in italics. It is worth noting that sh_ts_residual_coding_rice_idx can be coded in various ways and / or can have a maximum value. For example, for encoding / decoding the same syntax element, u(n), an unsigned integer using n bits, or f(n), a fixed pattern bit sequence using n bits with the left bit described first (from left to right), may also be used.
[0209] Sequence parameter set RBSP syntax [Table 37]
[0210] That sps_ts_residual_coding_rice_present_in_sh_flag is equal to 1 specifies that sh_ts_residual_coding_rice_idx may exist in the SH syntax structure referring to the SPS. That sps_ts_residual_coding_rice_present_in_sh_flag is equal to 0 specifies that sh_ts_residual_coding_rice_idx does not exist in the SH syntax structure referring to the SPS. When sps_ts_residual_coding_rice_present_in_sh_flag does not exist, the value of sps_ts_residual_coding_rice_present_in_sh_flag is inferred to be equal to 0.
[0211] Slice header syntax
Table 38
[0212] sh_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure.
Table 39
[0213] In yet another example, one syntax element is signaled for each conversion skip slice to indicate the Rice parameter of that slice. An example of the corresponding decoding process based on the VVC draft is shown below. The changes to the VVC draft are indicated in Table 40 in bold italic font. It is worth noting that sh_ts_residual_coding_rice_idx can be coded in various ways and / or may have a maximum value. For example, for coding / decoding the same syntax element, u(n), an unsigned integer using n bits, or f(n), a fixed pattern bit sequence using n bits with the left bit described first (from left to right), may also be used.
[0214] Slice header syntax [Table 40]
[0215] sh_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure. If sh_ts_residual_coding_rice_idx does not exist, the value of sh_ts_residual_coding_rice_idx is inferred to be equal to 0. [Table 41]
[0216] In yet another example, one control flag is signaled within the picture parameter set range extension syntax to indicate whether signaling of the Rice parameter for the conversion skip block is enabled or disabled. When the control flag is signaled as enabled, one more syntax element is signaled to indicate the Rice parameter of that picture. When the control flag is signaled as disabled (e.g., when set to be equal to "0"), no more syntax elements are signaled at a lower level to indicate the Rice parameter of the conversion skip slice, and a default Rice parameter (e.g., 1) is used for all conversion skip slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes to the VVC draft are shown in Table 42 in bold italic font. It is worth noting that pps_ts_residual_coding_rice_idx can be coded in various ways and / or can have a maximum value. For example, for coding / decoding the same syntax element, u(n), an unsigned integer using n bits, or f(n), a fixed pattern bit sequence using n bits with the left bit described first (from left to right), may also be used.
[0217] Picture parameter set range extension syntax [Table 42]
[0218] The fact that pps_ts_residual_coding_rice_flag is equal to 1 specifies that pps_ts_residual_coding_rice_index may exist within the current picture. The fact that pps_ts_residual_coding_rice_flag is equal to 0 specifies that pps_ts_residual_coding_rice_idx does not exist within the current picture. When pps_ts_residual_coding_rice_flag does not exist, the value of pps_ts_residual_coding_rice_flag is inferred to be equal to 0.
[0219] pps_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure.
Table 43
[0220] In yet another example, it is proposed to use only variable Rice parameters for the coding of the syntax element abs_remainder. The value of the Rice parameter to be applied may be determined according to specific coded information of the current block, such as block size, quantization parameter, bit depth, transform type, etc. In a particular embodiment, it is proposed to adjust the Rice parameter based on the coding bit depth and the quantization parameter applied to one CU. The corresponding decoding process based on the VVC draft is shown below, and the changes to the VVC draft are shown in bold italic font in Table 44, and the deleted content is shown in italic font. It is worth noting that in practice the same logic may be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter.
Table 44
[0221] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 33 or 34). The changes to the VVC draft are shown in bold italic font in Table 45, and the deleted content is shown in italics. It is worth noting that the same logic can actually be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter.
Table 45
[0222] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH A and TH B are predetermined thresholds (e.g., TH A = 8, TH B = 33 or 34). The changes to the VVC draft are shown in bold italics in Table 46, and the deleted content is shown in italics. It is worth noting that the same logic can actually be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter.
Table 46
[0223] In yet another example, it is proposed that when a new flag, for example, extended_precision_processing_flag is equal to 1, only a variable Rice parameter is used for encoding the syntax element of abs_remainder. The variable value may be determined according to specific encoded information of the current block, such as block size, quantization parameter, bit depth, transform type, etc. In a specific embodiment, it is proposed to adjust the Rice parameter based on the encoded bit depth and the quantization parameter applied to one CU. The corresponding decoding process based on the VVC draft is shown below. The changes to the VVC draft are shown in bold italic font in Table 47. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or look-up table may also be used to derive the same Rice parameter.
Table 47
[0224] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 18, 19). The changes to the VVC draft are shown in bold italic font in Table 48. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or look-up table may also be used to derive the same Rice parameter.
Table 48
[0225] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH A and TH B are predetermined thresholds (e.g., TH A = 8, TH B= 18 or 19). Changes to the VVC draft are indicated in Table 49 in bold italic font. It is worth noting that the same logic can be implemented in different ways in practice. For example, specific equations or look-up tables may also be used to derive the same Rice parameter.
Table 49
[0226] Figure 16 shows a method for video encoding. The method may be applied, for example, to an encoder. In step 1610, the encoder may receive a video input. The video input may be, for example, a live stream. In step 1612, the encoder may obtain a quantization parameter based on the video input. The quantization parameter may be calculated, for example, by a quantization unit in the encoder. In step 1614, the encoder may derive a Rice parameter based on at least one predetermined threshold, an encoded bit depth, and the quantization parameter. The Rice parameter is used, for example, to signal the syntax of abs_remainder and dec_abs_level. In step 1616, the encoder may entropy encode the video bitstream based on the Rice parameter. The video bitstream may be entropy encoded, for example, to generate a compressed video bitstream.
[0227] In yet another example, when BitDepth is greater than 10, it is proposed to use only fixed values (e.g., 2, 3, 4, 5, 6, 7, or 8) for the Rice parameter when encoding the abs_remainder syntax element. The fixed values may vary under different conditions according to specific encoded information of the current block, e.g., quantization parameters. The corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 18, 19). The changes to the VVC draft are shown in bold italic font in Table 50. It is worth noting that the same logic may actually be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter. [Table 50]
[0228] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH A and TH B are predetermined thresholds (e.g., TH A = 8, TH B = 18 or 19). The changes to the VVC draft are shown in bold italic font in Table 51. It is worth noting that the same logic may actually be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter. [Table 51]
[0229] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 33 or 34). The changes to the VVC draft are shown in bold italic font in Table 52. It is worth noting that the same logic may actually be implemented in different ways. For example, specific equations or look-up tables may also be used to derive the same Rice parameter.
Table 52
[0230] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH A and TH B are predetermined thresholds (e.g., TH A = 8, TH B = 33 or 34). The changes to the VVC draft are shown in bold italic font in Table 53. It is worth noting that the same logic can actually be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter.
Table 53
[0231] It is worth mentioning that in the above illustration, the equation used to calculate a specific Rice parameter is only used as an example to illustrate the proposed concept. For those skilled in the art of modern video coding techniques, other mapping functions (or equivalent mapping equations) are already applicable to the proposed concept (i.e., determining the Rice parameter of the transform skip mode based on the coded bits and the applied quantization parameter). On the other hand, it should also be mentioned that in the current VVC design, it is allowed that the value of the applied quantization parameter changes at the coded block group level. Therefore, the proposed Rice parameter adjustment method can provide flexible adaptation of the Rice parameter of the transform skip mode at the coded block group level.
[0232] Transmission information of normal residual coding and transform skip residual coding According to a 28th aspect of the present disclosure, signaling a Rice parameter of a binary codeword for encoding a specific syntax element, e.g., abs_remainder, in transform skip residual coding, a shift parameter and an offset parameter for deriving the Rice parameter used for abs_remainder / dec_abs_level in normal residual coding, and further, determining whether to signal, according to a quantization parameter or a coded bit depth associated with a specific coded information of a current block, e.g., TB / CB and / or slice / profile, and / or according to a new flag associated with TB / CB / slice / picture / sequence level, e.g., sps_residual_coding_info_present_in_sh_flag, is proposed.
[0233] In one example, one control flag is signaled within the slice header to indicate whether signaling of the Rice parameter for the conversion skip block and / or signaling of the shift parameter and / or offset parameter for derivation of the Rice parameter in the conversion block is effective or not. When the control flag is signaled as effective, one syntax element is further signaled for each conversion skip slice to indicate the Rice parameter of that slice, and two syntax elements are further signaled for each conversion slice to indicate the shift parameter and / or offset parameter for derivation of the Rice parameter of that slice. When the control flag is signaled as ineffective (e.g., when set to be equal to "0"), no further syntax elements are signaled at a lower level to indicate the Rice parameter of the conversion skip slice, a default Rice parameter (e.g., 1) is used for all conversion skip slices, no further syntax elements are signaled at a lower level to indicate the shift parameter and offset parameter for derivation of the Rice parameter of the conversion slice, and default shift parameter and / or offset parameter (e.g., 0) are used for all conversion slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes to the VVC draft are shown in Table 54 in bold italic font. It is worth noting that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_index can be coded in various ways and / or can have maximum values. For example, for encoding / decoding the same syntax element, a unsigned integer u(n) using n bits, or a fixed pattern bit sequence f(n) using n bits with the left bit described first (from left to right) may also be used.
[0234] Figure 17 shows a method for video decoding. The method may be applied to, for example, an encoder. In step 1710, the encoder may receive a video input. In step 1712, the encoder may signal a Rice parameter of a binary codeword for an encoded syntax element. The encoded syntax element may include an abs_remainder in transform skip residual coding. In step 1714, the encoder may entropy encode a video bitstream based on the Rice parameter and the video input.
[0235] Slice header syntax [Table 54]
[0236] That sh_residual_coding_rice_flag is equal to 1 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_residual_coding_rice_index may exist within the current slice. That sh_residual_coding_rice_flag is equal to 0 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_residual_coding_rice_index do not exist within the current slice.
[0237] sh_residual_coding_rice_shift specifies a shift parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_shift does not exist, the value of sh_residual_coding_rice_shift is inferred to be equal to 0.
[0238] sh_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_offset does not exist, the value of sh_residual_coding_rice_offset is inferred to be equal to 0.
[0239] sh_ts_residual_coding_rice_index specifies the Rice parameter used in the residual_ts_coding() syntax structure. When sh_ts_residual_coding_rice_index does not exist, the value of sh_ts_residual_coding_rice_index is inferred to be equal to 0.
Table 55
Table 56
[0240] In another example, one control flag is signaled within the sequence parameter set (or within the sequence parameter set range extension syntax) to indicate whether signaling of the Rice parameters for the conversion skip block and / or signaling of the shift parameter and / or offset parameter for derivation of the Rice parameters in the conversion block is effective or not. When the control flag is signaled as effective, one syntax element is further signaled for each conversion skip slice to indicate the Rice parameter of that slice, and two syntax elements are further signaled for each conversion slice to indicate the shift parameter and / or offset parameter for derivation of the Rice parameter of that slice. When the control flag is signaled as ineffective (e.g., when set to be equal to "0"), no further syntax elements are signaled at a lower level to indicate the Rice parameter of the conversion skip slice, a default Rice parameter (e.g., 1) is used for all conversion skip slices, no further syntax elements are signaled at a lower level to indicate the shift parameter and / or offset parameter for derivation of the Rice parameter of the conversion slice, and default shift and / or offset parameters (e.g., 0) are used for all conversion slices. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). The changes to the VVC draft are shown in bold italic font in Table 57. It is worth noting that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx can be coded in various ways and / or may have maximum values. For example, for coding / decoding the same syntax element, u(n), an unsigned integer using n bits, or f(n), a fixed pattern bit sequence using n bits with the left bit described first (from left to right), may also be used.
[0241] Sequence parameter set RBSP syntax
Table 57
[0242] The fact that sps_residual_coding_info_present_in_sh_flag is equal to 1 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx may exist in the SH syntax structure referring to the SPS. The fact that sps_residual_coding_info_present_in_sh_flag is equal to 0 specifies that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx do not exist in the SH syntax structure referring to the SPS. When sps_residual_coding_info_present_in_sh_flag does not exist, the value of sps_residual_coding_info_present_in_sh_flag is inferred to be equal to 0.
[0243] Slice header syntax
Table 58
[0244] sh_residual_coding_rice_shift specifies the shift parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_shift does not exist, the value of sh_residual_coding_rice_shift is inferred to be equal to 0.
[0245] sh_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_offset does not exist, the value of sh_residual_coding_rice_offset is inferred to be equal to 0.
[0246] sh_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure. When sh_ts_residual_coding_rice_index does not exist, the value of sh_ts_residual_coding_rice_index is inferred to be equal to 0.
Table 59
Table 60
[0247] In yet another example, one syntax element is signaled to indicate the Rice parameter of each transform skip slice, and two syntax elements are signaled to indicate the shift parameter and / or offset parameter for deriving the Rice parameter of each transform skip slice. An example of the corresponding decoding process based on the VVC draft is shown below. The changes to the VVC draft are indicated in Table 61 in bold italic font. It is worth noting that sh_residual_coding_rice_shift, sh_residual_coding_rice_offset, and sh_ts_residual_coding_rice_idx can be coded in various ways and / or may have maximum values. For example, for coding / decoding the same syntax element, a unsigned integer u(n) using n bits, or a fixed pattern bit sequence f(n) using n bits with the left bit described first (from left to right) may also be used.
[0248] Slice header syntax [Table 61]
[0249] sh_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure. When sh_ts_residual_coding_rice_idx does not exist, the value of sh_ts_residual_coding_rice_idx is inferred to be equal to 0.
[0250] sh_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When sh_residual_coding_rice_offset does not exist, the value of sh_residual_coding_rice_offset is inferred to be equal to 0.
[0251] The sh_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure. When the sh_ts_residual_coding_rice_index does not exist, it is inferred that the value of the sh_ts_residual_coding_rice_index is equal to 0. [Table 62] [Table 63]
[0252] In yet another example, one control flag is signaled within the picture parameter set range extension syntax to indicate whether signaling of the Rice parameter for the conversion skip block and / or signaling of the shift parameter and / or offset parameter for derivation of the Rice parameter in the conversion block is enabled or disabled. When the control flag is signaled as enabled, one syntax element is further signaled to indicate the Rice parameter for conversion skip residual coding of that picture, and two syntax elements are further signaled to indicate the shift parameter and / or offset parameter for derivation of the Rice parameter of that picture for each normal residual coding. When the control flag is signaled as disabled (e.g., when set equal to "0"), no further syntax elements are signaled at a lower level to indicate the Rice parameter for conversion skip residual coding, a default Rice parameter (e.g., 1) is used for all conversion skip residual codings, no further syntax elements are signaled at a lower level to indicate the shift parameter and / or offset parameter for derivation of the Rice parameter of normal residual coding, and default shift parameter and / or offset parameter (e.g., 0) are used for all normal residual codings. An example of a corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined value (e.g., 0, 1, 2). Changes to the VVC draft are shown in Table 64 in bold italic font. pps_residual_coding_rice_shift, pps_residual_coding_rice_offset, and pps_ts_residual_coding_rice_idx can be coded in various ways and / or may have maximum values. It is worth noting that, for example, for coding / decoding the same syntax element, a unsigned integer u(n) using n bits or a fixed pattern bit sequence f(n) using n bits with the left bit described first (from left to right) may also be used.
[0253] Picture parameter set range extension syntax
Table 64
[0254] The fact that pps_residual_coding_info_flag is equal to 1 specifies that pps_residual_coding_rice_shift, pps_residual_coding_rice_offset, and pps_ts_residual_coding_rice_index may exist within the current picture. The fact that pps_residual_coding_info_flag is equal to 0 specifies that pps_residual_coding_rice_shift, pps_residual_coding_rice_offset, and pps_ts_residual_coding_rice_idx do not exist within the current picture. When pps_residual_coding_info_flag does not exist, the value of pps_residual_coding_info_flag is inferred to be equal to 0.
[0255] pps_residual_coding_rice_shift specifies the shift parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When pps_residual_coding_rice_shift does not exist, the value of pps_residual_coding_rice_shift is inferred to be equal to 0.
[0256] pps_residual_coding_rice_offset specifies the offset parameter used in the Rice parameter derivation process for abs_remainder[] and dec_abs_level[]. When pps_residual_coding_rice_offset does not exist, the value of pps_residual_coding_rice_offset is inferred to be equal to 0.
[0257] The pps_ts_residual_coding_rice_idx specifies the Rice parameter used in the residual_ts_coding() syntax structure. When the pps_ts_residual_coding_rice_index does not exist, the value of the pps_ts_residual_coding_rice_index is inferred to be equal to 0. [Table 65] [Table 66]
[0258] According to the 29th aspect of the present disclosure, in transform skip residual coding, different Rice parameters for encoding specific syntax elements, such as abs_remainder, and shift and offset parameters for deriving the Rice parameters used for abs_remainder / dec_abs_level in normal residual coding are used, and which one to use is determined according to the quantization parameter or encoded bit depth associated with the specific encoded information of the current block, such as TB / CB and / or slice / profile, and / or according to a new flag associated with the TB / CB / slice / picture / sequence level, such as the sps_residual_coding_info_present_in_sh_flag.
[0259] In one example, one control flag is signaled within the slice header to indicate whether the process of deriving the Rice parameter of the conversion skip block and the process of deriving the shift parameter and / or offset parameter of the Rice parameter in the conversion block are valid or invalid. When the control flag is signaled as valid, the Rice parameter may vary under different conditions according to the specific encoded information of the current block, such as the quantization parameter and the bit depth. Also, the shift parameter and / or offset parameter for deriving the Rice parameter in normal residual encoding may vary under different conditions according to the specific encoded information of the current block, such as the quantization parameter and the bit depth. When the control flag is signaled as invalid (e.g., when set to be equal to "0"), a default Rice parameter (e.g., 1) is used for all conversion skip slices, and a default shift parameter and / or offset parameter (e.g., 0) is used for all conversion slices. An example of the corresponding decoding process based on the VVC draft is shown below, TH A and TH B are predetermined thresholds (e.g., TH A = 8, TH B = 18 or 19). The changes to the VVC draft are shown in bold italic font in Table 67. It is worth noting that the same logic can actually be implemented in different ways. For example, a specific equation or a look-up table may also be used to derive the same Rice parameter.
[0260] Slice header syntax
Table 67
[0261] That sh_residual_coding_rice_flag is equal to 1 specifies that the bit-depth-dependent Rice parameter derivation process is used for the current slice. That sh_residual_coding_rice_flag is equal to 0 specifies that the bit-depth-dependent Rice parameter derivation process is not used for the current slice.
Table 68
Table 69
[0262] In yet another example, the corresponding decoding process based on the VVC draft is shown below, where TH is a predetermined threshold (e.g., 18, 19). The changes to the VVC draft are shown in bold italic font in Table 70. It is worth noting that in practice the same logic may be implemented in different ways. For example, a specific equation or lookup table may also be used to derive the same Rice parameter.
Table 70
[0263] According to another aspect of the present disclosure, it is proposed to add a constraint that the values of these coding tools flag in order to provide the same general constraint control as other ones within the general constraint information.
[0264] For example, sps_ts_residual_coding_rice_present_in_sh_flag being equal to 1 specifies that sh_ts_residual_coding_rice_idx may exist in the SH syntax structure that refers to the SPS. When sps_ts_residual_coding_rice_present_in_sh_flag is equal to 0, it specifies that sh_ts_residual_coding_rice_idx does not exist in the SH syntax structure that refers to the SPS. According to the present disclosure, in order to provide the same general constraint control as other flags, it is proposed to add a syntax element gci_no_ts_residual_coding_rice_constraint_flag to the general constraint information syntax. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are emphasized in italics.
Table 71
Table 72
[0265] In another example, pps_ts_residual_coding_rice_flag being equal to 1 specifies that pps_ts_residual_coding_rice_index may exist within the current picture. pps_ts_residual_coding_rice_flag being equal to 0 specifies that pps_ts_residual_coding_rice_idx does not exist within the current picture. According to the present disclosure, in order to provide the same general constraint control as other flags, it is proposed to add a syntax element gci_no_ts_residual_coding_rice_constraint_flag to the general constraint information syntax. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are emphasized in italics.
Table 73
Table 74
[0266] In yet another example, when sps_rice_adaptation_enabled_flag is equal to 1, it indicates that the Rice parameters for the binarization of abs_remaining[] and dec_abs_level can be derived by an equation.
[0267] The equation may include RiceParam = RiceParam + shiftVal, and shiftVal = (localSumAbs < Tx[0])? Rx[0] : ((localSumAbs < Tx[1])? Rx[1] : ((localSumAbs < Tx[2])? Rx[2] : ((localSumAbs < Tx[3])? Rx[3] : Rx[4]))), where the lists Tx[] and Rx[] are specified as Tx[] = {32, 128, 512, 2048}>>(1523) Rx[] = {0, 2, 4, 6, 8}.
[0268] According to the present disclosure, in order to provide the same general constraint control as other flags, it is proposed to add a syntax element gci_no_rice_adaptation_constraint_flag to the general constraint information syntax. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are highlighted in italics.
Table 75
Table 76
[0269] Since the proposed Rice parameter adaptation method is only used for transform skip residual coding (TSRC), the proposed method can bring effects when TSRC is effective. Correspondingly, in one or more embodiments of the present disclosure, when the transform skip mode is invalid from the general constraint information level, for example, when the value of gci_no_transform_skip_constraint_flag is set to 1, it is proposed to add a one-bit stream constraint that requires the value of gci_no_rice_adaptation_constraint_flag to be 1.
[0270] In yet another example, sps_range_extension_flag being equal to 1 specifies that the sps_range_extension() syntax structure exists within the SPS RBSP syntax structure. sps_range_extension_flag being equal to 0 specifies that this syntax structure does not exist. According to the present disclosure, it is proposed to add a syntax element gci_no_range_extension_constraint_flag to the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are highlighted in italics.
Table 77
Table 78
[0271] Figure 19 shows a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 1902, the decoder may receive a sequence parameter set (SPS) range extension flag indicating whether a syntax structure sps_range_extension exists within the raw byte sequence payload (RBSP) syntax structure of a slice head (SH) based on the value of the SPS range extension flag.
[0272] In step 1904, the decoder may determine that sps_range_extension exists within the SH RBSP syntax structure in response to a determination that the value of the SPS range extension flag is equal to 1.
[0273] In step 1906, the decoder may determine that sps_range_extension does not exist within the SH RBSP syntax structure in response to a determination that the value of the range extension flag is equal to 0.
[0274] The sps_cabac_bypass_alignment_enabled_flag specifies that the value of ivlCurrRange can be aligned before the bypass decoding of the syntax elements sb_coded_flag[][], abs_remainder[], dec_abs_level[n], and coeff_sign_flag[]. That the sps_cabac_bypass_alignment_enabled_flag is equal to 0 specifies that the value of ivlCurrRange is not aligned before the bypass decoding. According to the present disclosure, it is proposed to add a syntax element gci_no_cabac_bypass_alignment_constraint_flag to the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are emphasized in italics.
Table 79
Table 80
[0275] Figure 20 shows a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 2002, the decoder may receive an SPS alignment valid flag indicating whether the index ivlCurrRange is aligned before bypass decoding based on the value of the syntax element sb_coded_flag, abs_remainder, dec_abs_level, and coeff_sign_flagn that is valid for SPS alignment.
[0276] In step 2004, the decoder may determine that ivlCurrRange is aligned before bypass decoding in response to a determination that the value of the SPS alignment valid flag is equal to 1.
[0277] In step 2006, the decoder may determine that ivlCurrRange is not aligned before bypass decoding in response to a determination that the value of the SPS alignment valid flag is equal to 0.
[0278] In yet another example, that the extended_precision_processing_flag is equal to 1 specifies that an extended dynamic range can be used for the transform coefficients and transform processing. That the extended_precision_processing_flag is equal to 0 specifies that the extended dynamic range is not used. According to the present disclosure, it is proposed to add a syntax element gci_no_extended_precision_processing_constraint_flag to the general constraint information syntax to provide the same general constraint control as other flags. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are highlighted in italics.
Table 81
Table 82
[0279] Figure 21 shows a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 2102, the decoder may include receiving a conversion coefficient and an extended precision processing flag indicating whether an extended dynamic range is adopted during the conversion process based on the value of the extended precision processing flag.
[0280] In step 2104, the decoder may determine that an extended dynamic range is adopted for the conversion coefficient and during the conversion process in response to a determination that the value of the extended precision processing flag is equal to 1.
[0281] In step 2106, the decoder may determine that an extended dynamic range is not adopted for the conversion coefficient or during the conversion process in response to a determination that the value of the extended precision processing flag is equal to 0.
[0282] In yet another example, the fact that the persistant_rice_adaptation_enabled_flag is equal to 1 specifies that at the start of each sub-block, the derivation of the Rice parameters for the binarization of abs_remaining[] and dec_abs_level can be initialized using the mode-dependent statistics accumulated from the previous sub-block. The fact that the persistant_rice_adaptation_enabled_flag is equal to 0 specifies that the previous sub-block state is not used in the derivation of the Rice parameters. According to the present disclosure, in order to provide the same general constraint control as for other flags, it is proposed to add a syntax element gci_no_persistent_rice_adaptation_constraint_flag within the general constraint information syntax. An example of the decoding process of the VVC draft is shown below. The changes to the VVC draft are highlighted. The added parts are highlighted in italics.
Table 83
Table 84
[0283] FIG. 22 shows a method for video encoding according to an example of the present disclosure. The method may be applied to, for example, a decoder. In step 2202, the decoder may receive a persistent Rice adaptation enable flag indicating whether, based on the value of the persistent Rice adaptation enable flag, the derivation of the Rice parameters for the binarization of abs_remaining and dec_abs_level is initialized at the start of each sub-block that employs the mode-dependent statistics accumulated from the previous sub-block.
[0284] In step 2204, the decoder may determine that the derivation of the Rice parameters for binarization is initialized at the start of each sub-block that employs the mode-dependent statistics accumulated from the previous sub-block in response to a determination that the value of the persistent Rice adaptation enable flag is equal to 1.
[0285] In step 2206, in response to determining that the value of the permanent Rice adaptation enable flag is equal to 0, the decoder may determine that the previous sub-block state is not adopted in Rice parameter derivation.
[0286] The above method may be implemented using one or more circuits including one or more of application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus may use the circuits in combination with other hardware components or software components to perform the methods described above. Each module, sub-module, unit, or sub-unit disclosed above may be at least partially implemented using one or more circuits.
[0287] FIG. 18 shows a computing environment 1810 coupled to a user interface 1860. The computing environment 1810 may be part of a data processing server. The computing environment 1810 includes a processor 1820, a memory 1840, and an I / O interface 1850.
[0288] The processor 1820 typically controls the overall operation of the computing environment 1810, such as operations related to display, data acquisition, data communication, and image processing. The processor 1820 may include one or more processors for executing instructions to perform all or some of the steps in the above methods. Further, the processor 1820 may include one or more modules to facilitate interaction between the processor 1820 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.
[0289] Memory 1840 is configured to store various types of data to support the operation of computing environment 1810. Memory 1840 may include a predetermined software 1842. Examples of such data include instructions for any application or method operating on computing environment 1810, video data sets, image data, and the like. Memory 1840 may be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0290] I / O interface 1850 provides an interface between processor 1820 and peripheral interface modules such as a keyboard, click wheel, buttons, and the like. The buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. I / O interface 1850 may be coupled to an encoder and a decoder.
[0291] In some embodiments, there is also provided a non-transitory computer-readable storage medium including a plurality of programs such as those included in memory 1840 and executable by processor 1820 within computing environment 1810 for performing the above-described method. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.
[0292] A non-transitory computer-readable storage medium stores a plurality of programs executable by a computing device having one or more processors therein, and the plurality of programs, when executed by the one or more processors, cause the computing device to execute the above method for motion prediction.
[0293] In some embodiments, the computing environment 1810 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for executing the above method.
[0294] FIG. 23 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 23, the system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. The source device 12 and the destination device 14 may include any of a variety of electronic devices including a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, or the like. In some implementations, the source device 12 and the destination device 14 have wireless communication capabilities.
[0295] In some implementations, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or communication device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless communication medium or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be useful in facilitating communication from the source device 12 to the destination device 14.
[0296] In some other implementations, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of data storage media that are either distributed or locally accessible, such as a hard drive, a Blu-ray disk, a digital versatile disk (DVD), a compact disk read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or download. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection that includes a combination of both a wireless channel (e.g., a wireless fidelity (Wi-Fi) connection) and a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.) suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0297] As shown in FIG. 23, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include a video capture device, such as a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a source such as a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the implementations described in this application may generally be applicable to video encoding and may be applicable to wireless applications and / or wired applications.
[0298] Captured video, previously captured video, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored on the storage device 32 so that it can be accessed later by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.
[0299] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and may receive encoded video data via link 16. The encoded video data communicated on link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use in decoding the video data by video decoder 30. Such syntax elements may be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0300] In some implementations, the destination device 14 may include a display device 34, which may be an integrated display device and / or an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to the user and may include any of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type.
[0301] The video encoder 20 and the video decoder 30 may operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and is applicable to other video encoding / decoding standards as well. In general, it is contemplated that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, in general, it is also contemplated that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of these current or future standards.
[0302] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the electronic device is implemented partially within software, instructions for the software are stored on a suitable non-transitory computer-readable medium and executed within the hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and either of them may be integrated as part of a combined encoder / decoder (CODEC) within their respective devices.
[0303] FIG. 24 is a block diagram showing an exemplary video encoder 20 according to some implementations described in this application. The video encoder 20 may perform intra prediction encoding and inter prediction encoding of video blocks within a video frame. Intra prediction encoding depends on spatial prediction to reduce or remove the spatial redundancy of video data within a given video frame or picture. Inter prediction encoding depends on temporal prediction to reduce or remove the temporal redundancy of video data within adjacent video frames or pictures of a video sequence. It should be noted that the term "frame" may be used synonymously with the terms "image" or "picture" in the field of video encoding.
[0304] As shown in FIG. 24, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63 such as a deblocking filter may be disposed between the adder 62 and the DPB 64 to filter block boundaries and remove blocky artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter such as a sample adaptive offset (SAO) filter and / or an adaptive loop filter (ALF) may also be used to filter the output of the adder 62. In some examples, the loop filter may be omitted, and the decoded video blocks may be directly provided to the DPB 64 by the adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.
[0305] The video data memory 40 may store video data to be encoded by components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18 shown in FIG. 23. The DPB 64 is a buffer that stores reference video data (e.g., a reference frame or a reference picture) for use when encoding video data by the video encoder 20 (e.g., in an intra prediction encoding mode or an inter prediction encoding mode). The video data memory 40 and the DPB 64 may be formed by any of various memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20 or may be off-chip with respect to those components.
[0306] As shown in FIG. 24, after receiving video data, the segmentation unit 45 in the prediction processing unit 41 segments the video data into video blocks. This segmentation also includes segmenting the video frame into slices, tiles (e.g., a set of video blocks), or other larger coding units (CUs) according to a predefined decomposition structure such as a quad-tree (QT) structure associated with the video data. The video frame may be or may be regarded as a two-dimensional array or two-dimensional matrix of samples having sample values. Samples in the array may also be called pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be divided into a plurality of video blocks, for example, by using QT segmentation. A video block, although also having a smaller dimension than the video frame, may be or may be regarded as a two-dimensional array or two-dimensional matrix of samples having sample values. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. The video block may be further segmented into one or more block segments or sub-blocks (which may form blocks again) by repeatedly using, for example, QT segmentation, binary-tree (BT) segmentation, or triple-tree (TT) segmentation, or any combination thereof. It should be noted that the term "block" or "video block" as used herein can be a part of a frame or picture, specifically a rectangular (square or non-square) portion. For example, with respect to HEVC and VVC, a block or video block may be or may correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or a corresponding block, e.g., a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB), and / or a sub-block, or may correspond to them.
[0307] Based on the error results (e.g., the coding rate and the level of distortion), the prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes. The prediction processing unit 41 provides the resulting intra prediction coded block or inter prediction coded block to the adder 50 to generate a residual block, and also provides the resulting intra prediction coded block or inter prediction coded block to the adder 62 to reconstruct the coded block for use as part of the reference frame later. Also, the prediction processing unit 41 provides syntax elements such as motion vectors, intra mode indicators, partition information, and other such syntax information to the entropy coding unit 56.
[0308] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 within the prediction processing unit 41 may perform intra prediction coding of the current video block with respect to one or more adjacent blocks within the same frame as the current block to be coded, providing spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter prediction coding of the current video block with respect to one or more prediction blocks within one or more reference frames, providing temporal prediction. The video encoder 20 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of video data.
[0309] In some embodiments, the motion estimation unit 42 determines an inter prediction mode for the current video frame by generating a motion vector indicating the displacement of a video block in the current video frame relative to a prediction block in a reference video frame according to a predetermined pattern within a sequence of video frames. Motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. For example, the motion vector may indicate the displacement of a video block in the current video frame or picture relative to a prediction block in a reference frame with respect to the current block being encoded in the current frame. The predetermined pattern designates video frames within the sequence as P-frames or B-frames. The intra BC unit 48 may determine a vector for intra BC encoding, such as a block vector, in a manner similar to the determination of the motion vector by the motion estimation unit 42 for inter prediction, or may utilize the motion estimation unit 42 to determine the block vector.
[0310] The prediction block of the video block may be a block or reference block in the reference frame that is considered to exactly match the video block to be encoded in terms of pixel difference, or may correspond to those blocks, and those blocks may be determined by sum of absolute differences (SAD), sum of square differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 may calculate values of sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 may interpolate values of quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation unit 42 may perform motion search for both integer pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0311] The motion estimation unit 42 calculates the motion vector of a video block in an inter-predicted coded frame by comparing the position of the video block with the position of the predicted block in a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1) that respectively identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44 and then to the entropy coding unit 56.
[0312] Motion compensation performed by the motion compensation unit 44 may include fetching or generating a predicted block based on the motion vector determined by the motion estimation unit 42. When receiving the motion vector of the current video block, the motion compensation unit 44 may identify the position of the predicted block pointed to by the motion vector in one of the reference frame lists, obtain the predicted block from the DPB 64, and transfer the predicted block to the adder 50. Next, the adder 50 forms a residual video block of pixel difference values by subtracting the pixel values of the predicted block provided by the motion compensation unit 44 from the pixel values of the currently encoded video block. The pixel difference values forming the residual video block may include a luminance difference component, a chrominance difference component, or both. The motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use when decoding the video blocks of the video frame by the video decoder 30. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predicted block, any flag indicating the prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for the purpose of illustration.
[0313] In some implementations, the intra BC unit 48 can generate vectors and fetch prediction blocks in a manner similar to the manner described above in relation to the motion estimation unit 42 and the motion compensation unit 44, where the prediction blocks are in the same frame as the current block being coded, and the vectors are called block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used for coding the current block. In some examples, the intra BC unit 48 may code the current block using various intra prediction modes, for example during a separate coding pass, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use from among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes and select, as the appropriate intra prediction mode to use, the intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was coded to create the coded block, and the bit rate (i.e., the number of bits) used to create the coded block. The intra BC unit 48 may calculate a ratio from the distortion and rate of the various coded blocks to determine which intra prediction mode shows the best rate-distortion value for that block.
[0314] In other examples, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra BC prediction according to the implementations described herein. In any case, for intra block copy, the predicted block may be a block that is considered to exactly match the block to be encoded, which can be determined by SAD, SSD, or other difference metrics with respect to the pixel difference, and the identification of the predicted block may include the calculation of the values of sub-pixel positions.
[0315] Whether the predicted block is a block from the same frame according to intra prediction or a block from different frames according to inter prediction, the video encoder 20 may form a residual video block by subtracting the pixel values of the predicted block from the pixel values of the currently encoded video block to form a pixel difference value. The pixel difference value for forming the residual video block may include both the luminance component difference and the chrominance component difference.
[0316] As described above, the intra prediction processing unit 46 may perform intra prediction on the current video block as an alternative to inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44, or as an alternative to intra block copy prediction executed by the intra BC unit 48. Specifically, the intra prediction processing unit 46 may determine an intra prediction mode to be used for encoding the current block. To do so, the intra prediction processing unit 46 may, for example, during a separate encoding pass, encode the current block using various intra prediction modes, and the intra prediction processing unit 46 (or, in some examples, the mode selection unit) may select an appropriate intra prediction mode to use from the tested intra prediction modes. The intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode into the bitstream.
[0317] After the prediction processing unit 41 determines a predicted block of the current video block by inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be included in one or more TUs and is provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0318] The conversion processing unit 52 may send the resulting conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0319] Following quantization, the entropy coding unit 56 entropy-codes the quantized conversion coefficients into the video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding methodology or technique. The encoded bitstream is then transmitted to the video decoder 30 as shown in FIG. 23, or may be archived in the storage device 32 as shown in FIG. 23 for later transmission to or retrieval by the video decoder 30. The entropy coding unit 56 may also entropy-code the motion vectors and other syntax elements of the currently encoded video frame.
[0320] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain in order to generate a reference block for prediction of other video blocks. As described above, the motion compensation unit 44 may generate a motion-compensated prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-pixel values for use in motion estimation.
[0321] The adder 62 may add the reconstructed residual block to the motion-compensated prediction block created by the motion compensation unit 44 to create a reference block for storage in the DPB 64. The reference block may then be used as a prediction block for inter prediction of another video block within a subsequent video frame by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44.
[0322] FIG. 25 is a block diagram showing an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is substantially inverse to the encoding process described above with respect to the video encoder 20 in relation to FIG. 24. For example, the motion compensation unit 82 may generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.
[0323] In some examples, tasks for implementing the embodiments of the present application may be assigned to units of the video decoder 30. Also, in some examples, the implementation forms of the present disclosure may be divided into one or more of the units of the video decoder 30. For example, the intra BC unit 85 may execute the implementation forms of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be executed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0324] The video data memory 79 may store video data such as an encoded video bit stream that is decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source such as a camera, via wired network communication or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bit stream. The DPB 92 of the video decoder 30 stores reference video data for use when decoding video data by the video decoder 30 (e.g., in an intra prediction coding mode or an inter prediction coding mode). The video data memory 79 and the DPB 92 may be formed by any of various memory devices such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. In FIG. 25, for illustration purposes, the video data memory 79 and the DPB 92 are depicted as two separate components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip with or off-chip from other components of the video decoder 30.
[0325] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. Video decoder 30 may receive syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Then, the entropy decoding unit 80 transfers the motion vectors or intra prediction mode indicators and other syntax elements to the prediction processing unit 81.
[0326] When the video frame is encoded as an intra prediction encoded (I) frame or for an intra encoded prediction block in another type of frame, the intra prediction unit 84 of the prediction processing unit 81 may generate prediction data for the video blocks of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.
[0327] When the video frame is encoded as an inter prediction encoded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 may generate one or more prediction blocks for the video blocks of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks may be created from a reference frame in one of the reference frame lists. Video decoder 30 may construct list 0 and list 1, which are the reference frame lists, using a default construction technique based on the reference frames stored in DPB 92.
[0328] In some examples, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 creates a prediction block of the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed area of the same picture as the current video block defined by the video encoder 20.
[0329] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by analyzing the motion vector and other syntax elements, and then use the prediction information to create a prediction block of the currently decoded video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to encode a video block of the video frame, an inter prediction frame type (e.g., B or P), construction information for one or more of the reference frame lists of the frame, the motion vector of each inter prediction encoded video block of the frame, the inter prediction status of each inter prediction encoded video block of the frame, and other information for decoding a video block within the current video frame.
[0330] Similarly, the intra BC unit 85 may determine some of the received syntax elements, such as a flag, information about whether the current video block is predicted using the intra BC mode, which video blocks of the frame are within the reconstructed area, which video blocks should be stored in the DPB 92, the block vector of each intra BC prediction video block of the frame, the intra BC prediction status of each intra BC prediction video block of the frame, and other information for decoding a video block within the current video frame.
[0331] The motion compensation unit 82 may also perform interpolation using the interpolation filter used by the video encoder 20 during the encoding of video blocks, and calculate the interpolation values of the sub-integer pixels of the reference blocks. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements, and create a prediction block using the interpolation filter.
[0332] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameter calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.
[0333] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block of the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block of the current video block by adding the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. In order to further process the decoded video block, a loop filter 91 such as a deblocking filter, an SAO filter, and / or an ALF may be arranged between the adder 90 and the DPB 92. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. Then, the decoded video blocks in a given frame are stored in the DPB 92 that stores the reference frames used for subsequent motion compensation of the next video block. Also, the DPB 92, or a memory device separate from the DPB 92, may store the decoded video for later presentation on a display device such as the display device 34 of FIG. 23.
[0334] The description of the present disclosure is presented for purposes of illustration and is not intended to be exhaustive or to limit the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and associated drawings.
[0335] The examples are selected and described in order to explain the principles of the present disclosure and to enable those skilled in the art to understand the present disclosure with respect to various embodiments, and to make the most of the underlying principles and various implementations by making various modifications suitable for the particular applications contemplated. Accordingly, it is to be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed implementations, and that modifications and other implementations are intended to be included within the scope of the present disclosure.
Claims
1. A method for video encoding, comprising: receiving, by a decoder, the extended precision processing flag indicating whether an extended dynamic range is adopted for a transform coefficient in a transform process based on a value of the extended precision processing flag; The method further comprises: receiving, by the decoder, an extended precision processing constraint flag in an extended precision information syntax to provide general constraint control for at least one other flag; and in response to a determination that the value of the extended precision processing constraint flag is equal to 1, determining that the value of the other flag is equal to 0.
2. in response to a determination that the value of the extended precision processing flag is equal to 1, determining that the extended dynamic range is adopted for the transform coefficient in the transform process; and in response to a determination that the value of the extended precision processing flag is equal to 0, determining that the extended dynamic range is not adopted for the transform coefficient in the transform process. The method for video encoding according to claim 1, further comprising:
3. in response to a determination that the value of the extended precision processing constraint flag is equal to 1, determining that the value of the extended precision processing flag is equal to 0. The method for video encoding according to claim 1, further comprising:
4. An apparatus for video encoding, comprising: one or more processors; and a memory configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 1 to 3 when executing the instructions.
5. A non-transitory computer-readable storage medium for video encoding, storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to execute the method according to any one of claims 1 to 3, process a video bitstream, and store the processed video bitstream in a non-transitory computer-readable storage medium.
6. A method for transmitting a video bitstream by an image encoding apparatus, wherein the video bitstream comprises: the extended precision processing flag indicating whether an extended dynamic range is adopted for a transform coefficient in a transform process based on a value of the extended precision processing flag; and including an extended precision processing constraint flag in a general constraint information syntax to provide general constraint control for at least one other flag A method, wherein the value of the extended precision processing constraint flag being equal to 1 indicates that the value of the at least one other flag is equal to 0
Citation Information
Patent Citations
Video encoding device, method and program, and video decoding device, method and program
WO2019181354A1
Cited By
Residual and coefficient coding for video coding
JP2025131817A