Adaptive Color Transformation in Image / Video Coding and Decoding

By introducing adaptive color transformation and weighted prediction in video encoding and decoding technology, combined with chroma residuals, the problem of low efficiency in combination with ACT mode and brightness BDPCM mode in the prior art is solved, and video encoding and decoding effects that are more efficient, flexible and support lossless decoding are achieved.

CN115176470BActive Publication Date: 2025-05-27DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180009720.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-18
Filing Date
2021-01-18
Publication Date
2025-05-27
Estimated Expiration
2041-01-18

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technology, the combination efficiency of ACT mode and brightness BDPCM mode is low. The quantization parameters may become negative when ACT is enabled. ACT does not support lossless encoding and decoding. Signaling notifications do not depend on block size. The flexibility of the palette mode is limited. The binarization of escaped samples does not depend on quantization parameters. The residuals in the YCgCo-R color space may exceed the dynamic range.

Method used

In the video processing method, through adaptive color transformation (ACT) and weighted prediction, combined with the joint codec (JCCR) mode of chroma residuals, quantization parameters are adjusted, lossless codec supports, dynamic block size dependence signaling notification is introduced, palette mode and escape sample processing are optimized, and the dynamic range of color space is expanded.

Benefits of technology

It improves the efficiency and flexibility of video encoding and decoding, avoids the problem of negative quantization parameters, supports lossless encoding and decoding, optimizes signaling notification and palette mode, expands the dynamic range of the color space, and improves the overall performance of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115176470B_ABST
    Figure CN115176470B_ABST
Patent Text Reader

Abstract

Methods, systems, and devices for implementing an Adaptive Color Transform (ACT) mode during image / video encoding and decoding are described. Example methods of video processing include: performing a conversion between a current video block of a video and a bitstream of the video, wherein the current video block is encoded / decoded using the ACT mode, wherein the conversion includes applying an inverse ACT transform to the current video block according to a rule, and wherein the rule specifies that a clipping operation based on the bit depth of the current video block is applied to the input of the inverse ACT transform.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] Under the provisions of the applicable patent laws and / or the Paris Convention, this application is based on International Patent Application No. PCT / CN2021 / 072396 filed on January 18, 2021, and claims the priority and benefits of International Patent Application No. PCT / CN2020 / 072900 filed on January 18, 2020 in a timely manner. For all purposes, the entire disclosure of the above - mentioned applications is incorporated by reference in its entirety as part of the disclosure of this application in accordance with the law. Technical field

[0003] This patent document relates to image and video encoding, decoding, and decoding. Background art

[0004] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the invention

[0005] This document discloses systems, methods, and devices for video encoding, decoding, and decoding, including an adaptive color transform (ACT).

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a current video block of a video and a bitstream of the video, wherein the current video block is encoded and decoded using an adaptive color transform (ACT) mode, wherein the conversion includes applying an inverse ACT transform to the current video block according to a rule, and wherein the rule stipulates that a clipping operation based on the bit - depth of the current video block is applied to the input of the inverse ACT transform.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a current video block and a bitstream of the video, wherein during the conversion, a weighted prediction of the current video block is determined using weights, and wherein the weights are included in the bitstream as information indicating the difference between the number of weights and a threshold K, where K is an integer.

[0008] In yet another example aspect, a video encoder device is disclosed. The video encoder device includes a processor configured to implement the above - mentioned method.

[0009] In yet another example aspect, a video decoder device is disclosed. The video decoder device includes a processor configured to implement the above - mentioned method.

[0010] In yet another example aspect, a non-transitory computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0011] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 Shows a screen content coding (SCC) decoder flow for loop adaptive color transform (ACT).

[0013] Figure 2 Illustrates the decoding process using ACT.

[0014] Figure 3 Shows an example of a block coded in palette mode.

[0015] Figure 4 Shows an example of signaling palette entries using a palette predictor.

[0016] Figure 5 Shows examples of horizontal and vertical traversal scans.

[0017] Figure 6 Shows an example of coding and decoding of palette indices.

[0018] Figure 7 Is a block diagram showing an example video processing system in which the various techniques disclosed herein may be implemented.

[0019] Figure 8 Is a block diagram of an example hardware platform for video processing.

[0020] Figure 9 Is a block diagram illustrating a video coding and decoding system according to some embodiments of the present disclosure.

[0021] Figure 10 Is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0022] Figure 11 Is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0023] Figures 12 - 13 Shows a flowchart of an example method of video processing. DETAILED DESCRIPTION

[0024] The use of chapter headings in this document is for ease of understanding and does not limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter. Additionally, the use of H.266 terms in some descriptions is for ease of understanding rather than to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.

[0025] 1 Preliminary Discussion

[0026] Video codec standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, and ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 video standard,

[0027] This patent document relates to image / video codec technology. Specifically, it relates to adaptive color transformation in image / video coding. It can be applied to the standard being developed, such as Versatile Video Coding. It can also be applicable to future video codec standards or video codecs.

[0028] 2 Introduction to Video Coding

[0029] Video codec standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, and ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec structure that uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) established the Joint Video Expert Team (JVET) to work on the VVC standard with the goal of reducing the bitrate by 50% compared to HEVC.

[0030] The latest version of the VVC draft, namely Versatile Video Coding (Draft 7), can be found at the following URL: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 16_Geneva / wg11 / JVET-P2001-v14.zip.

[0031] The latest reference software for VVC, named VTM, can be found at the following URL:

[0032] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-7.0

[0033] 2.1 Adaptive Color Transform (ACT) in HEVC-SCC

[0034] At the 18th JCT-VC meeting (June 30 - July 9, 2014, Sapporo, Japan), the Adaptive Color Transform (ACT) was adopted into the HEVC Screen Content Coding (SCC) Test Model 2. ACT performs loop color space conversion in the prediction residual domain using color transform matrices based on the YCoCg and YCoCg-R color spaces. The ACT is adaptively turned on or off at the CU level using the flag cu_residual_act_flag. ACT can be used in combination with Cross Component Prediction (CCP), which is another inter-component decorrelation method already supported in HEVC. When both are enabled, ACT is performed after CCP at the decoder, as Figure 1 shown.

[0035] 2.1.1 Color Space Conversion in ACT

[0036] The color space conversion in ACT is based on the YCoCg-R transform. The same inverse transform is used for both lossy and lossless coding (cu_transquant_bypass_flag = 0 or 1), but in the case of lossy coding, an additional 1-bit left shift is applied to the Co and Cg components. Specifically, the following color space transforms are used for forward and backward conversion in lossy and lossless coding:

[0037] Forward transform for lossy coding (non-normalized):

[0038]

[0039] Forward transform for lossless coding (non-normalized):

[0040] Co = R - B

[0041] t = B + (Co >> 1)

[0042] Cg = (G - t)

[0043] Y = t + (Cg >> 1)

[0044] Inverse transform (normalized):

[0045]

[0046] t = Y - (Cg >> 1)

[0047] G = Cg + t

[0048] B = t - (Co >> 1)

[0049] R = Co + b

[0050] The forward color transform is unnormalized and its norm is approximately equal for Y and Cg and equal for Co To compensate for the unnormalized nature of the forward transform, the increments QP of (-5, -3, -5) are applied to (Y, Co, Cg) respectively. In other words, for a given "normal" QP of the CU, if ACT is enabled, then for (Y, Co, Cg), the quantization parameters are set equal to (QP - 5, QP - 3, QP - 5) respectively. The adjusted quantization parameters only affect the quantization and inverse quantization of the residuals in the CU. For deblocking, the "normal" QP value is still used. Clipping to 0 is applied to the adjusted QP values to ensure that the adjusted QP values do not become negative. Note that this QP adjustment only applies to lossy coding / decoding, because quantization is not performed in lossless coding / decoding (cu_transquant_bypass_flag = 1). In SCM 4, additional PPS / slice-level signaling of QP offset values is introduced. When applying the adaptive color transform, these QP offset values can be used for the CU instead of (-5, -3, -5).

[0051] When the input bit depths of the color components are different, appropriate left shifts are applied during ACT to align the sample bit depth with the maximum bit depth, and appropriate right shifts are applied after ACT to restore the original sample bit depth.

[0052] 2.2 ACT in VVC

[0053] Figure 2 The decoding flow chart of the VVC to which ACT is applied is illustrated. As Figure 2As shown, the color space conversion is performed in the residual domain. Specifically, an additional decoding module, i.e., the inverse ACT, is introduced after the inverse transform to convert the residual from the YCgCo domain back to the original domain.

[0054] In VVC, unless the maximum transform size is smaller than the width or height of a coding unit (CU), a CU leaf node is also used as the unit for transform processing. Therefore, in the proposed implementation, an ACT flag is signaled for a CU to select the color space for encoding / decoding its residual. In addition, following the HEVC ACT design, for inter and IBC CUs, ACT is enabled only when there is at least one non-zero coefficient in the CU. For intra CUs, ACT is enabled only when the chrominance components select the same intra prediction mode as the luma component (i.e., the DM mode).

[0055] The core transform for color space conversion remains the same as that for HEVC. In addition, similar to the ACT design in HEVC, to compensate for the dynamic range change of the residual signaling before and after the color transform, a QP adjustment of (-5, -5, -3) is applied to the transform residual.

[0056] On the other hand, the forward and inverse color transforms require access to the residuals of all three components. Accordingly, in the proposed implementation, ACT is disabled in the following two cases where not all the residuals of the three components are available.

[0057] 1. Split tree partitioning: When using a split tree, the luma and chroma samples within a CTU are partitioned by different structures. This results in CUs in the luma tree containing only the luma component, while CUs in the chroma tree containing only the two chroma components.

[0058] 2. Intra sub-partition prediction (ISP): The ISP sub-partitioning is only applied to luma, while the chroma signaling is encoded / decoded without being partitioned. In the current ISP design, except for the last ISP sub-partition, other sub-partitions contain only the luma component.

[0059] The coding unit draft in the VVC draft is as follows.

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069] When cu_act_enabled_flag equals 1, it indicates that the residual of the current coding / decoding unit is coded / decoded in the YCgCo color space. When cu_act_enabled_flag equals 0, it indicates that the residual of the current coding / decoding unit is coded / decoded in the original color space. When cu_act_enabled_flag does not exist, it is inferred to be equal to 0.

[0070] 2.3 Transform Skip Mode in VVC

[0071] Similar to HEVC, the transform skip mode can be used to code / decode the residual of a block, which completely skips the transform process of the block. In addition, for transform skip blocks, the minimum allowable quantization parameter (QP) signaled in the SPS is used, which is set to be equal to 6*(internalBitDepth–inputBitDepth)+4 in VTM7.0.

[0072] 2.4 Block-based Delta Pulse Code Modulation (BDPCM)

[0073] In JVET-M0413, a block-based delta pulse code modulation (BDPCM) was proposed to efficiently code / decode screen content and then applied to VVC.

[0074] The prediction directions used in BDPCM can be vertical and horizontal prediction modes. Intra prediction of the entire block is performed by sample copying in the prediction direction similar to intra prediction (horizontal or vertical prediction). The residual is quantized, and the difference between the quantized residual and the quantized value of its predictor (horizontal or vertical) is coded / decoded. This can be described as follows: for a block of size M (rows) × N (columns), let r i,j , 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 be the prediction residual after performing intra prediction by copying the unfiltered samples from the upper or left block boundary samples horizontally (copying the left neighboring pixel values of the prediction block row by row) or vertically (copying the top neighboring line to each row in the prediction block).

[0075] Let Q(r i,j), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 represents the residual r i,j is the quantized version, where the residual is the difference between the original block and the predicted block. Then block DPCM is applied to the quantized residual samples, resulting in a modified M×N array with elements When signaling indicates vertical BDPCM: When signaling indicates vertical BDPCM:

[0076]

[0077] For horizontal prediction, a similar rule is applied, and the residual quantized samples are obtained by

[0078]

[0079] The remaining quantized samples are sent to the decoder.

[0080] At the decoder side, the above calculations are reversed to produce Q(r i,j ), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1.

[0081] For the vertical prediction case,

[0082]

[0083] For the horizontal case,

[0084]

[0085] The inverse quantized residual Q -1 (Q(r i,j )) is added to the in-block prediction value to produce the reconstructed sample value.

[0086] The main advantage of this scheme is that inverse BDPCM can be completed in real time during coefficient parsing, either by simply adding predictors when parsing the coefficients or by performing it after parsing.

[0087] In VTM7.0, BDPCM can also be applied to chroma blocks, and chroma BDPCM has different flags and BDPCM directions from the luma BDPCM mode.

[0088] 2.5 Scaling Process of Transform Coefficients

[0089] The text regarding the scaling process of transform coefficients in JVET-P2001-vE is as follows.

[0090] The inputs to this process are:

[0091] – The luma position (xTbY, yTbY), specifying the top-left sample of the current luma transform block relative to the top-left luma sample of the current picture,

[0092] – Variable nTbW, specifying the transform block width,

[0093] – Variable nTbH, specifying the transform block height,

[0094] – Variable predMode, specifying the prediction mode of the coding / decoding unit

[0095] – Variable cIdx, specifying the color component of the current block,

[0096] The output of this process is an (nTbW)×(nTbH) array d of scaled transform coefficients with elements d[x][y].

[0097] The quantization parameter qP is derived as follows:

[0098] – If cIdx is equal to 0, the following applies:

[0099] qP = Qp′ Y (1129)

[0100] – Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, the following applies:

[0101] qP = Qp′ CbCr (1130)

[0102] – Otherwise, if cIdx is equal to 1, the following applies:

[0103] qP = Qp′ Cb (1131)

[0104] – Otherwise (cIdx is equal to 2), the following applies:

[0105] qP = Qp′ Cr (1132)

[0106] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0107] – If transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0, the following applies:

[0108] qP = qP - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1133)

[0109] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1)? 1 : 0 (1134)

[0110] bdShift = BitDepth + rectNonTsFlag + ((Log2(nTbW) + Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag (1135)

[0111] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] equals 1), the following applies:

[0112] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1136)

[0113] rectNonTsFlag = 0 (1137)

[0114] bdShift = 10 (1138)

[0115] The variable bdOffset is derived as follows:

[0116] bdOffset = (1 << bdShift) >> 1 (1139)

[0117] Specify the list levelScale[][] as levelScale[j][k] = {{40, 45, 51, 57, 64, 72}, {57, 64, 72, 80, 90, 102}} where j = 0..1 and k = 0..5.

[0118] Set the (nTbW) x (nTbH) array dz to be equal to the (nTbW) x (nTbH) array TransCoeffLevel[xTbY][yTbY][cIdx].[[]]

[0119] For the derivation of the scaled transform coefficient d[x][y] where x = 0..nTbW - 1 and y = 0..nTbH - 1, the following applies:

[0120] – The intermediate scaling factor m[x][y] is derived as follows:

[0121] – If one or more of the following conditions are true, then m[x][y] is set to be equal to 16:

[0122] – sps_scaling_list_enabled_flag equals 0.

[0123] – pic_scaling_list_present_flag equals 0.

[0124] – The transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1.

[0125] – The scaling_matrix_for_lfnst_disabled_flag is equal to 1 and lfnst_idx[xTbY][yTbY] is not equal to 0.

[0126] – Otherwise, the following applies:

[0127] – Derive the variable id based on predMode, cIdx, nTbW, and nTbH specified in Table 36. The variable log2MatrixSize is derived as follows:

[0128] log2MatrixSize = (id < 2)? 1 : (id < 8)? 2 : 3 (1140)

[0129] – The scaling factor m[x][y] is derived as follows:

[0130] m[x][y] = ScalingMatrixRec[id][i][j], where i = (x << log2MatrixSize) >> Log2(nTbW), j = (y << log2MatrixSize) >> Log2(nTbH) (1141)

[0131] – If id is greater than 13 and both x and y are equal to 0, then m[0][0] is further modified as follows:

[0132] m[0][0] = ScalingMatrixDCRec[id - 14] (1142)

[0133] Note—When any of the following conditions is true, the quantization matrix element m[x][y] can be set to zero

[0134] – x is greater than 32

[0135] – y is greater than 32

[0136] – The decoded tu default transform mode is not coded (i.e., the transform type is not equal to 0) and x is greater than 16

[0137] – The decoded tu default transform mode is not coded (i.e., the transform type is not equal to 0) and y is greater than 16

[0138] – The scaling factor ls[x][y] is derived as follows:

[0139] – If pic_dep_quant_enabled_flag equals 1 and transform_skip_flag[xTbY][yTbY][cIdx] equals 0, then the following applies:

[0140] ls[x][y] = (m[x][y] * levelScale[rectNonTsFlag][(qP + 1) % 6]) << ((qP + 1) / 6) (1143)

[0141] – Otherwise (pic_dep_quant_enabled_flag equals 0 or transform_skip_flag[xTbY][yTbY][cIdx] equals 1), the following applies:

[0142] ls[x][y] = (m[x][y] * levelScale[rectNonTsFlag][qP % 6]) << (qP / 6) (1144)

[0143] – When BdpcmFlag[xTbY][yYbY][cIdx] equals 1, dz[x][y] is modified as follows:

[0144] – If BdpcmDir[xTbY][yYbY][cIdx] equals 0 and x is greater than 0, then the following applies:

[0145] dz[x][y] = Clip3(CoeffMin, CoeffMax, dz[x - 1][y] + dz[x][y]) (1145)

[0146] – Otherwise, if BdpcmDir[xTbY][yTbY][cIdx] equals 1 and y is greater than 0, then the following applies:

[0147] dz[x][y] = Clip3(CoeffMin, CoeffMax, dz[x][y - 1] + dz[x][y]) (1146)

[0148] – The value dnc[x][y] is derived as follows:

[0149] dnc[x][y] = (dz[x][y] * ls[x][y] + bdOffset) >> bdShift (1147)

[0150] – The scaled transform coefficient d[x][y] is derived as follows:

[0151] d[x][y] = Clip3(CoeffMin, CoeffMax, dnc[x][y]) (1148)

[0152] Table 36 - Specification of the scaling matrix identifier variable id according to predMode, cIdx, nTbW, and nTbH

[0153]

[0154] 2.6 Palette Mode

[0155] 2.6.1 Concept of Palette Mode

[0156] The basic idea behind the palette mode is that the pixels in a CU are represented by a small set of representative color values. This set is called the palette. Samples outside the palette can also be indicated by signaling an escape symbol for (possibly quantized) component values. Such pixels are called escape pixels. The palette mode is as Figure 3 shown. As Figure 3 shown, for each pixel with three coloc components (luminance and two chrominance components), a palette index is established, and the block can be reconstructed according to the values found in the palette.

[0157] 2.6.2 Coding and Decoding of Palette Entries

[0158] For the coding and decoding of palette entries, the palette predictor is retained. The maximum size of the palette and the palette predictor are signaled in the SPS. In HEVC - SCC, palette_predictor_initializer_present_flag is introduced in the PPS. When this flag is 1, the entries used to initialize the palette predictor are signaled in the bitstream. The palette predictor is initialized at the start of each CTU row, each slice, and each picture. Depending on the value of palette_predictor_initializer_present_flag, the palette predictor is reset to 0 or initialized using the palette predictor initializer entries signaled in the PPS. In HEVC - SCC, a palette predictor initializer of size 0 is enabled to allow explicit disabling of palette predictor initialization at the PPS level.

[0159] For each entry in the palette predictor, a reuse flag is sent to indicate whether it is part of the current palette. This is in Figure 4It is illustrated in []. The reuse flag is sent using zero-length encoding and decoding. After that, the number of new palette entries is signaled using an Exponential Golomb (EG) code of order 0 (i.e., EG-0). Finally, the component values of the new palette entries are signaled.

[0160] 2.6.3 Encoding and Decoding of Palette Indexes

[0161] The palette index is encoded and decoded using horizontal and vertical traversal scans, as Figure 5 shown. The scan order is signaled explicitly in the bitstream using palette_transpose_flag. For the remainder of this subsection, it is assumed that the scan is horizontal.

[0162] The palette index is encoded and decoded using two palette sample modes: "COPY_LEFT" and "COPY_ABOVE". In the "COPY_LEFT" mode, the palette index is assigned to the decoded index. In the "COPY_ABOVE" mode, the palette index of the samples in the previous row is copied. For both the "COPY_LEFT" and "COPY_ABOVE" modes, a run value is signaled, which specifies the number of subsequent samples that are also encoded and decoded using the same mode.

[0163] In the palette mode, the index value of the escape symbol is the number of palette entries. Also, when the escape symbol is part of a run in the "COPY_LEFT" or "COPY_ABOVE" mode, the escape component value is signaled for each escape symbol. The encoding and decoding of the palette index are as Figure 6 shown.

[0164] The syntax is completed as follows. First, the number of index values of the CU is signaled. Subsequently, the actual index values of the entire CU are signaled using truncated binary encoding. Both the number of indexes and the index values are encoded in bypass mode. This groups the bypass binary numbers (bins) related to the indexes together. Then, the palette sample mode (if required) and the run are signaled in an interleaved manner. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and encoded in bypass mode. The binarization of the escape symbol is EG encoding of order 3, i.e., EG-3.

[0165] After signaling the index values, the additional syntax element last_run_type_flag is signaled. This syntax element, combined with the number of indexes, eliminates the need to signal the run value corresponding to the last run in the block.

[0166] In HEVC-SCC, the palette mode is also enabled for 4:2:2, 4:2:0, and monochrome chroma formats. For all chroma formats, the signaling of palette entries and palette indices is almost the same. In the case of non-monochrome formats, each palette entry consists of 3 components. For monochrome formats, each palette entry consists of a single component. For subsampled chroma directions, chroma samples are associated with luma sample indices divisible by 2. After the palette index of the reconstructed CU, if the sample has only one associated component, only the first component of the palette entry is used. The only difference in signaling is the escape component value. For each escape sample, the number of escape component values signaled may vary depending on the number of components associated with that sample.

[0167] 2.6.4 Palette in the dual-tree

[0168] In VVC, a dual-tree coding structure is used for intra-strip coding, so the luma component and the two chroma components can have different palettes and palette indices. Additionally, the two chroma components share the same palette and palette index.

[0169] 2.6.5 Line-based CG palette mode

[0170] VVC adopts a line-based CG palette mode. In this method, each CU in the palette mode is divided into multiple segments of m samples (m = 16 in this test) based on a traversal scan mode. The encoding order of palette run coding and decoding in each segment is as follows: For each pixel, a context-coded binary run_copy_flag = 0 is signaled to indicate whether the pixel has the same mode as the previous pixel, that is, whether the previously scanned pixel and the current pixel are both of the run type COPY_ABOVE, or whether the previously scanned pixel and the current pixel are both of the run type INDEX and have the same index value. Otherwise, run_copy_flag = 1 is signaled. If the pixel and the previous pixel have different modes, a context-coded binary number copy_above_palette_indices_flag is signaled to indicate the run type of the pixel, that is, INDEX or COPY_ABOVE. Similar to the palette mode in VTM6.0, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to parse the run type because the INDEX mode is used by default. Additionally, if the previously parsed run type is COPY_ABOVE, the decoder also does not have to parse the run type. After palette run coding and decoding of the pixels in a segment, the index value (for the INDEX mode) and the quantized escape color are bypass-coded and grouped separately from the coding / parsing of the context-coded binary numbers to improve the throughput within each row of CG. Since the index value is now coded / parsed after run coding and decoding, instead of being processed before palette run coding and decoding as in VTM, the encoder does not have to signal the number of index values num_palette_indices_minus1 and the last run type copy_above_indices_for_final_run_flag.

[0171] 2.7 Picture Header and Strip Header

[0172] In this section, the newly added text is highlighted with highlighting. The deleted text is marked with italic text.

[0173] The following text about the Picture Parameter Set, Picture Header, and Strip Header comes from the latest status of the VVC HLS design and has some differences from JVET-P2001_vE, including the following two aspects:

[0174] 1) Signaling of collocated pictures, which is in the PH when the RPL is in the PH and in the SH when the RPL is in the SH;

[0175] 2) Signaling of the WP table, when RPL is in PH, it is in PH, and when RPL is SH, it is in SH.

[0176] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0177]

[0178]

[0179]

[0180]

[0181] 7.3.2.6 Picture Header RBSP Syntax

[0182]

[0183] 7.3.2.7 Picture Header Structure Syntax

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190] 7.3.7.1 General Strip Header Syntax

[0191]

[0192]

[0193]

[0194]

[0195] 7.3.7.2 Weighted Prediction Parameter Syntax

[0196]

[0197]

[0198] When a picture consists of more than one VCL NAL unit, the PH NAL unit shall be present in the PU.

[0199] If a PH NAL unit is present in a PU, the first VCL NAL unit of the picture is the first VCL NAL unit after the PH NAL unit in the decoding order of the picture. Otherwise (no PH NAL unit in the PU), the first VCL NAL unit of the picture is the only VCL NAL unit of the picture.

[0200] 2.8. Embodiments of JVET-Q0513 and JVET-Q0820

[0201] In JVET-Q0513, a clipping problem was raised where the residual might exceed the 16-bit dynamic range after inverse ACT. One solution is to clip the input of inverse ACT according to the range of [-(1<<BitDepth),(1<<BitDepth)–1].

[0202] However, when using the YCgCo-R color space described in, for example, JVET-Q0820-v3, the dynamic ranges of Co and Cg are doubled and may exceed the range of [-(1<<BitDepth),(1<<BitDepth)–1].

[0203] 3 Technical problems solved by the embodiments and solutions described in this paper

[0204] 1. In the current design, ACT and the luminance BDPCM mode can be enabled for a block. However, the chrominance BDPCM mode is always disabled on the blocks encoded and decoded using the ACT mode. Therefore, the prediction signaling can be derived differently for the luminance and chrominance blocks in the same coding unit, which is less efficient.

[0205] 2. When ACT is enabled, the quantization parameter (QP) of a block can become negative.

[0206] 3. The current design of ACT does not support lossless coding and decoding.

[0207] 4. The signaling of using ACT does not depend on the block size.

[0208] 5. The maximum palette size and the maximum predictor size are fixed numbers, which may limit the flexibility of the palette mode.

[0209] 6. The escape samples use the third-order exponential Golomb (EG) as the binarization method, but the binarization of the escape samples does not depend on the quantization parameter (QP).

[0210] 7. Since when signaling a reference picture list (RPL) in a Picture header (PH), the number of active entries for each slice is still signaled in the slice header, the co-located signaling is sub-optimal and the signaling of the weighted prediction (WP) table is disrupted.

[0211] 8. For the YCgCo-R color space, the range of [-(1<<BitDepth),(1<<BitDepth)–1] may not be sufficient to accommodate the residuals.

[0212] 4 Technical Solutions

[0213] The technical solutions described below should be regarded as examples to explain general concepts. These technical solutions should not be interpreted narrowly. In addition, these technical solutions can be combined in any way.

[0214] In the following description, the term "block" may represent a video region, such as a coding unit (CU), a prediction unit (PU), or a transform unit (TU), which may contain samples of three color components. The term "BDPCM" is not limited to the design in VVC and may also represent a technique for coding and decoding residuals using different prediction signaling generation methods.

[0215] In the following description, a video block coded using the Joint Coding of Chroma Residuals (JCCR) mode includes signaling only one chroma residual block (e.g., the Cb residual block), and deriving another chroma residual block (e.g., the Cr residual block) based on the signaled chroma residual block and one or more flags indicating a specific JCCR mode (e.g., at the transform unit level). As described above, the JCCR mode utilizes the correlation between the Cb residual and the Cr residual to improve the coding and decoding efficiency.

[0216] Interaction between ACT and BDPCM (Items 1 - 4)

[0217] 1. Whether to enable the chroma BDPCM mode may depend on the use of ACT and / or the luma BDPCM mode. a. In one example, when ACT is enabled on a block, the indication of the use of the chroma BDPCM mode (e.g., intra_bdpcm_chroma_flag) may be inferred as the indication of the use of the luma BDPCM mode (e.g., intra_bdpcm_luma_flag).

[0218] i. In one example, the inferred value of the chroma BDPCM mode is defined as (Is ACT and the luma BDPCM mode enabled? True: False).

[0219] 1. In one example, when intra_bdpcm_luma_flag is false, intra_bdpcm_chroma_flag can be set to false.

[0220] a. Alternatively, when intra_bdpcm_luma_flag is true, intra_bdpcm_chroma_flag can be set to true.

[0221] ii. Alternatively, in one example, if the use indication of the luminance BDPCM mode and ACT for a block is true, then the use indication of the chrominance BDPCM mode can be inferred to be true.

[0222] b. Alternatively, it can be conditionally checked whether the use of ACT for a block is signaled, such as using the same BDPCM prediction direction for luminance and chrominance samples in the block.

[0223] i. Alternatively, in addition, the use indication of ACT is signaled after using the BDPCM mode.

[0224] 2. When ACT is enabled for a block, the indication of the prediction direction of the chrominance BDPCM mode (e.g., intra_bdpcm_chroma_dir_flag) can be inferred as the indication of the prediction direction of the luminance BDPCM mode used (e.g., intra_bdpcm_luma_dir_flag).

[0225] a. In one example, define the inferred value of intra_bdpcm_chroma_dir_flag as (ACT enabled? intra_bdpcm_luma_dir_flag : 0).

[0226] i. In one example, if the indication of the prediction direction of the luminance BDPCM mode is horizontal, the indication of the prediction direction of the chrominance BDPCM mode can be inferred to be horizontal.

[0227] ii. Alternatively, in one example, if the indication of the prediction direction of the luminance BDPCM mode is vertical, the indication of the prediction direction of the chrominance BDPCM mode can be inferred to be vertical.

[0228] 3. ACT and BDPCM mode can be applied separately.

[0229] a. In one example, when the ACT mode is enabled for a block, the BDPCM mode can be disabled for the block.

[0230] i. Alternatively, in addition, the usage indication of the BDPCM mode may be signaled after the usage indication of the ACT mode is signaled.

[0231] ii. Alternatively, in addition, the usage indication of the BDPCM mode may not be signaled and may be inferred as false (0).

[0232] b. In one example, when the BDPCM mode is enabled on a block, the ACT mode may be disabled on the block.

[0233] i. Alternatively, in addition, the usage indication of the ACT mode may be signaled after the usage indication of the BDPCM mode is signaled.

[0234] ii. Alternatively, in addition, the usage indication of the ACT mode may not be signaled and may be inferred as false (0).

[0235] c. In one example, the BDPCM mode in the above example may represent a luminance BDPCM mode and / or a chrominance BDPCM mode.

[0236] 4. Inverse ACT may be applied before inverse BDPCM in the decoder.

[0237] a. In one example, ACT may be applied even when the luminance and chrominance BDPCMs have different prediction modes.

[0238] b. Alternatively, at the encoder, forward ACT may be applied after BDPCM.

[0239] QP Setting when ACT is Enabled (Item 5)

[0240] 5. It is proposed to clip the QP when ACT is enabled.

[0241] a. In one example, the clipping function may be defined as (l, h, x), where l is the lowest possible value of the input x and h is the highest possible value of the input x.

[0242] i. In one example, l may be set to be equal to 0.

[0243] ii. In one example, h may be set to be equal to 63.

[0244] b. In one example, the QP may be the qP given in Section 2.5.

[0245] c. In one example, clipping may be performed after the QP adjustment for the ACT mode.

[0246] d. In one example, when transform skipping is applied, l may be set to be equal to the minimum allowed QP for the transform skipping mode.

[0247] Related to Palette Mode (Items 6 - 7)

[0248] 6. The value of the maximum allowable palette size and / or the maximum allowable predictor size can depend on the codec characteristics. Suppose S 1 is the maximum palette size (or palette predictor size) associated with the first codec characteristic; S 2 is the maximum palette size (or palette predictor size) associated with the second codec characteristic.

[0249] a. In one example, the codec characteristic can be a color component.

[0250] i. In one example, the maximum allowable palette size and / or the maximum allowable predictor size for different color components can have different values.

[0251] ii. In one example, the value of the maximum allowable palette size and / or the maximum allowable predictor size for the first color component (e.g., Y in YCbCr, G in RGB) can be different from that of the other two color components (e.g., Cr and Cb in YCbCr, B and R in RGB), excluding the first color component.

[0252] b. In one example, the codec characteristic can be a quantization parameter (QP).

[0253] i. In one example, if QP 1 is greater than QP 2 , then S 1 for QP 1 and / or S 2 should be less than S 2 for QP 1 and / or S 2

[0254] ii. In one example, the QP can be a slice-level QP or a block-level QP.

[0255] c. In one example, S 2 can be greater than or equal to S 1 .

[0256] d. For the first and second codec characteristics, the indication of the maximum palette size / palette predictor size can be signaled separately or inferred from one another.

[0257] i. In one example, S 1 can be signaled and S 1 can be derived based on S 2 .

[0258] 1. In one example, S2 can be inferred as S 1 –n.

[0259] 2. In one example, S 2 can be inferred as S 1 >>n.

[0260] 3. In one example, S 2 can be inferred as floor(S 1 / n), where floor(x) represents the largest integer not greater than x.

[0261] e. In one example, S 1 and / or S 2 can be signaled at a high level (e.g., SPS / PPS / PH / strip header) and adjusted at a lower level (e.g., CU / block).

[0262] i. How to adjust S 1 and / or S 2 can depend on the codec information.

[0263] 1. How to adjust S 1 and / or S 2 can depend on the current QP.

[0264] a. In one example, if the current QP increases, then S 1 and / or S 2 should decrease.

[0265] 2. How to adjust S 1 and / or S 2 can depend on the block dimension.

[0266] a. In one example, if the current block size increases, then S 1 and / or S 2 .

[0267] f. S 1 and / or S 2 can depend on whether LMCS is used.

[0268] 7. Parameters related to the binarization method of escape samples / pixels can depend on the codec information, such as the quantization parameter (QP).

[0269] a. In one example, the EG binarization method can be used, and the order of EG binarization represented by k can depend on the codec information.

[0270] i. In one example, k can decrease when the current QP increases.

[0271] Signaling Notification of ACT Mode (Items 8 - 10)

[0272] 8. An indication of the maximum and / or minimum allowable ACT size may be signaled or derived at the sequence / video / strip / slice / sub - picture / tile / other video processing unit level based on codec information.

[0273] a. In one example, they may be signaled in the SPS / PPS / picture header / strip header.

[0274] b. In one example, they may be signaled conditionally, e.g., according to whether ACT is enabled.

[0275] c. In one example, the N level of the maximum and / or minimum allowable ACT size may be signaled / defined, e.g., N = 2.

[0276] i. In one example, the maximum and / or minimum allowable ACT size may be set to K0 or K1 (e.g., K0 = 64, K1 = 32).

[0277] ii. Alternatively, in addition, an indication of the level may be signaled, e.g., when N = 2, a flag may be signaled.

[0278] d. In one example, an indication of the difference between the maximum and / or minimum allowable ACT size and the maximum and / or minimum allowable transform (or transform skip) size (e.g., for the luminance component) may be signaled.

[0279] e. In one example, the maximum and / or minimum allowable ACT size may be derived from the maximum and / or minimum (or transform skip) size (e.g., for the luminance component).

[0280] f. Alternatively, in addition, whether and / or how to signal an indication of the use of ACT and other auxiliary information related to ACT may depend on the maximum and / or minimum allowable values.

[0281] 9. When a block is larger than the maximum allowable ACT size (or the maximum allowable transform size), the block may be automatically divided into multiple sub - blocks, where all sub - blocks share the same prediction mode (e.g., all sub - blocks are intra - coded), and ACT may be enabled at the sub - block level instead of the block level.

[0282] 10. An indication of the use of the ACT mode may be signaled conditionally based on the block dimensions (e.g., block width and / or height, block width times height, ratio between block width and height, maximum / minimum of block width and height) and / or the maximum allowable ACT size.

[0283] a. In one example, when certain conditions (e.g., according to block dimensions) are met, an indication of the use of the ACT mode can be signaled.

[0284] i. In one example, the condition is whether the current block width is less than or equal to m and / or whether the current block height is less than or equal to n.

[0285] ii. In one example, the condition is whether the product of the current block width and height is less than or not greater than m.

[0286] iii. In one example, the condition is whether the product of the current block width and height is greater than or not less than m.

[0287] b. Alternatively, in one example, when certain conditions (e.g., according to block dimensions) are not met, the usage pattern of ACT may not be signaled.

[0288] i. In one example, the condition is whether the current block width is greater than m and / or whether the current block height is greater than n.

[0289] ii. In one example, the condition is whether the product of the current block width and height is less than or not greater than m.

[0290] iii. In one example, the condition is whether the product of the current block width and height is greater than or not less than m.

[0291] iv. Alternatively, in addition, the indication of the use of the ACT mode can be inferred as 0.

[0292] c. In the above examples, the variables m, n can be predefined (e.g., 4, 64, 128), signaled, or derived on-the-fly.

[0293] i. In one example, m and / or n can be derived based on the decoding messages of SPS / PPS / APS / CTU row / CTU group / CU / block.

[0294] 1. In one example, m and / or n can be set to be equal to the maximum allowed transform size (e.g., MaxTbSizeY).

[0295] Signaling Notification of Constraint Flags in General Constraint Information Syntax (Items 11 - 16)

[0296] The following constraint flags can be signaled in video units other than SPS. For example, they can be signaled in the general constraint information syntax specified in JVET-P2001-vE.

[0297] 11. It is proposed to have a constraint flag to specify whether the SPS ACT enable flag (e.g., sps_act_enabled_flag) should be equal to 0.

[0298] a. In one example, the flag can be represented as no_act_constraint_flag

[0299] i. When the flag equals 1, the SPS ACT enable flag (e.g., sps_act_enabled_flag) shall equal 0.

[0300] ii. When the flag equals 0, no such constraint is imposed.

[0301] 12. It is proposed to have a constraint flag to specify whether the SPS BDPCM enable flag (e.g., sps_bdpcm_enabled_flag) should equal 0.

[0302] a. In one example, the flag can be represented as no_bdpcm_constraint_flag.

[0303] i. When the flag equals 1, the SPS BDPCM enable flag (e.g., sps_bdpcm_enabled_flag) shall equal 0.

[0304] ii. When the flag equals 0, no such constraint is imposed.

[0305] 13. It is proposed to have a constraint flag to specify whether the SPS chroma BDPCM enable flag (e.g., sps_bdpcm_chroma_enabled_flag) should equal 0.

[0306] a. In one example, the flag can be represented as no_bdpcm_chroma_constraint_flag.

[0307] i. When the flag equals 1, the SPS chroma BDPCM enable flag (e.g., sps_bdpcm_chroma_enabled_flag) shall equal 0.

[0308] ii. When the flag equals 0, no such constraint is imposed.

[0309] 14. It is proposed to have a constraint flag to specify whether the SPS palette enable flag (e.g., sps_palette_enabled_flag) should equal 0.

[0310] a. In one example, the flag can be represented as no_palette_constraint_flag.

[0311] i. When this flag equals 1, the SPS palette enable flag (e.g., sps_palette_enabled_flag) shall equal 0.

[0312] ii. When this flag equals 0, such a constraint is not imposed.

[0313] 15. It is proposed to have a constraint flag to specify whether the SPS RPR enable flag (e.g., ref_pic_resampling_enabled_flag) shall equal 0.

[0314] a. In one example, this flag can be expressed as

[0315] no_ref_pic_resampling_constraint_flag.

[0316] i. When this flag equals 1, the SPS RPR enable flag (e.g., ref_pic_resampling_enabled_flag) shall equal 0.

[0317] ii. When this flag equals 0, such a constraint is not imposed.

[0318] 16. In the above example (bullets 11 - 15), such a constraint flag can be signaled conditionally, e.g., according to the chroma format (e.g., chroma_format_idc) and / or separate plane coding or ChromaArrayType.

[0319] ACT QP Offset (Items 17 - 19)

[0320] 17. It is proposed that when applying ACT to a block, the ACT offset can be applied after applying other chroma offsets (e.g., chroma offsets in PPS and / or Picture Header (PH) and / or Slice Header (SH)).

[0321] 18. It is proposed that when applying the YCgCo color transform to a block, for JCbCr mode 2, set the PPS and / or PH offset other than -5.

[0322] a. In one example, the offset can be other than -5.

[0323] b. In one example, the offset can be indicated in the PPS (e.g., as pps_act_cbcr_qp_offset_plus6), and the offset can be set to pps_act_cbcr_qp_offset_plus6 - 6.

[0324] c. In one example, the offset can be indicated in the PPS (e.g., as pps_act_cbcr_qp_offset_plus7), and the offset can be set to pps_act_cbcr_qp_offset_plus7 - 7.

[0325] 19. It is proposed that when applying YCgCo - R to a block, for JCbCr mode 2, the PPS and / or PH offset is not 1.

[0326] a. In one example, the offset can be not - 1.

[0327] b. In one example, the offset can be indicated in the PPS (e.g., as pps_act_cbcr_qp_offset), and the offset can be set to pps_act_cbcr_qp_offset.

[0328] c. In one example, the offset can be indicated in the PPS (e.g., as pps_act_cbcr_qp_offset_plus1), and the offset can be set to pps_act_cbcr_qp_offset_plus1 - 1.

[0329] 20. It is proposed that when using JCCR, the QP offset of the ACT with YCgCo transform, denoted as act_qp_offset, can depend on the JCCR mode.

[0330] a. In one example, when the JCCR mode is 1, act_qp_offset can be - 5.

[0331] i. Alternatively, in one example, when the JCCR mode is 1, act_qp_offset can be - 6.

[0332] b. In one example, when the JCCR mode is 2, act_qp_offset can be - 7.

[0333] i. Alternatively, in one example, when the JCCR mode is 2, act_qp_offset can be (-7 - pps_joint_cbcr_qp_offset - slice_joint_cbcr_qp_offset).

[0334] ii. Alternatively, in one example, when the JCCR mode is 2, act_qp_offset can be (-7 + pps_joint_cbcr_qp_offset + slice_joint_cbcr_qp_offset).

[0335] c. In one example, when the JCCR mode is 3, the ACT offset can be -4.

[0336] i. Alternatively, in one example, when the JCCR mode is 3, the ACT offset can be -5.

[0337] 21. It is proposed that when using JCCR, the QP offset of the ACT with YCgCo-R transform, denoted as act_qp_offset, can depend on the JCCR mode.

[0338] a. In one example, when the JCCR mode is 1, act_qp_offset can be 1.

[0339] i. Alternatively, in one example, when the JCCR mode is 1, act_qp_offset can be 0.

[0340] b. In one example, when the JCCR mode is 2, act_qp_offset can be -1.

[0341] i. Alternatively, in one example, when the JCCR mode is 2, act_qp_offset can be (-1 - pps_joint_cbcr_qp_offset - slice_joint_cbcr_qp_offset).

[0342] ii. Alternatively, in one example, when the JCCR mode is 2, act_qp_offset can be (-1 + pps_joint_cbcr_qp_offset + slice_joint_cbcr_qp_offset).

[0343] c. In one example, when the JCCR mode is 3, the ACT offset can be 2.

[0344] i. Alternatively, in one example, when the JCCR mode is 3, the ACT offset can be 1.

[0345] 22. Whether to signal the QP offset (such as slice_act_cbcr_qp_offset) of the ACT and JCCR coding / decoding blocks can depend on whether the condition check of JCCR is enabled.

[0346] a. Alternatively, in addition, whether there is an indication of the QP offset for the ACT and JCCR encoding / decoding blocks (e.g., pps_joint_cbcr_qp_offset_present_flag / pps_slice_cbcr_qp_offset_present_flag) can depend on whether the condition check for JCCR is enabled.

[0347] 23. How to derive / signal the delta QP can depend on the encoding / decoding information, such as the encoding / decoding mode and the use of encoding / decoding tools.

[0348] a. In one example, how to derive / signal the delta QP can depend on the use of ACT for the current block and / or the previously encoded / decoded blocks (e.g., the upper and left neighboring blocks).

[0349] b. In one example, deriving / signaling the delta QP can depend on whether the current block and the previously encoded / decoded blocks used for deriving the QP predictor share the same encoding / decoding mode or the enabled / disabled state of an encoding / decoding tool (e.g., ACT).

[0350] c. In one example, for the current block encoded using the encoding / decoding tool X (e.g., ACT / transform skip), the QP predictor derivation process can be different from that of other blocks where X is disabled.

[0351] i. In one example, the QP predictor for the current block encoded using the encoding / decoding tool X can be derived from those blocks where the encoding / decoding tool X is enabled.

[0352] ii. In one example, the QP predictor for the current block not encoded using the encoding / decoding tool X can be derived from those blocks where the encoding / decoding tool X is disabled.

[0353] 24. How to signal the QP offset in the PPS can be independent of the SPS.

[0354] a. Alternatively, in addition, an indication of the color format and / or an indication of the enabling of ACT can be signaled in the PPS.

[0355] i. Alternatively, in addition, the QP offset can be signaled under the indicated condition check.

[0356] 25. Whether to signal the QP offset (e.g., applied to the ACT encoded blocks) at the first level (e.g., picture level)

[0357] a. It is proposed that if the QP offset (e.g., applied to the ACT encoded blocks) exists at the second level (e.g., slice level), the signaling of the QP offset at the first level is skipped.

[0358] 26. The QP mentioned in the above example can also exist at the picture / strip / slice / sub - picture level (e.g., picture header / PPS / strip header).

[0359] 27. An override mechanism can be used to signal the available QP offset. That is, the QP offset can be signaled at the first level (e.g., PPS), but overridden at the second level (e.g., picture header).

[0360] a. Alternatively, in addition, a flag can be signaled at the first level or the second level to indicate whether the override is enabled.

[0361] b. Alternatively, when the override is enabled, there is no need to signal the relevant information at the first level.

[0362] i. In one example, if an override is applied to the second level (e.g., strip), the QP offset (e.g., applied to the ACT coding / decoding block) may not be signaled at the first level (e.g., PPS).

[0363] c. Alternatively, in addition, the difference in QP offset between the first and second levels can be signaled at the second level.

[0364] 28. In one example, for the first ACT coding / decoding block, the first QP and the second QP of the i - th (e.g., i is from 0 to 2) color component are derived. The first QP of the i - th color component is used for reconstruction and / or de - blocking, and the second QP of the i - th color component is used to predict the QP of subsequent coding / decoding blocks. The first QP and the second QP of the i - th color component can be different.

[0365] a. In one example, the first luma QP can be equal to the second luma QP plus one or more offsets.

[0366] i. In one example, the offset can be signaled, e.g., in the SPS / PPS / picture header / strip header / CTU / CU / other video processing units (e.g., pps_act_y_qp_offset, pps_act_cb_qp_offset, pps_act_cb_qp_offset, pps_act_cbcr_qp_offset, slice_act_y_qp_offset, slice_act_cb slice_qp_offset, slice_act_cb slice_qp_offset).

[0367] ii. In one example, the offset can be fixed (e.g., - 5, - 3, 1, 3).

[0368] b. In one example, it is not allowed to consider the QP offsets signaled in the PPS / strip header / picture header of the ACT codec block (e.g., pps_act_y_qp_offset, pps_act_cb_qp_offset, pps_act_cb_qp_offset, pps_act_cbcr_qp_offset, slice_act_y_qp_offset, slice_act_cb_qp_offset, predict_act_c_offset)

[0369] to predict the QP value of the block to be coded / decoded subsequently.

[0370] Related to Deblocking Filter

[0371] 29. It is proposed to consider the QP offsets signaled / derived in the SPS / PPS / picture header / strip header during the deblocking filter process (e.g., in determining the Tc and / or Beta parameters used in the deblocking filter process) and / or during the predictive coding of the incremental QP of other blocks.

[0372] a. In one example, the QP used in the deblocking filter process can depend on the QP offsets signaled / derived in the SPS / PPS / picture header / strip header.

[0373] b. In one example, the QP offset is related to the luminance component (or the G component in RGB coding).

[0374] Regarding Picture Header

[0375] 30. When the reference picture list (RPL) is in the picture header (PH), for all strips of the picture, the number of active entries of reference picture list 0 and / or 1 is required to be the same.

[0376] a. In one alternative, for i in the range from 0 to 1 (including 0 and 1), the value of NumRefIdxActive[i] is set to be equal to num_ref_idx_default_active_minus1[i] + 1.

[0377] b. In another alternative, when the RPL is in the PH, another syntax element (e.g., a flag) can be further signaled in the PH to indicate how to derive the number of active entries of the reference picture list.

[0378] i. In one example, according to the "another syntax element", the number of active entries can be derived with or without additional signaling.

[0379] ii. In one example, the following can be applied:

[0380] 1. If the flag is equal to X (e.g., X = 1), then for i in the range from 0 to 1 (including 0 and 1), NumRefIdxActive[i] is set to be equal to num_ref_idx_default_active_minus1[i] + 1.

[0381] 2. Otherwise (this flag is equal to (1 - X)), one or more than two syntax elements are signaled in the PH (e.g., num_active_entries_minus1[i] for i equal to 0 and 1) and NumRefIdxActive[i] is set to be equal to num_active_entries_minus1[i] + 1.

[0382] 3. Alternatively, predictive coding can be applied, where the difference in the number of active entries can be signaled, e.g., the difference in the number of active entries is signaled relative to the default value signaled by num_ref_idx_default_active_minus1[i] signaled in the PPS.

[0383] iii. Alternatively, only when the flag is used to control both, two flags can be signaled in the PH to respectively control whether the information on the number of active entries is signaled in the PH of reference picture lists 0 and 1 in a similar manner as above.

[0384] c. In the above example, it is assumed that B slices can be used. If all slices referring to the same PH or all slices between are P slices, then the number of active entries in reference picture list 0 can be derived only.

[0385] d. Alternatively, in addition, how to signal new syntax elements (e.g., "another syntax element" in bullet 29.b; "one or two more syntax elements" in bullet 30.b.ii.2 and bullet 30.b.iii) can depend on the allowed slice types referring to the PH.

[0386] i. In one example, new syntax elements can be signaled under the condition check that at least one slice is inter - coded or a P slice or a B slice; or whether all slices are intra - slices.

[0387] 31. It is proposed to use the following syntax in the PH syntax structure

[0388]

[0389] Change to the following:

[0390]

[0391] 32. The WP table syntax is changed from the following (where the yellow highlighted parts are new compared to the latest VVC draft text in JVET-Q0041-v2):

[0392]

[0393]

[0394] As follows (i.e., the same as JVET-Q0041-v2):

[0395]

[0396]

[0397] 33. In one example, the bullets 30, 31, 32 can be used together as a combined solution.

[0398] a. Alternatively, when signaling the WP table in the strip header (i.e., when wp_info_in_ph_flag is equal to 0), the changes in entry 30 above are not made (i.e., when the RPL information is in the SH, the number of active entries is still signaled in the SH), but all strips of the constrained picture should have the same value of NumRefIdxActive[0] and the same value of NumRefIdxActive[0], and the constrained WP syntax table is also changed to be the same as in the latest VVC draft text in JVET-Q0041-v2.

[0399] 34. For the above example, instead of directly encoding and decoding the number of weights to be signaled, the difference between the number of weights and a threshold K (e.g., K = 1) can be encoded in the bitstream.

[0400] a. In one example, num_l0_weights in the WP table should be encoded as num_l0_weights_minus1 (i.e., the number minus 1 is encoded), and the syntax table is changed accordingly.

[0401] b. Alternatively, predictive coding of the number of weights of list X can be used. For example, the difference between the number of weights of list 0 and list 1 can be encoded.

[0402] Cropping of ACT Residual

[0403] 35. When applying the inverse ACT to a block, the inverse ACT input at the decoder can be clipped based on (BitDepth + K), where BitDepth represents the bit depth of the block (e.g., the internal bit depth, which can be represented by bit_depth_minus8), and K is not equal to 0.

[0404] a. In one example, the clipping operation can be represented by f(a, b, x), where x represents the value to be clipped, a is the lower bound of x, and b is the upper bound of x.

[0405] i. In one example, f(a, b, x) can be equal to (x < a? a : (x > b? b : x)).

[0406] ii. In one example, a can be set to equal –(1 << (BitDepth + K)).

[0407] iii. In one example, b can be set to equal (1 << (BitDepth + K)) - 1.

[0408] b. In one example, K can be set to equal 1.

[0409] i. Alternatively, K can be set to equal 0.

[0410] c. In one example, the clipping based on (BitDepth + K) can be applied only to the second color component and / or the third color component.

[0411] 36. When applying the inverse ACT to a block, the input clipping of the inverse ACT may not be applied to the first color component or the luminance component.

[0412] 37. The output of the inverse ACT can be clipped based on the internal bit depth (i.e., BitDepth).

[0413] a. The clipping operation can represent the operation described in item 22.a, where K is set to 0.

[0414] b. The clipping may apply only to the first color component or the luminance component.

[0415] c. The clipping may apply only to the second color component and the third color component.

[0416] General Techniques (Items 20 - 21)

[0417] 38. In the above examples, S 1 、S 2 、l, h, m, n, and / or k are integers and can depend on

[0418] a. Messages signaled in DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header / Largest coding unit (LCU) / Coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit

[0419] b. Positions of CU / PU / TU / block / video coding unit

[0420] c. Coding modes of blocks containing samples along edges

[0421] d. Transformation matrices applied to blocks containing samples along edges

[0422] e. Block size / block shape of the current block and / or its neighboring blocks

[0423] f. Indication of color format (e.g., 4:2:0, 4:4:4, RGB, or YUV)

[0424] g. Coding tree structure (e.g., binary tree or single tree)

[0425] h. Strip / slice group type and / or picture type

[0426] i. Color components (e.g., may apply only to Cb or Cr)

[0427] j. Temporal layer ID

[0428] k. Profile / level / layer of the standard

[0429] l. Alternatively, S 1 、S 2 、l、h、m、n, and / or k can be signaled to the decoder

[0430] 39. The proposed method above can be applied under certain conditions

[0431] a. In one example, the condition is that the color format is 4:2:0 and / or 4:2:2

[0432] b. In one example, an indication of the use of the method above can be signaled at the sequence / picture / strip / slice / tile / video region level (e.g., SPS / PPS / picture header / strip header)

[0433] c. In one example, the use of the method above can depend on

[0434] i. Video content (e.g., screen content or natural content)

[0435] ii. Messages signaled in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Slice Group Header / Largest Coding Unit (LCU) / Coding Unit (CU) / LCU Row / LCU Group / TU / PU Block / Video Coding Unit

[0436] iii. Positions of CU / PU / TU / Block / Video Coding Unit

[0437] iv. Coding mode of blocks containing samples along the edge

[0438] v. Transformation matrix applied to blocks containing samples along the edge

[0439] vi. Block size of the current block and / or its neighboring blocks

[0440] vii. Block shape of the current block and / or its neighboring blocks

[0441] viii. Indication of color format (e.g., 4:2:0, 4:4:4, RGB, or YUV)

[0442] ix. Coding tree structure (e.g., binary tree or single tree)

[0443] x. Strip / Slice Group type and / or Picture type

[0444] xi. Color component (e.g., may apply only to Cb or Cr)

[0445] xii. Temporal layer ID

[0446] xiii. Standard profile / level / layer

[0447] xiv. Alternatively, m and / or n may be signaled to the decoder.

[0448] 5 Examples

[0449] These examples are based on JVET-P2001-vE. Newly added text is highlighted with highlighting. Deleted text is marked with italic text.

[0450] 5.1 Example #1

[0451] This example relates to the interaction between ACT and BDPCM modes.

[0452]

[0453]

[0454] 5.2 Example #2

[0455] This embodiment relates to the interaction between the ACT and BDPCM modes.

[0456] When intra_bdpcm_chroma_flag is equal to 1, it is specified that BDPCM is applied to the current chroma coding / decoding block at position (x0, y0), i.e., the transform is skipped, and the intra-chroma prediction mode is specified by intra_bdpcm_chroma_dir_flag. When intra_bdpcm_chroma_flag is equal to 0, it is specified that BDPCM is not applied to the current chroma coding / decoding block at position (x0, y0).

[0457] When intra_bdpcm_chroma_flag does not exist it is inferred to be equal to 0.

[0458]

[0459] For x = x0..x0 + cbWidth - 1, y = y0..y0 + cbHeight - 1, and cIdx = 1..2, set the variable BdpcmFlag[x][y][cIdx] to be equal to intra_bdpcm_chroma_flag.

[0460] When intra_bdpcm_chroma_dir_flag is equal to 0, it indicates that the BDPCM prediction direction is horizontal. When intra_bdpcm_chroma_dir_flag is equal to 1, it indicates that the BDPCM prediction direction is vertical.

[0461] For x = x0..x0 + cbWidth - 1, y = y0..y0 + cbHeight - 1, and cIdx = 1..2, set the variable BdpcmDir[x][y][cIdx] to be equal to intra_bdpcm_chroma_dir_flag.

[0462] 5.3 Embodiment #3

[0463] This embodiment relates to QP setting.

[0464] 8.7.3 Scaling Process of Transform Coefficients

[0465] The inputs of this process are:

[0466] – The luminance position (xTbY, yTbY), which specifies the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture

[0467] slice,

[0468] – The variable nTbW, which specifies the transform block width,

[0469] – The variable nTbH, which specifies the transform block height,

[0470] – The variable predMode, which specifies the prediction mode of the coding / decoding unit

[0471] – The variable cIdx, which specifies the color component of the current block,

[0472] The output of this process is an (nTbW)×(nTbH) array d of scaled transform coefficients with elements d[x][y].

[0473] …

[0474] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0475] – If transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0, the following applies:

[0476] qP = qP - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1133)

[0477]

[0478] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1)? 1 : 0 (1134)

[0479] bdShift = BitDepth + rectNonTsFlag + (1135)

[0480] ((Log2(nTbW) + Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag

[0481] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1), the following applies:

[0482] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0)(1136)

[0483]

[0484] rectNonTsFlag = 0 (1137)

[0485] bdShift = 10 (1138)

[0486] …

[0487] 5.4 Example #4

[0488] 8.7.1 Derivation Process of Quantization Parameters

[0489] …

[0490] The chrominance quantization parameters (Qp′Cb and Qp′Cr) for the Cb and Cr components and the chrominance quantization parameter Qp′CbCr for joint Cb - Cr encoding and decoding are derived as follows:

[0491] Qp′Cb = Clip3(-QpBdOffset, 63, qPCb + pps_cb_qp_offset + slice_cb_qp_offset + CuQpOffsetCb)

[0492] +QpBdOffset (1122)

[0493] Qp′Cr = Clip3(-QpBdOffset, 63, qPCr + pps_cr_qp_offset + slice_cr_qp_offset +

[0494] CuQpOffsetCr)+QpBdOffset (1123)

[0495] Qp′CbCr = Clip3(-QpBdOffset, 63, qPCbCr + pps_joint_cbcr_qp_offset + slice_joint_cbcr_qp_offset + CuQpOffsetCbCr)+QpBdOffset (1124) 5.5 Example #5

[0496] 7.3.9.5 Coding and Decoding Unit Syntax

[0497]

[0498]

[0499] 5.6 Example #6

[0500] 7.3.9.5 Coding and Decoding Unit Syntax

[0501]

[0502] 5.7 Example #7

[0503] 8.7.3 Scaling Process of Transform Coefficients

[0504] The inputs to this process are as follows:

[0505] – The luminance position (xTbY, yTbY), which specifies the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture,

[0506] – The variable nTbW, which specifies the transform block width,

[0507] – The variable nTbH, which specifies the transform block height,

[0508] – The variable predMode, which specifies the prediction mode of the coding / decoding unit

[0509] – The variable cIdx, which specifies the color component of the current block,

[0510] The output of this process is an (nTbW) × (nTbH) array d of scaled transform coefficients with elements d[x][y].

[0511] The quantization parameter qP is derived as follows:

[0512] – If cIdx is equal to 0, then the following applies:

[0513]

[0514] qP = Qp′ Y (1129)

[0515] – Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, then the following applies:

[0516]

[0517] qP = Qp′ CbCr (1130)

[0518] – Otherwise, if cIdx is equal to 1, then the following applies:

[0519]

[0520] qP = Qp′ Cb (1131)

[0521] – Otherwise (cIdx is equal to 2), the following applies:

[0522]

[0523] qP = Qp′ Cr (1132)

[0524] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0525] – If transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0, the following applies:

[0526] qP = qP - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1133)

[0527]

[0528] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1)? 1 : 0 (1134)

[0529] bdShift = BitDepth + rectNonTsFlag + (1135)

[0530] ((Log2(nTbW) + Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag

[0531] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1), the following applies:

[0532] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1136)

[0533]

[0534] rectNonTsFlag = 0 (1137)

[0535] bdShift = 10 (1138)

[0536] The variable bdOffset is derived as follows:

[0537] bdOffset = (1 << bdShift) >> 1 (1139)

[0538] Specify the list levelScale[][] as levelScale[j][k] = {{40, 45, 51, 57, 64, 72}, {57, 64, 72, 80, 90, 102}} where j = 0..1 and k = 0..5.

[0539] …

[0540] 5.8 Embodiment #8

[0541] 8.7.3 Scaling Process of Transformation Coefficients

[0542] The inputs to this process are as follows:

[0543] – The luminance position (xTbY, yTbY), which specifies the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture,

[0544] – The variable nTbW, which specifies the transform block width,

[0545] – The variable nTbH, which specifies the transform block height,

[0546] – The variable predMode, which specifies the prediction mode of the coding / decoding unit

[0547] – The variable cIdx, which specifies the color component of the current block,

[0548] The output of this process is an (nTbW) × (nTbH) array d of scaled transformation coefficients with elements d[x][y].

[0549] – If cIdx is equal to 0, then the following applies:

[0550] qP = Qp′ Y (1129)

[0551]

[0552] – Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, then the following applies:

[0553] –

[0554]

[0555] –

[0556] qP = Qp′ CbCr (1130)

[0557] – Otherwise, if cIdx is equal to 1, then the following applies:

[0558] –

[0559]

[0560] –

[0561] qP = Qp' Cb (1131)

[0562] – Otherwise (cIdx equals 2), the following applies:

[0563] –

[0564]

[0565] –

[0566] qP = Qp' Cr (1132)

[0567] The quantization parameter qP is modified and the variables rectNonTsFlag and bdShift are derived as follows:

[0568] – If transform_skip_flag[xTbY][yTbY][cIdx] equals 0, the following applies:

[0569] qP = qP - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1133)

[0570] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1)? 1 : 0 (1134)

[0571] bdShift = BitDepth + rectNonTsFlag + (1135)

[0572] ((Log2(nTbW) + Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag

[0573] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] equals 1), the following applies:

[0574] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1136)

[0575]

[0576] rectNonTsFlag = 0 (1137)

[0577] bdShift = 10 (1138)

[0578] The variable bdOffset is derived as follows:

[0579] bdOffset = (1 << bdShift) >> 1 (1139)

[0580] Specify the list levelScale[][] as levelScale[j][k] = {{40, 45, 51, 57, 64, 72}, {57, 64, 72, 80, 90, 102}}, where j = 0..1 and k = 0..5.

[0581] …

[0582] 5.9 Example #9

[0583] 7.4.3.4 Picture Parameter Set RBSP Semantics

[0584]

[0585]

[0586]

[0587]

[0588]

[0589] 8.7.3 Scaling Process of Transform Coefficients

[0590] The inputs to this process are:

[0591] – The luminance position (xTbY, yTbY), which specifies the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture,

[0592] – The variable nTbW, which specifies the transform block width,

[0593] – The variable nTbH, which specifies the transform block height,

[0594] – The variable predMode, which specifies the prediction mode of the coding / decoding unit

[0595] – The variable cIdx, which specifies the color component of the current block,

[0596] The output of this process is an (nTbW) × (nTbH) array d of scaled transform coefficients with elements d[x][y].

[0597] The quantization parameter qP is derived as follows:

[0598] – If cIdx is equal to 0, then the following applies:

[0599]

[0600] qP = Qp' Y (1129)

[0601] – Otherwise, if TuCResMode[xTbY][yTbY] equals 2, the following applies:

[0602]

[0603] qP = Qp' CbCr (1130)

[0604] – Otherwise, if cIdx equals 1, the following applies:

[0605]

[0606] qP = Qp' Cb (1131)

[0607] – Otherwise (cIdx equals 2), the following applies:

[0608]

[0609] qP = Qp' Cr (1132)

[0610] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0611] – If transform_skip_flag[xTbY][yTbY][cIdx] equals 0, the following applies:

[0612] qP = qP - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1133)

[0613]

[0614] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1)? 1 : 0 (1134)

[0615] bdShift = BitDepth + rectNonTsFlag + (1135)

[0616] ((Log2(nTbW) + Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag

[0617] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] equals 1), the following applies:

[0618] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1136)

[0619]

[0620] rectNonTsFlag = 0 (1137)

[0621] bdShift = 10 (1138)

[0622] The variable bdOffset is derived as follows:

[0623] bdOffset = (1 << bdShift) >> 1 (1139)

[0624] Specify the list levelScale[][] as levelScale[j][k] = {{40, 45, 51, 57, 64, 72}, {57, 64, 72, 80, 90, 102}} where j = 0..1 and k = 0..5.

[0625] …

[0626] 5.10 Example #10

[0627] 7.4.3.4 Picture Parameter Set RBSP Semantics

[0628]

[0629]

[0630]

[0631]

[0632]

[0633] 8.7.3 Scaling Process of Transform Coefficients

[0634] The inputs to this process are:

[0635] – Luminance position (xTbY, yTbY), specifying the top-left sample of the current luminance transform block relative to the top-left luminance sample of the current picture,

[0636] – Variable nTbW, specifying the transform block width,

[0637] – The variable nTbH specifies the transform block height.

[0638] – The variable predMode specifies the prediction mode of the coding / decoding unit.

[0639] – The variable cIdx specifies the color component of the current block.

[0640] The output of this process is an (nTbW)×(nTbH) array d of scaled transform coefficients with elements d[x][y].

[0641] – If cIdx is equal to 0, the following applies:

[0642] qP = Qp′ Y (1129)

[0643]

[0644] – Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, the following applies:

[0645] –

[0646]

[0647] –

[0648] qP = Qp′ CbCr (1130)

[0649] – Otherwise, if cIdx is equal to 1, the following applies:

[0650] –

[0651]

[0652] –

[0653] qP = Qp′ Cb (1131)

[0654] – Otherwise (cIdx is equal to 2), the following applies:

[0655] –

[0656]

[0657] –

[0658] qP = Qp′ Cr(1132)

[0659] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0660] – If transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0, the following applies:

[0661] qP = qP - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1133)

[0662] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1)? 1 : 0 (1134)

[0663] bdShift = BitDepth + rectNonTsFlag + (1135)

[0664] ((Log2(nTbW) + Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag

[0665] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1), the following applies:

[0666] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5 : 0) (1136)

[0667]

[0668] rectNonTsFlag = 0 (1137)

[0669] bdShift = 10 (1138)

[0670] The variable bdOffset is derived as follows:

[0671] bdOffset = (1 << bdShift) >> 1 (1139)

[0672] Specify the list levelScale[][] as levelScale[j][k] = {{40, 45, 51, 57, 64, 72}, {57, 64, 72, 80, 90, 102}} where j = 0..1 and k = 0..5.

[0673] …

[0674] 5.11 Example #11

[0675] Syntax change in PPS:

[0676]

[0677] Syntax change in strip header:

[0678]

[0679] When cu_act_enabled_flag is 1, pps_act_y_qp_offset, pps_act_cb_qp_offset, pps_act_cr_qp_offset, and pps_act_cbcr_qp_offset specify the offsets of quantization parameters Qp′Y, Qp′Cb, Qp′Cr, and Qp′CbCr, respectively. The values of pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset shall be in the range of -12 to +12, inclusive of the endpoints. When not present, the values of pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset are inferred to be equal to 0.

[0680] pps_slice_act_qp_offsets_present_flag being equal to 1 specifies that slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset are present in the strip header. pps_slice_act_qp_offsets_present_flag being equal to 0 indicates that slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset are not present in the strip header. When not present, the value of pps_slice_act_qp_offsets_present_flag is inferred to be equal to 0.

[0681] When determining the values of the Qp'Y, Qp'Cb, Qp'Cr, and Qp'CbCr quantization parameters, slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset specify the offsets. The values of slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset shall be in the range of -12 to +12, inclusive of the endpoints. When absent, these values are inferred to be equal to 0. 8.7.1 Derivation process of quantization parameters

[0682] …

[0683] The variable Qp Y is derived as follows:

[0684]

[0685] The luma quantization parameter Qp' Y is derived as follows:

[0686] Qp′ Y = Qp Y + QpBdOffset (1117)

[0687] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the following applies:

[0688] – When treeType is equal to DUAL_TREE_CHROMA, set the variable Qp Y equal to the luma quantization parameter Qp of the luma coding unit covering the luma position (xCb + cbWidth / 2, yCb + cbHeight / 2) Y .

[0689] – The derivation of the variables qP Cb , qP Cr and qP CbCr is as follows:

[0690] qP Chroma = Clip3(-QpBdOffset, 63, Qp Y ) (1118)

[0691] qP Cb = ChromaQpTable[0][qP Chroma(1119)

[0692] qP Cr = ChromaQpTable[1][qP Chroma (1120)

[0693] qP CbCr = ChromaQpTable[2][qP Chroma (1121)

[0694] The chroma quantization parameters Qp'Cb and Qp'Cr for the Cb and Cr components, and the chroma quantization parameter Qp'CbCr for joint Cb-Cr encoding / decoding are derived as follows:

[0695]

[0696]

[0697]

[0698] 8.7.3 Scaling Process of Transform Coefficients

[0699] …

[0700]

[0701] The quantization parameter qP is derived as follows:

[0702] – If cIdx is equal to 0, the following applies:

[0703] qP = Qp' Y (1129)

[0704] – Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, the following applies:

[0705] qP = Qp' CbCr (1130)

[0706] – Otherwise, if cIdx is equal to 1, the following applies:

[0707] qP = Qp' Cb (1131)

[0708] – Otherwise (cIdx is equal to 2), the following applies:

[0709] qP = Qp' Cr (1132)

[0710]

[0711] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0712] – If transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0, the following applies:

[0713]

[0714] rectNonTsFlag = (((Log2(nTbW)+Log2(nTbH)) & 1) == 1)? 1:0 (1134)

[0715] bdShift = BitDepth + rectNonTsFlag + (1135)

[0716] ((Log2(nTbW)+Log2(nTbH)) / 2) - 5 + pic_dep_quant_enabled_flag

[0717] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1), the following applies:

[0718] qP = Max(QpPrimeTsMin, qP) - (cu_act_enabled_flag[xTbY][yTbY]? 5:0) (1136)

[0719]

[0720] 5.12 Example #12

[0721] Syntax changes in PPS:

[0722]

[0723]

[0724] Syntax changes in slice header:

[0725]

[0726] When cu_act_enabled_flag is 1, pps_act_y_qp_offset, pps_act_cb_qp_offset, pps_act_cr_qp_offset, and pps_act_cbcr_qp_offset specify the offsets of quantization parameters Qp′Y, Qp′Cb, Qp′Cr, and Qp′CbCr, respectively. The values of pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset shall be in the range of -12 to +12, inclusive of the endpoints. When not present, the values of pps_act_y_qp_offset, pps_act_cb_qp_offset, and pps_act_cr_qp_offset are inferred to be equal to 0.

[0727] pps_slice_act_qp_offsets_present_flag being equal to 1 specifies that slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset are present in the slice header. pps_slice_act_qp_offsets_present_flag being equal to 0 indicates that slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset are not present in the slice header. When not present, the value of pps_slice_act_qp_offsets_present_flag is inferred to be equal to 0.

[0728] slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset specify the offsets when determining the values of Qp'Y, Qp'Cb, Qp'Cr, and Qp'CbCr quantization parameters. The values of slice_act_y_qp_offset, slice_act_cb_qp_offset, slice_act_cr_qp_offset, and slice_act_cbcr_qp_offset shall be in the range of -12 to +12, inclusive of the endpoints. When not present, these values are inferred to be equal to 0.

[0729] 8.7.1 Derivation Process of Quantization Parameters

[0730] …

[0731] The variable Qp Y is derived as follows:

[0732] Qp Y = ((qP Y_PRED + CuQpDeltaVal

[0733] + 64 + 2 * QpBdOffset) % (64 + QpBdOffset)) - QpBdOffset (1116)

[0734] The luminance quantization parameter Qp' Y is derived as follows:

[0735]

[0736] When ChromaArrayType is not equal to 0 and treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the following applies:

[0737] – When treeType is equal to DUAL_TREE_CHROMA, set the variable Qp Y equal to the luminance quantization parameter Qp of the luminance coding / decoding unit covering the luminance position (xCb + cbWidth / 2, yCb + cbHeight / 2) Y .

[0738] – The variables qP Cb , qP Cr and qP CbCr are derived as follows:

[0739] qP Chroma = Clip3(-QpBdOffset, 63, Qp Y ) (1118)

[0740] qP Cb = ChromaQpTable[0][qP Chroma (1119)

[0741] qP Cr = ChromaQpTable[1][qP Chroma (1120)

[0742] qP CbCr = ChromaQpTable[2][qP Chroma (1121)

[0743] – Chrominance quantization parameter Qp′ for Cb and Cr components Cb and Qp′ Cr as well as the chrominance quantization parameter Qp′ for joint Cb - Cr encoding and decoding CbCr is derived as follows:

[0744] –

[0745]

[0746]

[0747]

[0748] –

[0749] Qp′ Cb = Clip3(-QpBdOffset, 63, qP Cb + pps_cb_qp_offset + slice_cb_qp_offset + CuQpOffset Cb )

[0750] + QpBdOffset (1122)

[0751] Qp′ Cr = Clip3(-QpBdOffset, 63, qP Cr + pps_cr_qp_offset + slice_cr_qp_offset + CuQpOffset Cr )

[0752] + QpBdOffset (1123)

[0753] Qp′ CbCr = Clip3(-QpBdOffset, 63, qP CbCr + pps_joint_cbcr_qp_offset + slice_joint_cbcr_qp_offset + CuQpOffset CbCr ) + QpBdOffset

[0754] 8.7.3 Scaling process of transform coefficients

[0755] …

[0756]

[0757] The quantization parameter qP is derived as follows:

[0758] – If cIdx is equal to 0, the following applies:

[0759] qP = Qp′ Y (1129)

[0760] – Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, the following applies:

[0761] qP = Qp′ CbCr (1130)

[0762] – Otherwise, if cIdx is equal to 1, the following applies:

[0763] qP = Qp′ Cb (1131)

[0764] – Otherwise (cIdx is equal to 2), the following applies:

[0765] qP = Qp′ Cr (1132)

[0766]

[0767] The quantization parameter qP is modified, and the variables rectNonTsFlag and bdShift are derived as follows:

[0768] – If transform_skip_flag[xTbY][yTbY][cIdx] is equal to 0, the following applies:

[0769]

[0770] rectNonTsFlag = (((Log2(nTbW)+Log2(nTbH))&1)==1)?1:0 (1134)

[0771] bdShift = BitDepth+rectNonTsFlag + (1135)

[0772] ((Log2(nTbW)+Log2(nTbH)) / 2)-5+pic_dep_quant_enabled_flag

[0773] – Otherwise (transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1), the following applies:

[0774] qP = Max(QpPrimeTsMin,qP)-(cu_act_enabled_flag[xTbY][yTbY]?5:0) (1136)

[0775]

[0776] 5.13 Example #13

[0777] 8.7.4.6 Residual Modification Process for Blocks Using Color Space Conversion

[0778] …

[0779] Residual sample point r Y 、r Cb and r Cr 's (nTbW) x (nTbH) array is modified as follows:

[0780]

[0781]

[0782]

[0783]

[0784]

[0785]

[0786] r Cr [x][y]± = tmp + - r Cr [x][y] (1194)

[0787] 5.14 Example #14

[0788] 8.7.4.6 Residual Modification Process for Blocks Using Color Space Conversion

[0789] …

[0790] Residual sample point r Y 、r Cb and r Cr 's (nTbW) x (nTbH) array is modified as follows:

[0791]

[0792]

[0793]

[0794]

[0795] r Y [x][y] = r Y [x][y] tmp+r Cb [x][y] (1192)

[0796]

[0797] 5.15 Example #15

[0798] 8.7.4.6 Residual modification process for blocks using color space conversion

[0799] …

[0800] Residual sample r Y 、r Cb and r Cr The (nTbW) x (nTbH) array of, and is modified as follows:

[0801]

[0802]

[0803]

[0804]

[0805]

[0806] Figure 7 is a block diagram showing an example video processing system 700 in which various techniques disclosed herein can be implemented. Various implementations may include some or all components of video processing system 700. System 700 may include an input 702 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be received in a compressed or encoded format. Input 702 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0807] System 700 may include a codec component 704 that may implement various encoding or decoding methods described in this document. The codec component 704 may reduce the average bit rate of a video from the input 702 to the output of the codec component 704 to produce a coded representation of the video. Coding and decoding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 704 may be stored or transmitted via a communication connection represented by component 706. The stored or communicated bitstream (or coded) representation of the video received at input 702 may be used by component 708 to generate pixel values or a displayable video for transmission to a display interface 710. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are referred to as "coding and decoding" operations or tools, it should be understood that coding and decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the coding and decoding results will be performed by the decoder.

[0808] Examples of a peripheral bus interface or a display interface may include a universal serial bus (USB), or a high definition multimedia interface (HDMI), or a Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.

[0809] Figure 8 is a block diagram of a video processing apparatus 800. The apparatus 800 may be used to implement one or more methods described herein. The apparatus 800 may be implemented in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 800 may include one or more processors 802, one or more memories 804, and video processing hardware 806 (also known as video processing circuitry). The (multiple) processors 802 may be configured to implement one or more methods described in this document (e.g., in Figure 5 ). The memory (multiple memories) 804 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 806 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the hardware 806 may be partially or fully located within the processor 802 (e.g., a graphics processor).

[0810] Figure 9 is a block diagram showing an example of a video coding and decoding system 100 that can utilize the technology of the present disclosure. As Figure 9 shown, the video coding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, and the source device 110 may be referred to as a video coding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the destination device 120 may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0811] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 120 via the I / O interface 116 through a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0812] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0813] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120 that may be configured to interface with an external display device.

[0814] The video encoder 114 and the video decoder 124 may operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0815] Figure 10 is a block diagram showing an example of a video encoder 200, and the video encoder 200 may beFigure 9 The video encoder 114 in the illustrated video coding system 100.

[0816] The video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In Figure 10 an example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0817] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206), a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0818] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0819] In addition, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but are shown separately in Figure 10 an example for purposes of explanation.

[0820] The splitting unit 201 may split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0821] The mode selection unit 203 may select one of the coding modes (intra or inter) based on, for example, an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP), where the prediction is based on inter prediction signaling and intra prediction signaling. In the case of inter prediction, the mode selection unit 203 may also select the resolution of the motion vector of the block (e.g., sub-pixel or integer pixel accuracy).

[0822] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0823] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.

[0824] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search for a reference picture in list 0 or list 1 of reference video blocks for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1, which includes the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0825] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block, the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1, which includes the reference video block and a motion vector indicating a spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0826] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder.

[0827] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0828] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, which indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0829] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0830] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector predication (AMVP) and Merge mode signaling.

[0831] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0832] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0833] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0834] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0835] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0836] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0837] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce video blocking artifacts in the video block.

[0838] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0839] Figure 11 is a block diagram showing an example of a video decoder 300, and the video decoder 300 can be Figure 9 the video decoder 114 in the video codec system 100 shown.

[0840] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 11 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.

[0841] In Figure 11 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described for the video encoder 200 ( Figure 10 ).

[0842] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and merge mode signaling.

[0843] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.

[0844] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during encoding of a video block to compute interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information and may use the interpolation filter to generate a prediction block.

[0845] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of an encoded video sequence, partitioning information that describes how each macroblock of a picture describing the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0846] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in a bitstream. The inverse quantization unit 304 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301, i.e., dequantizes. The inverse transform unit 305 applies an inverse transform.

[0847] The reconstruction unit 306 may add a residual block to a corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303 to form a decoded block. If needed, deblocking filtering may also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra prediction and also to produce decoded video for presentation on a display device.

[0848] Figures 12 - 13 An example method for implementing the above technical solution in an embodiment such as Figures 7 - 11 is shown.

[0849] Figure 12 A flowchart of an example method 1200 for video processing is shown, including, at operation 1210, performing a conversion between a current video block of a video and a bitstream of the video, encoding and decoding the current video block using an Adaptive Color Transform (ACT) mode, the conversion including applying an inverse ACT transform to the current video block according to a rule that specifies that a clipping operation based on the bit depth of the current video block is applied to the input of the inverse ACT transform.

[0850] Figure 13A flowchart showing an example method 1300 for video processing is presented, including, at operation 1310, performing a conversion between a video including a current video block and a bitstream of the video, wherein, during the conversion, a weighted prediction of the current video block is determined using weights, and wherein the weights are included in the bitstream as information indicating a difference between the number of weights and a threshold K, where K is an integer.

[0851] Next, a list of preferred solutions for some embodiments is provided.

[0852] 1. A video processing method, including: performing a conversion between a current video block of a video and a bitstream of the video, wherein the current video block is encoded and decoded using an Adaptive Color Transform (ACT) mode, and wherein the conversion includes applying an inverse ACT transform to the current video block according to a rule, and wherein the rule stipulates that a first clipping operation based on the bit depth of the current video block (denoted as BitDepth) is applied to the input of the inverse ACT transform.

[0853] 2. The method according to solution 1, wherein the first clipping operation is based on the sum of the bit depth and K, and wherein K is a non-zero integer.

[0854] 3. The method according to solution 2, wherein K = 1.

[0855] 4. The method according to solution 1, wherein the first clipping operation is based on the sum of the bit depth and K, and wherein K = 0.

[0856] 5. The method according to any one of solutions 1 to 4, wherein the first clipping operation is represented as f(a, b, x), where x is the value to be clipped, a is the lower bound of the first clipping operation, b is the upper bound of the first clipping operation, and wherein a and b are integers.

[0857] 6. The method according to solution 5, wherein f(a, b, x) = (x < a? a : (x > b? b : x)).

[0858] 7. The method according to solution 5 or 6, wherein a = -(1 << (BitDepth + K)).

[0859] 8. The method according to any one of solutions 5 to 7, wherein b = (1 << (BitDepth + K)) - 1.

[0860] 9. The method according to solution 1, wherein the input is associated with at least one sample point of the Cb color component or the Cr color component of the video.

[0861] 10. The method according to solution 1, wherein the input is associated with sample points excluding sample points of the Y color component or the luminance component of the video.

[0862] 11. The method according to any one of Solutions 1 to 10, wherein the bit depth is the input bit depth or the internal bit depth.

[0863] 12. The method according to any one of Solutions 1 to 11, wherein the rule further stipulates that a second clipping operation based on the bit depth of the current video block is applied to the output of the inverse ACT transform.

[0864] 13. The method according to Solution 12, wherein the output is associated with at least one sample of the Cb color component or the Cr color component of the video.

[0865] 14. The method according to Solution 12, wherein the output is associated with at least one sample of the Y color component of the video or the luminance component of the video.

[0866] 15. A video processing method, comprising: performing a conversion between a video including a current video block and a bitstream of the video, wherein, during the conversion, a weighted prediction of the current video block is determined using weights, and wherein the weights are included in the bitstream as information indicating a difference between the number of weights and a threshold K, where K is an integer.

[0867] 16. The method according to Solution 15, wherein K = 1.

[0868] 17. The method according to Solution 15, wherein the number of weights associated with reference picture list 0 minus one is included in the bitstream.

[0869] 18. The method according to Solution 15, wherein the predictive coding of the number of weights associated with reference picture list X is included in the bitstream.

[0870] 19. The method according to Solution 18, wherein the predictive coding includes the difference between the number of weights associated with reference picture list X and the number of weights associated with reference picture list (1 - X).

[0871] 20. The method according to Solution 18 or 19, wherein X = 0 or X = 1.

[0872] 21. The method according to any one of Solutions 1 to 20, wherein the conversion includes decoding the video from the bitstream.

[0873] 22. The method according to any one of Solutions 1 to 20, wherein the conversion includes encoding the video into the bitstream.

[0874] 23. A method according to any one of Solutions 1 to 20, wherein the conversion includes generating a bitstream from a video, and wherein the method further includes storing the bitstream in a non-transitory computer-readable recording medium.

[0875] 24. A method of storing a bitstream representing a video in a computer-readable recording medium, comprising: generating a bitstream from a video according to the method of any one or more of Solutions 1 to 20; and writing the bitstream to a computer-readable recording medium.

[0876] 25. A video processing apparatus, comprising a processor configured to implement the method of any one or more of Solutions 1 to 24.

[0877] 26. A computer-readable medium storing instructions that, when executed, cause a processor to implement the method of one or more of Solutions 1 to 24.

[0878] 27. A computer-readable medium storing a bitstream generated according to any one or more of Solutions 1 to 24.

[0879] 28. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method of any one or more of Solutions 1 to 24.

[0880] 29. A bitstream generated according to the method or system disclosed in this document.

[0881] Next, another list of preferred solutions of some embodiments is provided.

[0882] P1. A video processing method, comprising: determining whether to enable chrominance block differential pulse code modulation (BDPCM) mode for a video block of a video based on whether an adaptive color transform (ACT) mode and / or a luminance BDPCM mode of the video block is enabled; and performing a conversion between the video block and a bitstream representation of the video according to the determination.

[0883] P2. The method of Solution P1, wherein a signaling notification of a first value of a first flag associated with enabling chrominance BDPCM mode is determined based on a signaling notification of enabling the ACT mode for the video block and based on a signaling notification of a second value of a second flag associated with the use of the luminance BDPCM mode.

[0884] P3. The method of Solution P2, wherein, in response to the ACT mode being enabled and the second value of the second flag having a dummy value, the first value of the first flag has a dummy value.

[0885] P4. A method of Solution P2, wherein in response to the second value of the second flag having a true value, the first value of the first flag has a true value.

[0886] P5. A method of Solution P1, wherein the signaling of the ACT mode for a video block is conditionally based on the same BDPCM prediction direction for the luminance samples and chrominance samples of the video block.

[0887] P6. A method of Solution P5, wherein the signaling of the ACT mode is indicated after the signaling of the chrominance BDPCM mode and the luminance BDPCM mode.

[0888] P7. A method of Solution P1, wherein in response to the use of the ACT mode being enabled, a first value indicating a first prediction direction of the chrominance BDPCM mode is derived from a second value indicating a second prediction direction of the luminance BDPCM mode.

[0889] P8. A method of Solution P7, wherein the first value indicating the first prediction direction of the chrominance BDPCM mode is the same as the second value indicating the second prediction direction of the luminance BDPCM mode.

[0890] P9. A method of Solution P8, wherein the first prediction direction of the chrominance BDPCM mode and the second prediction direction of the luminance BDPCM mode are horizontal directions.

[0891] P10. A method of Solution P8, wherein the first prediction direction of the chrominance BDPCM mode and the second prediction direction of the luminance BDPCM mode are vertical directions.

[0892] P11. A method of Solution P1, wherein in response to the use of the ACT mode being disabled, the first value indicating the first prediction direction of the chrominance BDPCM mode is zero.

[0893] P12. A video processing method, comprising: determining whether to enable a block-based differential pulse code modulation (BDPCM) mode for a video block of a video based on whether to enable the use of a video block adaptive color transform (ACT) mode; and performing a conversion between the video block and a bitstream representation of the video according to the determination.

[0894] P13. A method of Solution P12, wherein in response to the ACT mode of the video block being enabled, the BDPCM mode of the video block is disabled.

[0895] P14. A method of Solution P13, wherein a first flag indicating the BDPCM mode is signaled after a second flag indicating the ACT mode.

[0896] P15. A method of solution P13, wherein a flag indicating the BDPCM mode is not signaled, and wherein the flag is determined to be a false value or zero.

[0897] P16. A method of solution P12, wherein in response to the BDPCM mode of a video block being enabled, the ACT mode of the video block is disabled.

[0898] P17. A method of solution P16, wherein a first flag indicating the BDPCM mode is signaled before a second flag indicating the ACT mode.

[0899] P18. A method of solution P16, wherein a flag indicating the ACT mode is not signaled, and wherein the flag is determined to be a false value or zero.

[0900] P19. A method of solution P12 to P18, wherein the BDPCM mode includes a luminance BDPCM mode and / or a chrominance BDPCM mode.

[0901] P20. A method of solution P1, wherein when the chrominance BDPCM mode and the luminance BDPCM mode are associated with different prediction modes, the ACT mode is applied.

[0902] P21. A method of solution P20, wherein a forward ACT mode is applied after the chrominance BDPCM mode or the luminance BDPCM mode.

[0903] P22. A method of solution P1 to P21, wherein in response to the ACT mode being enabled, the quantization parameter (QP) of a video block is clipped.

[0904] P23. A method of solution P22, wherein a clipping function for clipping the QP is defined as (l, h, x), where l is the lowest possible value of the input x and h is the highest possible value of the input x.

[0905] P24. A method of solution P23, wherein l is equal to 0.

[0906] P25. A method of solution P23, wherein h is equal to 63.

[0907] P26. A method of solution P22, wherein the QP of a video block is clipped after adjusting the QP for the ACT mode.

[0908] P27. A method of solution P23, wherein in response to applying transform skip to a video block, l is equal to the minimum allowable QP of the transform skip mode.

[0909] P28. The method of any one of Solutions P23 to P26, wherein l, h, m, n, and / or k are integers that depend on (i) a message signaled in a DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header / largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit, (ii) the position of the CU / PU / TU / block / video coding unit, (iii) the coding mode of a block containing samples along an edge, (iv) the transform matrix applied to a block containing samples along an edge, (v) the block dimension / block shape of the current block and / or its neighboring blocks, (vi) an indication of the color format (such as 4:2:0, 4:4:4, RGB, or YUV), (vii) the coding tree structure (such as a binary tree or a single tree), (viii) the strip / slice group type and / or the picture type, (ix) the color component (for example, it may only apply to Cb or Cr), (x) the temporal layer ID, or (xi) the profile / level / layer of the standard.

[0910] P29. The method of any one of Solutions P23 to P26, wherein l, h, m, n, and / or k are signaled to the decoder.

[0911] P30. The method of any one of Solutions P30, wherein the color format is 4:2:0 or 4:2:2.

[0912] P31. The method of any one of Solutions P1 to P30, wherein an indication of the ACT mode or the BDPCM mode or the chrominance BDPCM mode or the luminance BDPCM mode is signaled at the sequence / picture / strip / slice / tile or video region level.

[0913] P32. A video processing method, comprising: determining to use a joint chrominance residual coding (JCCR) tool on a video block of a video in which the adaptive color transform (ACT) mode is enabled; and performing a conversion between the video block and the bitstream representation of the video based on the determination, wherein the quantization parameter (QP) of the ACT mode is based on the mode of the JCCR tool.

[0914] P33. The method of Solution P32, wherein when the mode of the JCCR tool is determined to be 1, the QP is -5 or -6.

[0915] P34. The method of Solution P32, wherein when the mode of the JCCR tool is determined to be 3, the QP is -4 or -5.

[0916] P35. The method of any one of Solutions P1 to P34, wherein the conversion includes parsing and decoding the coded representation to generate video pixels.

[0917] P36. A method according to any one of solutions P1 to P34, wherein the conversion includes generating an encoded representation by encoding a video.

[0918] P37. A video decoding apparatus, including a processor configured to implement the method according to any one or more of solutions P1 to P36.

[0919] P38. A video encoding apparatus, including a processor configured to implement the method according to any one or more of solutions P1 to P36.

[0920] P39. A computer program product having computer code stored thereon, which code, when executed by a processor, causes the processor to implement the method according to any one or more of solutions P1 to P36.

[0921] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. The bitstream representation of the current video block may, for example, correspond to bits located at different positions or distributed at different positions within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on the transformed and encoded error residual values, and bits in the header and other fields in the bitstream may also be used.

[0922] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for running by a data processing apparatus or controlling the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to the hardware, the apparatus may also include code for creating a runtime environment for the computer program under discussion, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.

[0923] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0924] The processes and logical flows described in this document can be performed by one or more programmable processors running one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit).

[0925] Processors suitable for running a computer program include, for example, any one or more processors of general and special-purpose microprocessors, as well as any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, e.g., magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to or both receive and transfer data to and from the one or more mass storage devices. However, a computer does not necessarily require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0926] Although this patent document contains many details, these details should not be construed as limitations on any subject matter or the scope that may be claimed, but rather as descriptions of features of particular embodiments specific to a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0927] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0928] Only some implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing video data, comprising: performing a conversion between a current video block of a video and a bitstream of the video, wherein when the current video block is encoded and decoded using an Adaptive Color Transform (ACT) mode, the conversion includes applying an inverse ACT transform to the current video block according to a rule, wherein the rule specifies applying a clipping operation f(a, b, x) to the input of the inverse ACT transform, where f(a, b, x) = (x < a? a : (x > b? b : x)), x is the value to be clipped, a is the integer lower bound of the clipping operation, and b is the integer upper bound of the clipping operation, where a = -(1 << (BitDepth + K)), and BitDepth is the bit depth of the current video block, and where b = (1 << (BitDepth + K)) – 1, and where K is a non - zero integer.

2. The method according to claim 1, wherein K = 1.

3. The method according to claim 1 or 2, wherein the conversion includes decoding the video from the bitstream.

4. The method according to claim 1 or 2, wherein the conversion includes encoding the video into the bitstream.

5. The method according to claim 1, wherein, the input is associated with at least one sample of the Cb color component or the Cr color component of the video.

6. The method according to claim 1, wherein, the input is associated with samples excluding samples of the Y color component or the luminance component of the video.

7. The method according to claim 1, wherein, the bit depth is the input bit depth or the internal bit depth.

8. The method according to claim 1, wherein, the rule further specifies that a second clipping operation based on the bit depth of the current video block is applied to the output of the inverse ACT transform.

9. The method according to claim 8, wherein, the output is associated with at least one sample of the Cb color component or the Cr color component of the video.

10. The method according to claim 8, wherein, the output is associated with at least one sample of the Y color component or the luminance component of the video.

11. The method according to any one of claims 5 to 10, wherein the conversion includes generating the bitstream from the video, and wherein, the method further includes: storing the bitstream in a non - transitory computer - readable recording medium.

12. An apparatus for processing video data, comprising a processor and a non - transitory memory having instructions thereon, wherein when the instructions are executed by the processor, the processor is caused to: perform a conversion between a current video block of a video and a bitstream of the video, wherein when the current video block is encoded and decoded using an Adaptive Color Transform (ACT) mode, the conversion includes applying an inverse ACT transform to the current video block according to a rule, wherein the rule specifies applying a clipping operation f(a, b, x) to the input of the inverse ACT transform, where f(a, b, x) = (x < a? a : (x > b? b : x)), x is the value to be clamped, a is the integer lower bound of the clamping operation, and b is the integer upper bound of the clamping operation, where a = -(1 << (BitDepth + K)), and BitDepth is the bit depth of the current video block, and where b = (1 << (BitDepth + K)) – 1, and where K is a non-zero integer.

13. The apparatus according to claim 12, wherein K = 1.

14. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a current video block of a video and a bitstream of the video, where when the current video block is encoded or decoded using an Adaptive Color Transform (ACT) mode, the conversion includes applying an inverse ACT transform to the current video block according to a rule, where the rule specifies applying a clamping operation f(a, b, x) to the input of the inverse ACT transform, where f(a, b, x) = (x < a? a : (x > b? b : x)), x is the value to be clamped, a is the integer lower bound of the clamping operation, and b is the integer upper bound of the clamping operation, where a = -(1 << (BitDepth + K)), and BitDepth is the bit depth of the current video block, and where b = (1 << (BitDepth + K)) – 1, and where K is a non-zero integer.

15. The non-transitory computer-readable storage medium according to claim 14, wherein K = 1.

16. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing apparatus, wherein the method comprises: generating a bitstream of a current video block of the video, where when the current video block is encoded or decoded using an Adaptive Color Transform (ACT) mode, the generating of the bitstream includes applying an inverse ACT transform to the current video block according to a rule, where the rule specifies applying a clamping operation f(a, b, x) to the input of the inverse ACT transform, where f(a, b, x) = (x < a? a : (x > b? b : x)), x is the value to be clamped, a is the integer lower bound of the clamping operation, and b is the integer upper bound of the clamping operation, where a = -(1 << (BitDepth + K)), and BitDepth is the bit depth of the current video block, and where b = (1 << (BitDepth + K)) – 1, and where K is a non-zero integer.

17. The non-transitory computer-readable recording medium according to claim 16, wherein K = 1.

18. A method of storing a bitstream of a video, comprising: generating a bitstream of a current video block of the video, storing the bitstream in a non-transitory computer-readable recording medium, where when the current video block is encoded or decoded using an Adaptive Color Transform (ACT) mode, the generating of the bitstream includes applying an inverse ACT transform to the current video block according to a rule, wherein the rule provides for applying a clipping operation f(a, b, x) to the input of the inverse ACT transform, and where f(a, b, x) = (x < a? a : (x > b? b : x)), x is the value to be clipped, a is the integer lower bound of the clipping operation, and b is the integer upper bound of the clipping operation, where a = -(1 << (BitDepth + K)), and BitDepth is the bit depth of the current video block, where b = (1 << (BitDepth + K)) – 1, and where K is a non-zero integer.

19. A video processing apparatus, comprising a processor configured to implement the method according to any one of claims 5 to 10.

20. A computer-readable medium having instructions stored thereon, which when executed cause a processor to implement the method according to any one of claims 5 to 10.

21. A computer-readable medium storing the bitstream generated according to any one of claims 5 to 10.

22. A video processing apparatus for storing a bitstream, wherein, the video processing apparatus is configured to implement the method according to any one of claims 5 to 10.

Citation Information

Patent Citations

  • Color-space inverse transform both for lossy and lossless encoded video

    CN106105202A

  • Clipping for cross-component prediction and adaptive color transform for video coding

    US20160227224A1