Method, device and computer-readable medium for processing video data

By introducing transform skip mode and quantized residual block differential pulse coding modulation (QR-BDPCM) technology in video coding, the problem of low screen content coding efficiency in the existing technology is solved, and more efficient video coding is achieved.

CN113711613BActive Publication Date: 2025-09-12DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080029880.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-19
Filing Date
2020-04-17
Publication Date
2025-09-12
Estimated Expiration
2040-04-17

AI Technical Summary

Technical Problem

Existing video coding and decoding standards have the problem of low coding efficiency when processing screen content. Especially in intra-frame coding and decoding, it is difficult to effectively utilize repeated patterns in screen content, resulting in low coding efficiency.

Method used

A transform skip mode is adopted to encode the prediction error residual between the current video block and the reference video block without using transform in the bitstream representation, and include syntax elements of the maximum allowed size in the bitstream, allowing the current video block to be encoded without transform, combining quantized residual block differential pulse coding modulation (QR-BDPCM) technology and adaptive motion vector resolution (AMVR) to improve coding efficiency.

Benefits of technology

Improves the efficiency of video encoding, especially when processing screen content, reduces redundant data transmission, and improves encoding efficiency and compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113711613B_ABST
    Figure CN113711613B_ABST
Patent Text Reader

Abstract

Apparatus, systems, and methods for coefficient encoding and decoding in a transform skip mode are described. An exemplary method for video processing includes determining a maximum allowed size for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, within which maximum allowed size a current video block of the one or more video blocks is allowed to be encoded using a transform skip mode such that a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and including a syntax element in the bitstream representation indicating the maximum allowed size.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is intended to claim priority and the benefit, in a timely manner, under applicable patent laws and / or regulations under the Paris Convention to International Patent Application No. PCT / CN2019 / 083366, filed on April 19, 2019. The entire disclosure of the aforementioned application is incorporated herein by reference as a part of the disclosure of this application for all purposes prescribed by law. Technical Field

[0003] This document covers video and image encoding and decoding technologies, systems, and devices. Background Art

[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of networked user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] Devices, systems, and methods related to digital video codecs are described, and more particularly, devices, systems, and methods related to coefficient coding in transform skip modes for video codecs are described. The described methods can be applied to existing video codec standards (e.g., High Efficiency Video Coding (HEVC)) and future video codec standards (e.g., Versatile Video Coding (VVC)) or codecs.

[0006] In one exemplary aspect, a method for visual media processing is disclosed. The method includes: determining a maximum allowed size for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, within which a current video block of the one or more video blocks is allowed to be encoded using a transform skip mode such that a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and including a syntax element in the bitstream representation indicating the maximum allowed size.

[0007] In another exemplary aspect, a method for visual media processing is disclosed. The method includes parsing syntax elements from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax elements indicate a maximum allowed size within which a current video block of the one or more blocks of the video region is allowed to be encoded or decoded using a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and generating a decoded video region from the bitstream representation by decoding the one or more video blocks according to the maximum allowed size.

[0008] In yet another exemplary aspect, a method for visual media processing is disclosed. The method includes determining to codec a current video block of visual media data using a transform skip mode; and based on the determination, performing a conversion between the current video block and a bitstream representation of the visual media data, wherein during the conversion, the current video block is divided into a plurality of coefficient groups, and signaling of a codec block flag for at least one of the plurality of coefficient groups is excluded from the bitstream representation, wherein in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform.

[0009] In another exemplary aspect, a method for visual media processing is disclosed. The method includes determining to encode and decode a current video block of visual media data using a transform skip mode; and based on the determination, performing a conversion between the current video block and a bitstream representation of the visual media data, wherein during the conversion, the current video block is divided into a plurality of coefficient groups, wherein in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform, and wherein during the conversion, a coefficient scan order of the plurality of coefficient groups is determined based at least in part on an indication in the bitstream representation.

[0010] In yet another exemplary aspect, a method for visual media processing is disclosed. The method includes encoding a current video block in a video region of visual media data into a bitstream representation of the visual media data using a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and selecting a context for a sign flag of the current video block based on partitioning the current video block into a plurality of coefficient groups and depending on sign flags of one or more neighboring video blocks in the coefficient groups associated with the current video block.

[0011] In yet another exemplary aspect, a method for visual media processing is disclosed. The method includes parsing a bitstream representation of visual media data including a video region including a current video block to identify a context for a sign flag used in a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and generating a decoded video region from the bitstream representation such that the context for the sign flag is based on partitioning the current video block into a plurality of coefficient groups and on sign flags of one or more neighboring video blocks in the coefficient groups associated with the current video block.

[0012] In yet another exemplary aspect, a method for visual media processing is disclosed. The method includes determining a position of a current coefficient associated with the current video block of visual media data when the current video block is divided into a plurality of coefficient positions; deriving a context for a sign flag of the current coefficient based at least on sign flags of one or more neighboring coefficients; and generating a sign flag for the current coefficient based on the context, wherein the sign flag of the current coefficient is used in a transform skip mode in which the current video block is encoded or decoded without applying a transform.

[0013] In another exemplary aspect, a method for visual media processing is disclosed. The method includes: for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, determining that a chroma transform skip mode is applicable to a current video block based on satisfying at least one rule, wherein in the chroma transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform; and including a syntax element in the bitstream representation indicating the chroma transform skip mode.

[0014] In another exemplary aspect, a method for visual media processing is disclosed. The method includes parsing a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, the syntax element indicating that at least one rule associated with application of a chroma transform skip mode is satisfied, wherein in the chroma transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform; and generating a decoded video region from the bitstream representation by decoding the one or more video blocks.

[0015] In yet another exemplary aspect, a method for visual media processing is disclosed. The method includes: for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, making a decision based on a condition to selectively apply a transform skip mode to a current video block; and including a syntax element in the bitstream representation indicating the condition, wherein in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform.

[0016] In another exemplary aspect, a method for visual media processing is disclosed. The method includes parsing syntax elements from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax elements indicate a condition related to use of a transform skip mode in which a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation without applying a transform; and generating a decoded video region from the bitstream representation by decoding the one or more video blocks according to the condition.

[0017] In another exemplary aspect, a method for visual media processing is disclosed. The method includes: for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, based on an indication of a transform skip mode in the bitstream representation, making a decision regarding the selective application of a Quantized Residual Block Differential Pulse-Code Modulation (QR-BDPCM) technique, wherein in the transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform, wherein in the QR-BDPCM technique, the residual of the prediction error in the horizontal direction and / or the vertical direction is quantized and entropy coded; and including an indication of the selective application of the QR-BDPCM technique in the bitstream representation.

[0018] In another exemplary aspect, a method for visual media processing is disclosed. The method includes parsing syntax elements from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax elements indicate selective application of a quantized residual block differential pulse coding modulation (QR-BDPCM) technique based on an indication of a transform skip mode in the bitstream representation, wherein in the transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform, wherein in the QR-BDPCM technique, the residual of the prediction error in the horizontal direction and / or the vertical direction is quantized and entropy coded; and generating a decoded video region from the bitstream representation by decoding the one or more video blocks according to a maximum allowed size.

[0019] In yet another exemplary aspect, a method for visual media processing is disclosed, comprising: determining, based on a condition, selective application of a single or dual tree for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data; and including a syntax element in the bitstream representation indicating the selective application of the single or dual tree.

[0020] In another exemplary aspect, a method for visual media processing is disclosed, comprising: parsing syntax elements from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax elements indicate selective application of a single tree or a dual tree based on a condition or inferred from a condition; and generating a decoded video region from the bitstream representation by decoding the one or more video blocks according to the syntax elements.

[0021] In yet another example aspect, a video encoder or decoder apparatus is disclosed, comprising a processor configured to implement the above method.

[0022] In another exemplary aspect, a computer-readable program medium is disclosed, wherein the medium stores code embodying processor-executable instructions for implementing one of the disclosed methods.

[0023] These and other aspects are described further throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 An example of intra block copy is shown.

[0025] Figure 2 An example of a block encoded and decoded in palette mode is shown.

[0026] Figure 3 An example of using palette prediction values ​​to signal palette entries is shown.

[0027] Figure 4 Examples of horizontal and vertical traversal scans are shown.

[0028] Figure 5 An example of encoding and decoding of palette indexes is shown.

[0029] Figure 6 An example of a multi-type tree partitioning pattern is shown.

[0030] Figure 7 An example of sample points used to derive parameters in a Cross-Component Linear Model (CCLM) prediction mode is shown.

[0031] Figure 8 An exemplary architecture for luma mapping with chroma scaling is shown.

[0032] Figures 9A-9E is a flowchart of an example of a video processing method.

[0033] Figure 10 is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.

[0034] Figure 11 is a block diagram of an example video processing system in which the disclosed technology may be implemented.

[0035] Figure 12 is a flow chart of an example method for visual media processing.

[0036] Figure 13 is a flow chart of an example method for visual media processing.

[0037] Figure 14 is a flow chart of an example method for visual media processing.

[0038] Figure 15 is a flow chart of an example method for visual media processing.

[0039] Figure 16 is a flow chart of an example method for visual media processing.

[0040] Figure 17 is a flow chart of an example method for visual media processing.

[0041] Figure 18 is a flow chart of an example method for visual media processing.

[0042] Figure 19is a flow chart of an example method for visual media processing.

[0043] Figure 20 is a flow chart of an example method for visual media processing.

[0044] Figure 21 is a flow chart of an example method for visual media processing.

[0045] Figure 22 is a flow chart of an example method for visual media processing.

[0046] Figure 23 is a flow chart of an example method for visual media processing.

[0047] Figure 24 is a flow chart of an example method for visual media processing.

[0048] Figure 25 is a flow chart of an example method for visual media processing.

[0049] Figure 26 is a flow chart of an example method for visual media processing. DETAILED DESCRIPTION

[0050] This document provides various techniques that can be used by decoders of image or video bitstreams to improve the quality of decompressed or decoded digital video or images. For brevity, the term "video" as used herein includes both sequences of pictures (traditionally referred to as video) and individual images. In addition, video encoders can also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.

[0051] For ease of understanding, section headings are used in this document and do not limit the embodiments and techniques to the corresponding sections. Thus, embodiments from one section can be combined with embodiments from other sections.

[0052] 1. Overview

[0053] This document relates to video coding / decoding techniques. Specifically, it relates to coefficient coding and decoding in transform skip mode in video codecs. It can be applied to existing video codec standards such as HEVC, or to the upcoming standard (Universal Video Codec). It may also be applicable to future video codec standards or video codecs.

[0054] 2. Preliminary Discussion

[0055] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4S Visual, and the two organizations jointly produced H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC [1,2]. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction and transform coding. To advance future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Expert Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC.

[0056] The latest version of the VVC draft, the Universal Video Codec (Draft 4), can be found at:

[0057] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / current_document.php?id=5755VVC

[0058] The latest reference software for VVC, called VTM, can be found at:

[0059] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-4.0

[0060] Intra-block copy

[0061] Intra block copy (IBC), also known as current picture reference, has been adopted by HEVC Screen Content Coding extension (HEVC-SCC) and the current VVC test model (VTM-4.0). IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. Figure 1 As shown in the figure, when IBC is applied, the current block is predicted by a reference block in the same picture. Before encoding or decoding the current block, the samples in the reference block must have been reconstructed. Although IBC is not efficient for most camera-captured sequences, it shows significant codec gains for screen content. The reason is that there are many repeated patterns in screen content pictures, such as icons and text characters. IBC can effectively remove the redundancy between these repeated patterns. In HEVC-SCC, if the coding unit (CU) of the inter-frame codec selects the current picture as its reference picture, it can apply IBC. In this case, the MV is renamed to the block vector (BV), and the BV always has integer pixel precision. For compatibility with main profile HEVC, the current picture is marked as a "long-term" reference picture in the decoded picture buffer (DPB). It should be noted that similarly, in the multi-view / 3D video codec standard, inter-view reference pictures are also marked as "long-term" reference pictures.

[0062] After the BV finds its reference block, it can generate a prediction by copying the reference block. The residual can be obtained by subtracting the reference pixels from the original signal. Transformation and quantization can then be applied as in other codec modes.

[0063] However, when the reference block is outside the picture, overlaps with the current block, is outside the reconstructed area, or is outside the valid area subject to certain constraints, some or all pixel values ​​are undefined. There are basically two ways to solve this problem. One is to disallow this situation, such as in bitstream conformance. The other is to apply padding to those undefined pixel values. The following subsections describe these methods in detail.

[0064] IBC in HEVC Screen Content Codec Extension

[0065] In the screen content codec extension of HEVC, when a block uses the current picture as a reference, it should ensure that the entire reference block is within the available reconstruction area, as indicated by the following protocol text:

[0066] The variables offsetX and offsetY are derived as follows:

[0067] offsetX=(ChromaArrayType==0)? 0:(mvCLX[0]&0x7?2:0) (8-106)

[0068] offsetY=(ChromaArrayType==0)? 0:(mvCLX[1]&0x7?2:0) (8-107)

[0069] The bitstream conformance requirement is that when the reference picture is the current picture, the luma motion vector mvLX shall obey the following constraints:

[0070] - When the derivation procedure for z-scan order block availability specified in clause 6.4.1 is called with (xCurr,yCurr) set equal to (xCb,yCb) and the adjacent luma position (xNbY,yNbY) set equal to (xPb+(mvLX[0]>>2)-offsetX,yPb+(mvLX[1]>>2)-offsetY) as input, the output shall be equal to TRUE.

[0071] - When the derivation procedure for z-scan order block availability specified in clause 6.4.1 is called with as input (xCurr,yCurr) set equal to (xCb,yCb) and the adjacent luma position (xNbY,yNbY) set equal to (xPb+(mvLX[0]>>2)+nPbW-1+offsetX,yPb+(mvLX[1]>>2)+nPbH-1+offsetY), the output shall be equal to TRUE.

[0072] - One or both of the following conditions should be true:

[0073] The value of (mvLX[0]>>2)+nPbW+xB1+offsetX is less than or equal to 0.

[0074] The value of (mvLX[1]>>2)+nPbH+yB1+offsetY is less than or equal to 0.

[0075] - The following conditions should hold:

[0076] (xPb+(mvLX[0]>>2)+nPbSw-1+offsetX) / CtbSizeY-xCurr / CtbSizeY<=yCurr / CtbSizeY-(yPb+(mvLX[1]>>2)+nPbSh-1+offsetY) / CtbSizeY (8-108)

[0077] Therefore, there is no chance that the reference block overlaps with the current block or that the reference block is outside the picture. No padding of the reference or prediction blocks is required.

[0078] IBC in the VVC test model

[0079] In the current VVC test model, that is, the VTM-4.0 design, the entire reference block should be together with the current coding tree unit (CTU) and does not overlap with the current block. Therefore, there is no need to fill the reference or prediction block. The IBC flag is encoded and decoded as the prediction mode of the current CU. Therefore, for each CU, there are a total of three prediction modes: MODE_INTRA, MODE_INTER, and MODE_IBC.

[0080] IBC Merge mode

[0081] In IBC Merge mode, the index pointing to the entry in the IBC Merge candidate list is parsed from the bitstream. The construction of the IBC Merge list can be summarized as the following sequence of steps:

[0082] Step 1: Derivation of spatial candidates

[0083] Step 2: Insert HMVP

[0084] Step 3: Insert pairwise average candidates

[0085] In the derivation of spatial Merge candidates, up to four Merge candidates are selected from the candidates at positions A1, B1, B0, A0, and B2. The order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only when any PU at position A1, B1, B0, A0 is unavailable (for example, because it belongs to another slice or piece) or is not coded in IBC mode. After adding the candidate at position A1, the insertion of the remaining candidates is subject to redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency.

[0086] After inserting the spatial candidates, if the IBC Merge list size is still less than the maximum IBC Merge list size, then IBC candidates from the HMVP table can be inserted. When inserting HMVP candidates, a redundancy check is performed.

[0087] Finally, the pairwise average candidates are inserted into the IBC Merge list.

[0088] When the reference block identified by a Merge candidate is outside the picture, overlaps with the current block, is outside the reconstruction area, or is outside the valid area subject to certain constraints, the Merge candidate is called an invalid Merge candidate.

[0089] Note that invalid merge candidates can be inserted into the IBC merge list.

[0090] IBC AMVP model

[0091] In IBC AMVP mode, the AMVP index pointing to the entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized as the following sequence of steps:

[0092] Step 1: Derivation of spatial candidates

[0093] Check A0, A1 until a usable candidate is found.

[0094] Check B0, B1, B2 until a usable candidate is found.

[0095] Step 2: Insert HMVP candidates

[0096] Step 3: Insert zero candidates

[0097] After inserting the spatial candidates, if the IBC AMVP list size is still smaller than the maximum IBC AMVP list size, the IBC candidates from the HMVP table may be inserted.

[0098] Finally, the zero candidate is inserted into the IBC AMVP list.

[0099] Adaptive Motion Vector Resolution (AMVR)

[0100] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the CU's motion vector and the predicted motion vector) is signaled in units of 1 / 4 luma samples. In VVC, the CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the CU's MVD to be encoded and decoded with different precisions. Depending on the current CU mode (conventional AMVP mode or affine AVMP mode), the current CU's MVD can be adaptively selected as follows:

[0101] - Conventional AMVP mode: 1 / 4 luminance sample, integer luminance sample or four luminance samples.

[0102] -Affine AMVP mode: 1 / 4 luma samples, integer luma samples, or 1 / 16 luma samples.

[0103] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., horizontal and vertical MVD for reference list L0 and reference list L1) are zero, a 1 / 4 luma sample MVD resolution is inferred.

[0104] For CUs with at least one non-zero MVD component, a first flag is signaled to indicate whether 1 / 4 luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and 1 / 4 luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether integer luma samples or four luma sample MVD precision is used for regular AMVP CUs. The same second flag is used to indicate whether integer luma samples or 1 / 16 luma sample MVD precision is used for affine AMVP CUs. To ensure that the reconstructed MV has the expected precision (1 / 4 luma sample, integer luma sample, or four luma samples), the motion vector prediction value of the CU will be rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded towards zero (i.e., negative motion vector prediction values ​​are rounded towards positive infinity, while positive motion vector prediction values ​​are rounded towards negative infinity).

[0105] The encoder uses RD check to determine the motion vector resolution of the current CU. In order to avoid always performing three CU-level RD checks for each MVD resolution, in VTM4, the RD check for MVD precision other than 1 / 4 luma sample is only conditionally called. For the conventional AVMP mode, the RD cost of 1 / 4 luma sample MVD precision and integer luma sample MV precision is first calculated. Then, the RD cost of integer luma sample MVD precision is compared with the RD cost of 1 / 4 luma sample MVD precision to decide whether it is necessary to further check the RD cost of four luma sample MVD precision. When the RD cost of 1 / 4 luma sample MVD precision is much smaller than the RD cost of integer luma sample MVD precision, the RD check of four luma sample MVD precision is skipped. For the affine AMVP mode, if the affine inter mode is not selected after checking the rate-distortion cost of the affine merge / skip mode, merge / skip mode, 1 / 4 luma sample MVD accuracy regular AMVP mode, and 1 / 4 luma sample MVD accuracy affine AMVP mode, the 1 / 16 luma sample MV accuracy and 1 pixel MV accuracy affine inter modes are not checked. In addition, in the 1 / 16 luma sample and 1 / 4 luma sample MV accuracy affine inter modes, the affine parameters obtained in the 1 / 4 luma sample MV accurate affine inter mode are used as the starting search point.

[0106] Palette Mode

[0107] The basic idea behind palette mode is that samples in a CU are represented by a small set of representative color values. This set is called a palette. Samples outside the palette can also be indicated by signaling an escape symbol after the component value (possibly quantized). This is done in Figure 2 Shown in.

[0108] Palette Mode in HEVC Screen Content Codec Extension (HEVC-SCC)

[0109] In the palette mode of HEVC-SCC, a prediction method is used to encode and decode the palette and index map.

[0110] Encoding and decoding of palette entries

[0111] For encoding and decoding of palette entries, palette prediction values ​​are maintained. The maximum size of the palette and the palette prediction values ​​are signaled in the SPS. In HEVC SCC, palette_predictor_initializer_present_flag is introduced in the PPS. When this flag is 1, the entry for initializing the palette prediction values ​​is signaled in the bitstream. The palette prediction values ​​are initialized at the beginning of each CTU row, each slice, and each slice. Depending on the value of palette_predictor_initializer_present_flag, the palette prediction values ​​are reset to 0 or initialized using the palette prediction value initialization value entry signaled in the PPS. In HEVC-SCC, palette prediction value initialization values ​​of size 0 are enabled to allow palette prediction value initialization to be explicitly disabled at the PPS level.

[0112] For each entry in the palette prediction, a reuse flag is signaled to indicate whether it is part of the current palette. Figure 3 The reuse flag is signaled using run-length coding of zero. After that, the number of new palette entries is signaled using exponential Golomb coding of order 0. Finally, the component values ​​of the new palette entries are signaled.

[0113] Palette index encoding and decoding

[0114] Use horizontal and vertical traversal scans to encode and decode palette indices, such as Figure 4The scanning order is explicitly signaled in the bitstream using palette_transpose_flag. For the rest of this subclause, it is assumed that the scanning is horizontal.

[0115] The palette index is encoded and decoded using two main palette sample modes: "INDEX" and "COPY_ABOVE". As mentioned before, the escape symbol is also signaled in "INDEX" mode and is assigned an index equal to the maximum palette size. Flags are used to signal the mode, except when the top row or the previous mode is "COPY_ABOVE". In "COPY_ABOVE" mode, the palette index of the sample in the row above is copied. In "INDEX" mode, the palette index is signaled explicitly. For both "INDEX" and "COPY_ABOVE" modes, a run value is signaled that specifies the number of subsequent samples to be encoded and decoded using the same mode. When an escape symbol is part of a run in "INDEX" or "COPY_ABOVE" mode, the escape component value is signaled for each escape symbol. The encoding and decoding of the palette index is as follows Figure 5 shown.

[0116] The syntax sequence is done as follows. First, the number of index values ​​for the CU is signaled. Next, the actual index values ​​for the entire CU are signaled using truncated binary codec. Both the number of indices and the index values ​​are coded in bypass mode. This groups the index-related bypass bins together. Then the palette sample mode (if necessary) and run are signaled in an interleaved manner. Finally, the component escape values ​​corresponding to the escape samples for the entire CU are grouped together and coded in bypass mode.

[0117] An additional syntax element, last_run_type_flag, is signaled after the index value is signaled. This syntax element, combined with the number of indices, eliminates the need to signal the run value corresponding to the last run in the block.

[0118] In HEVC-SCC, palette mode is also enabled for 4:2:2, 4:2:0 and monochrome chroma formats. The signaling of palette entries and palette indices is almost the same for all chroma formats. In the case of non-monochrome formats, each palette entry consists of 3 components. For monochrome formats, each palette entry consists of a single component. For the subsampled chroma direction, chroma samples are associated with luma sample indices that are divisible by 2. After reconstructing the palette index of the CU, if a sample has only one component associated with it, only the first component of the palette entry is used. The only difference in the signaling is the escape component values. For each escape sample, the number of escape component values ​​signaled may be different depending on the number of components associated with the sample.

[0119] Coefficient encoding and decoding in transform skip mode

[0120] In JVET-M0464 and JVET-N0280, several modifications to the coefficient coding in transform skip (TS) mode are proposed in order to adapt the residual coding to the statistical and signal characteristics of the transform skip level.

[0121] The proposed modifications are as follows.

[0122] No last valid scan position : Since the residual signal reflects the spatial residual after prediction and no energy compression by transform is performed for TS, it no longer gives a higher probability of trailing zeros or non-significant levels at the bottom right corner of the transform block. Therefore, in this case, the signaling of the last valid scan position is omitted.

[0123] Sub-block CBF : The absence of signaling of the last valid scan position requires that the sub-block CBF signaling with coded_sub_block_flag for TS be modified as follows:

[0124] Due to quantization, the aforementioned non-significant sequence may still appear locally inside the transform block. Therefore, as described above, the last valid scan position is removed and coded_sub_block_flag is encoded and decoded for all sub-blocks.

[0125] ○ The coded_sub_block_flag of the sub-block covering the DC frequency position (the upper left sub-block) represents a special case. In VVC Draft 3, the coded_sub_block_flag of this sub-block is never signaled and is always inferred to be equal to 1. When the last valid scan position is in another sub-block, this means that there is at least one valid level outside the DC sub-block. Therefore, although the coded_sub_block_flag of this sub-block is inferred to be equal to 1, the DC sub-block may only contain zero / non-valid levels. In the absence of information about the last scan position in the TS, the coded_sub_block_flag of each sub-block is signaled. This also includes the coded_sub_block_flag of the DC sub-block, unless all other coded_sub_block_flag syntax elements are already equal to 0. In this case, the DC coded_sub_block_flag is inferred to be equal to 1 (inferDcSbCbf=1). Because there must be at least one valid level in this DC sub-block, if all other sig_coeff_flag syntax elements in this DC sub-block are equal to 0, the first position sig_coeff_flag syntax element at (0,0) is not signaled and is derived to be equal to 1 (inferSbDcSigCoeffFlag=1).

[0126] o The context modeling of coded_sub_block_flag is changed. The context model index is calculated as the sum of the coded_sub_block_flag to the left of the current sub-block and the coded_sub_block_flag above, rather than the logical disjunction of the two.

[0127] sig_coeff_flag context modeling The local template in the sig_coeff_flag context modeling is modified to include only the neighbors to the left (NB0) and above (NB1) of the current scan position. The context model offset is simply the number of valid neighbors sig_coeff_flag[NB0] + sig_coeff_flag[NB1]. Thus, the selection of different context sets depending on the diagonal d within the current transform block is removed. This results in three context models and a single context model set for encoding and decoding sig_coeff_flag.

[0128] abs_level_gt1_flag and par_level_flag context modeling : Use a single context model for abs_level_gt1_flag and par_level_flag.

[0129] abs_remainder codec : Although the empirical distribution of the absolute levels of the transform skip residuals is generally still fit to a Laplacian or geometric distribution, there is greater instability than in the absolute levels of the transform coefficients. In particular, for the absolute levels of the residuals, the variance within the window of consecutive realizations is higher. This motivates the following modifications to the abs_remainder syntax binarization and context modeling:

[0130] ○ Using higher cutoff values ​​in binarization, i.e., the transition point from a codec with sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag to a Rice codec with abs_remainder, and a dedicated context model for each bin position, yields higher compression efficiency. Increasing the cutoff value will result in more "greater than X" flags, such as introducing abs_level_gt5_flag, abs_level_gt7_flag, etc. until the cutoff value is reached. The cutoff value itself is fixed to 5 (numGtFlags=5).

[0131] o The template used for rice parameter derivation is modified, i.e., only the neighbors to the left of the current scan position and the neighbors above the current scan position are considered similar to the local template used for sig_coeff_flag context modeling.

[0132] coeff_sign_flag context modeling Due to the instability of the sequence of symbols and the fact that the prediction residuals are often biased, a context model can be used to encode and decode the symbol even when the global empirical distribution is almost uniform. A single dedicated context model is used for encoding and decoding the symbol, and the symbol is parsed after sig_coeff_flag to keep all context encoding and decoding bins together.

[0133] Quantized Residual Block Differential Pulse Coded Modulation (QR-BDPCM)

[0134] In JVET-M0413, Quantized Residual Block Differential Pulse Codec Modulation (QR-BDPCM) is proposed to efficiently encode and decode screen content.

[0135] The prediction direction used in QR-BDPCM can be vertical and horizontal prediction mode. Similar to intra prediction, intra prediction is performed on the entire block by copying the samples in the prediction direction (horizontal or vertical prediction). The residual is quantized and the deviation between the quantized residual and its predicted value (horizontal or vertical) is encoded and decoded. This can be described as follows: For a block of size M (rows) × N (columns), let r i,j, 0≤i≤M-1,0≤j≤N-1 is the prediction residual after performing intra prediction horizontally (copying the left neighboring pixel values ​​on the prediction block row by row) or vertically (copying the top neighboring row to each row in the prediction block) using the unfiltered samples from the block boundary samples on the upper or left side. Let Q(r i,j ), 0≤i≤M-1,0≤j≤N-1 represents the residual r i,j The quantized version of , where the residual is the difference between the original block and the predicted block value. Then, block DPCM is applied to the quantized residual samples to obtain a sample with elements The modified M×N array When vertical BDPCM is signaled:

[0136]

[0137] For horizontal prediction, similar rules apply and the residual quantization samples are obtained as follows:

[0138]

[0139] Residual quantization samples is sent to the decoder.

[0140] At the decoder side, the above calculation is inversely performed to produce Q(r i,j ),0≤i≤M-1,0≤j≤N-1, for vertical prediction,

[0141]

[0142] For the horizontal case,

[0143]

[0144] Inverse quantization residual, Q -1 (Q(r i,j )) is added to the intra block prediction value to produce the reconstructed sample value.

[0145] The main advantage of this scheme is that inverse DPCM can be done on the fly during coefficient parsing by simply adding prediction values ​​as the coefficients are parsed, or it can be performed after parsing.

[0146] The changes to the draft text of QR-BDPCM are as follows.

[0147] 7.3.6.5 Codec unit syntax

[0148]

[0149]

[0150] bdpcm_flag[x0][y0] equal to 1 specifies that bdpcm_dir_flag is present in the codec unit including the luma codec block at position (x0, y0).

[0151] bdpcm_dir_flag[x0][y0] equal to 0 specifies that the prediction direction to be used in the bdpcm block is horizontal, otherwise it is vertical.

[0152] Split structure

[0153] Use tree structure to split CTU

[0154] In HEVC, CTU is divided into CUs using a quadtree structure represented as a codec tree to adapt to various local characteristics. The decision whether to use inter-frame (temporal domain) or intra-frame (spatial domain) prediction to encode and decode the picture area is made at the leaf CU level. Depending on the PU partition type, each leaf CU can be further divided into one, two or four PUs. Within a PU, the same prediction process is applied, and relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the leaf CU can be divided into transform units (TUs) according to another quadtree structure similar to the codec tree of the CU. An important feature of the HEVC structure is that it has a multi-partition concept, including CU, PU and TU.

[0155] In VVC, a quadtree with nested multi-type trees using a binary and ternary partitioning segment structure replaces the concept of multiple partition unit types, that is, it removes the separation of CU, PU, ​​and TU concepts, except for the need for CUs that are too large for the maximum transform length, and supports greater flexibility in CU partition shapes. In the codec tree structure, a CU can be square or rectangular. The codec tree unit (CTU) is first partitioned by a quadtree (also known as a quadtree) structure. Then, the quadtree leaf nodes can be further partitioned by a multi-type tree structure. As Figure 6As shown, there are four types of splitting in the multi-type tree structure, vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER) and horizontal ternary splitting (SPLIT_TT_HOR). The multi-type tree leaf nodes are called codec units (CUs), and unless the CU is too large for the maximum transform length, the segment is used for prediction and transform processing without any further splitting. This means that in most cases, in a quadtree with a nested multi-type tree codec block structure, the CU, PU and TU have the same block size. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CU. In addition, the luminance and chrominance components have separate split structures on the I slice.

[0156] Cross-component linear model prediction

[0157] To reduce cross-component redundancy, the Cross-Component Linear Model (CCLM) prediction mode is used in VTM4. For this mode, the chroma samples are predicted based on the reconstructed luma samples of the same CU using the following linear model:

[0158] pred C (i,j)=α·rec L ′(i,j)+β

[0159] Among them, pred C (i, j) represents the predicted chroma sample in CU, rec L (i, j) represents the downsampled reconstructed luma sample of the same CU. The linear model parameters α and β are derived from the relationship between the luma and chroma values ​​of two samples, which are the luma sample with the minimum sample value and the luma sample with the maximum sample value in the set of downsampled adjacent luma samples, and their corresponding chroma samples. The linear model parameters α and β are obtained according to the following equations:

[0160]

[0161] β=Y b -α·X b

[0162] Among them, Y a and X a Indicates the luminance and chrominance values ​​of the luminance sample with the maximum luminance sample value. b and Y b Represents the luminance value and chrominance value of the luminance sample with the minimum luminance sample, respectively. Figure 7 An example of the positions of the left and upper samples and the samples of the current block involved in the CCLM mode is shown.

[0163] Luma Mapping with Chroma Scaling (LMCS)

[0164] In VTM4, a codec tool called Luma Mapping and Chroma Scaling (LMCS) was added as a new processing module before the loop filter. LMCS has two main components: 1) loop mapping of the luma component based on an adaptive piecewise linear model; 2) for the chroma components, luma-dependent chroma residual scaling is applied. Figure 8 The LMCS architecture is shown from the decoder's perspective. Figure 8 The shaded blocks in indicate where processing is applied in the mapped domain; these include inverse quantization, inverse transform, luma intra prediction, and the addition of the luma prediction to the luma residual. Figure 8 The non-shaded blocks in indicate where processing is applied in the original (i.e., non-mapped) domain; and these include loop filters such as deblocking, ALF and SAO, motion compensated prediction, chroma intra prediction, addition of chroma prediction with chroma residual, and storage of decoded pictures as reference pictures. Figure 8 The light yellow shaded block in the figure is the new LMCS functional block, which includes the forward mapping and inverse mapping of the luma signal and the chroma scaling process that depends on the luma. Like most other tools in VVC, LMCS can be enabled / disabled at the sequence level using the SPS flag.

[0165] 3. Examples of Problems Solved by the Embodiments

[0166] Although the coefficient coding in JVET-N0280 can achieve coding advantages over screen content coding, the coefficient coding and transform skip (TS) mode still have some disadvantages.

[0167] (1) The maximum allowed width or height of the TS mode is controlled by a common value in the PPS, which may limit flexibility.

[0168] (2) For TS mode, each coding group (CG) needs to signal the cbf flag, which may increase the overhead cost.

[0169] (3) The coefficient scanning order does not take into account the intra prediction mode.

[0170] (4) The symbol flag encoding and decoding uses only one context.

[0171] (5) Transform skipping on chroma components is not supported.

[0172] (6) The transform skip flag is applied to all prediction modes, which increases the overhead cost and coding complexity.

[0173] 4. Examples of Embodiments

[0174] The following detailed inventions should be considered as examples to explain the general concept. These inventions should not be interpreted narrowly. In addition, these inventions can be combined in any way.

[0175] 1. The indication of the maximum allowed width and maximum allowed height of transform skip may be signaled in SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group.

[0176] a. In one example, the maximum allowed width and maximum allowed height of the transform skip can be indicated by different messages signaled in SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group.

[0177] b. In one example, the maximum allowed width and / or height may be first signaled in SPS / PPS and then updated in picture header / slice header / slice group header / LCU row / LCU group.

[0178] 2. The TS codec block can be divided into several coefficient groups (CGs), and the signaling notification of the coded block flag (Cbf) of at least one CG can be skipped.

[0179] a. In one example, signaling of the Cbf flag for all CGs may be skipped, e.g., for TS-encoded blocks.

[0180] b. In one example, for TS mode, the skipped cbf flag of the CG can be inferred to be 1.

[0181] c. In one example, whether some or all of the Cbf flags for the CG are skipped depends on the codec mode.

[0182] i. In one example, for intra blocks of TS codec, signaling of all Cbf flags of CG is skipped.

[0183] d. In one example, the skipped Cbf flag for a CG can be inferred based on:

[0184] i. Messages signaled via SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group / LCU / CU

[0185] ii. CG location

[0186] iii. The block size of the current block and / or its neighboring blocks

[0187] iv. Block shape of the current block and / or its neighboring blocks

[0188] v. The most likely mode of the current block and / or its neighboring blocks

[0189] vi. Prediction mode of the neighboring blocks of the current block (intra / inter)

[0190] vii. Intra-frame prediction mode of the neighboring blocks of the current block

[0191] viii. Motion vector of the adjacent block of the current block

[0192] ix. Indication of the QR-BDPCM mode of the neighboring blocks of the current block

[0193] x. Current quantization parameters of the current block and / or its neighboring blocks

[0194] xi. Indication of the color format (e.g. 4:2:0, 4:4:4)

[0195] xii. Single / Dual Codec Tree Structure

[0196] xiii. Slice / slice group type and / or picture type

[0197] 3. The coefficient scanning order in a TS-coded block may depend on the message signaled in SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group / LCU / CU.

[0198] a. Alternatively, when TS is adopted, the CG and / or coefficient scanning order can depend on the intra prediction mode.

[0199] i. In one example, if the intra prediction mode is horizontally dominant, the scanning order may be vertical.

[0200] 1. In one example, if the intra prediction mode index ranges from 2 to 34, the scanning order may be vertical.

[0201] 2. In one example, if the intra prediction mode index ranges from 2 to 33, the scanning order may be vertical.

[0202] ii. In one example, if the intra prediction mode is vertical dominant, the scanning order may be vertical.

[0203] 1. In one example, if the intra prediction mode index ranges from 34-66, the scanning order may be vertical.

[0204] 2. In one example, if the intra prediction mode index ranges from 35-66, the scanning order may be vertical.

[0205] iii. In one example, if the intra prediction mode is vertically dominant, the scanning order may be horizontal.

[0206] 1. In one example, if the intra prediction mode index ranges from 34 to 66, the scanning order may be vertical.

[0207] 2. In one example, if the intra prediction mode index ranges from 35 to 66, the scanning order may be vertical.

[0208] iv. In one example, if the intra prediction mode is horizontal-dominant, the scanning order may be horizontal.

[0209] 1. In one example, if the intra prediction mode index ranges from 2 to 34, the scanning order may be vertical.

[0210] 2. In one example, if the intra prediction mode index ranges from 2 to 33, the scanning order may be vertical.

[0211] 4. It is proposed that for TS mode, the context of symbol flag encoding and decoding can depend on the adjacent information in the coefficient block.

[0212] a. In one example, for TS mode, the context for encoding and decoding the current symbol flag may depend on the value of the adjacent symbol flag.

[0213] i. In one example, the context for encoding and decoding the current symbol flag may depend on the value of the symbol flag of the left and / or upper neighbor.

[0214] 1. In one example, the context of the current symbol can be derived as C=(L+A), where C is the context id, L is the symbol of its left neighbor, and A is the symbol of its upper neighbor.

[0215] 2. In one example, the context of the current symbol can be derived as C=(L+A*2), where C is the context id, L is the symbol of its left neighbor, and A is the symbol of its upper neighbor.

[0216] 3. In one example, the context of the current symbol can be derived as C=(L*2+A), where C is the context id, L is the symbol of its left neighbor, and A is the symbol of its upper neighbor.

[0217] ii. In one example, the context for encoding and decoding the current symbol flag may depend on the values ​​of the symbol flags of the left neighbor, the above neighbor, and the above-left neighbor.

[0218] iii. In one example, the context for encoding and decoding the current symbol flag may depend on the values ​​of the symbol flags of the left neighbor, the upper neighbor, the upper-left neighbor, and the upper-right neighbor.

[0219] b. In one example, the context of encoding the current symbol flag may depend on the position of the coefficient.

[0220] i. In one example, the context of a symbolic marker may be different in different locations.

[0221] ii. In one example, the context of a symbol mark may depend on x+y, where x and y are the horizontal and vertical positions of a location.

[0222] iii. In one example, the context of the symbol flag may depend on min(x,y), where x and y are the horizontal and vertical positions of the location.

[0223] iv. In one example, the context of a symbol marker may depend on max(x,y), where x and y are the horizontal and vertical positions of the location.

[0224] 5. It is proposed to support chroma transform skip mode.

[0225] a. In one example, the use of chroma transform skip mode can be based on messages signaled in SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group / LCU / CU / video data unit.

[0226] b. Alternatively, the use of chroma transform skip mode may be based on decoded information of one or more representative previously coded blocks in the same color component or other color components.

[0227] i. In one example, if the indication of the representative block's TS flag is false, then the indication of the chroma TS flag can be inferred to be false. Alternatively, if the indication of the representative block's TS flag is true, then the indication of the chroma TS flag can be inferred to be true.

[0228] ii. In one example, the representative block may be a luma block or a chroma block.

[0229] iii. In one example, the representative block may be any block within the collocated luma block.

[0230] iv. In one example, the representative block may be one of the neighboring chroma blocks of the current chroma block.

[0231] v. In one example, the representative block may be a block of corresponding luma samples covering the center chroma sample within the current chroma block.

[0232] vi. In one example, the representative block may be a block covering the corresponding luma samples of the lower right chroma sample in the current chroma block.

[0233] 6. Whether and / or how the transform skip mode is applied may depend on a message signaled in SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group / LCU / CU / video data unit.

[0234] a. In one example, the indication of when and / or how to apply transform skip mode may depend on:

[0235] i. The block size of the current block and / or its neighboring blocks

[0236] ii. Block shape of the current block and / or its neighboring blocks

[0237] iii. The most likely mode of the current block and / or its neighboring blocks

[0238] iv. Prediction mode of the neighboring blocks of the current block (intra / inter)

[0239] v. Intra-frame prediction mode of the neighboring blocks of the current block

[0240] vi. Motion vector of the adjacent block of the current block

[0241] vii. Indication of the QR-BDPCM mode of the neighboring blocks of the current block

[0242] viii. Current quantization parameters of the current block and / or its neighboring blocks

[0243] ix. Indication of color format (such as 4:2:0, 4:4:4)

[0244] x. Single / Dual Codec Tree Structure

[0245] xi. Slice / slice group type and / or picture type

[0246] xii. Time domain layer ID

[0247] b. In one example, when the prediction mode is IBC mode and the block width and / or height is less than / greater than / equal to a threshold, transform skip mode may be applied.

[0248] i. In one example, the threshold value may be 4, 8, 16, or 32.

[0249] ii. In one example, the threshold value may be signaled in the bitstream.

[0250] iii. In one example, the threshold may be based on:

[0251] 1. Messages signaled via SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU

[0252] 2. The block size of the current block and / or its neighboring blocks

[0253] 3. Block shape of the current block and / or its neighboring blocks

[0254] 4. The most likely mode of the current block and / or its neighboring blocks

[0255] 5. Prediction mode of the neighboring blocks of the current block (intra-frame / inter-frame)

[0256] 6. Intra-frame prediction mode of the neighboring blocks of the current block

[0257] 7. Motion vector of the adjacent block of the current block

[0258] 8. Indication of the QR-BDPCM mode of the neighboring blocks of the current block

[0259] 9. Current quantization parameters of the current block and / or its neighboring blocks

[0260] 10. Indication of color format (such as 4:2:0, 4:4:4)

[0261] 11. Single / Dual Codec Tree Structure

[0262] 12. Strip / slice group type and / or picture type

[0263] 13. Time domain layer ID

[0264] 7. Whether the indication of TS mode is signaled may depend on the decoded / derived intra prediction mode.

[0265] a. Alternatively, furthermore, it may depend on the allowed intra prediction modes / directions used in the QR-BDPCM codec block and the use of QR-BDPCM.

[0266] b. For a decoded or derived intra prediction mode, signaling of the TS flag may be skipped if it is part of the allowed set of intra prediction modes / directions used in the QR-BDPCM codec block.

[0267] i. In one example, if QR-BDPCM is allowed to be used for encoding and decoding one slice / picture / slice / brick, vertical and horizontal modes are the two modes allowed in the QR-BDPCM process, and the decoded / derived intra mode is vertical or horizontal mode, then the indication of the TS mode is not signaled.

[0268] c. In one example, when the indication of QR-BDPCM mode (eg, bdpcm_flag) is 1, it can be inferred that transform skip mode is enabled.

[0269] d. The above methods can be applied based on the following:

[0270] i. Messages signaled via SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group / LCU / CU

[0271] ii. The block size of the current block and / or its neighboring blocks

[0272] iii. Block shape of the current block and / or its neighboring blocks

[0273] iv. The most likely mode of the current block and / or its neighboring blocks

[0274] v. Prediction mode of the neighboring blocks of the current block (intra / inter)

[0275] vi. Intra-frame prediction mode of the neighboring blocks of the current block

[0276] vii. Motion vector of the adjacent block of the current block

[0277] viii. Indication of the QR-BDPCM mode of the neighboring blocks of the current block

[0278] ix. Current quantization parameters of the current block and / or its neighboring blocks

[0279] x. Indication of the color format (e.g. 4:2:0, 4:4:4)

[0280] xi. Single / Dual Codec Tree Structure

[0281] xii. Slice / slice group type and / or picture type

[0282] xiii. Time domain layer ID

[0283] 8. Whether and / or how QR-BDPCM is applied may depend on the indication of TS mode.

[0284] a. In one example, the indication of whether to apply QR-BDPCM may be signaled at the transform unit (TU) level, rather than signaled in CU.

[0285] i. In one example, after the indication of TS mode is applied to the TU, an indication of whether to apply QR-BDPCM may be signaled.

[0286] b. In one example, QR-BDPCM is considered as a special case of TS mode.

[0287] i. When a block is coded in TS mode, another flag may be further signaled to indicate whether QR-BDPCM or legacy TS mode is applied. If it is coded in QR-BDPCM, the prediction direction used in QR-BDPCM may be further signaled.

[0288] ii. Alternatively, when a block is coded in TS mode, another flag may be further signaled to indicate which QR-BDPCM mode (eg, QR-BDPCM based on horizontal / vertical prediction direction) or conventional TS mode is applied.

[0289] c. In one example, whether to indicate QR-BDPCM can be inferred based on the indication of TS mode.

[0290] i. In one example, if the indication of whether the transform skip flag is applied to the luma and / or chroma block is true, then the indication of whether QR-BDPCM is applied to the same block can be inferred to be true. Alternatively, if the indication of whether the transform skip flag is applied to the luma and / or chroma block is true, then the indication of whether QR-BDPCM is applied to the same block can be inferred to be true.

[0291] ii. In one example, if the indication of whether the transform skip flag is applied to the luma and / or chroma block is false, the indication of whether QR-BDPCM is applied to the same block can be inferred to be false. Alternatively, if the indication of whether the transform skip flag is applied to the luma and / or chroma block is false, the indication of whether QR-BDPCM is applied to the same block can be inferred to be false.

[0292] 9. Whether and / or how to apply single / dual tree may depend on the message signaled in SPS / VPS / PPS / picture header / slice header / slice group header / LCU row / LCU group / LCU / CU / video data unit.

[0293] a. In one example, the indication of whether to apply a single / dual tree may depend on whether a slice / slice / LCU / LCU row / LCU group / video data unit is determined to be screen content.

[0294] i. In addition, in one example, whether a slice / slice / LCU / LCU row / LCU group / video data unit is determined to be screen content depends on:

[0295] 1. Messages / flags signaled via SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0296] 2. The block size of the current CTU and / or its adjacent CTUs

[0297] 3. Block shape of the current CTU and / or its adjacent CTUs

[0298] 4. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0299] 5. Indication of color format (such as 4:2:0, 4:4:4)

[0300] 6. Single / dual codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0301] 7. Strip / slice group type and / or picture type

[0302] 8. Time domain layer ID

[0303] b. In one example, whether to apply a single / dual tree indication may be inferred, which may depend on:

[0304] i. Messages signaled with SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0305] ii. Hash hit rate of IBC / inter mode in the previous coded picture / slice / strip / reconstructed region

[0306] iii. The block size of the current CTU and / or its adjacent CTUs

[0307] iv. Block shape of the current CTU and / or its neighboring CTUs

[0308] v. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0309] vi. Indication of color format (such as 4:2:0, 4:4:4)

[0310] vii. Single / dual codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0311] viii. Slice / slice group type and / or picture type

[0312] ix. Time domain layer ID

[0313] c. In one example, the indication of whether CCLM and / or LMCS is applied may depend on the single / dual codec tree structure type.

[0314] i. In one example, when single / dual tree is used, the indication of CCLM and / or LMCS can be inferred to be false.

[0315] d. In one example, whether to apply the single / dual tree indication (ie, CTU-level adaptation of a separate tree or a dual tree) may depend on whether the dual tree enable indication (eg, qtbtt_dual_tree_intra_flag) is true.

[0316] i. Alternatively, furthermore, if dual tree for sequence / slice / slice / brick is not enabled (e.g., qtbtt_dual_tree_intra_flag is false), then the indication of whether single / dual tree is applied at a lower level (e.g., CTU) can be inferred to be false.

[0317] 1. Alternatively, if dual tree is enabled (eg, qtbtt_dual_tree_intra_flag is true), an indication of whether single / dual tree is applied for video units smaller than a sequence / slice / slice / brick may be signaled.

[0318] e. In one example, the indication of whether single / dual tree is applied (ie, CTU-level adaptation of single or dual tree) may depend on whether a flag at a level higher than the CTU level (eg, sps_asdt_flag and / or slice_asdt_flag) is true.

[0319] i. Alternatively, furthermore, if a flag at a level higher than the CTU level (e.g., sps_asdt_flag and / or slice_asdt_flag) is false, the indication of whether single / dual tree is applied (i.e., CTU-level adaptation of single tree or dual tree) may be inferred to be false.

[0320] f. The above method can also be applied to single tree splitting, or single / dual codec tree structure type. 10. Whether to enable IBC can depend on the codec tree structure type.

[0321] a. In one example, for a given codec tree structure type (eg, dual tree), signaling of an indication of the IBC mode and / or block vector used in IBC mode may be skipped and the indication may be inferred.

[0322] b. In one example, when the dual codec tree structure type is applied, the indication of IBC mode can be inferred to be false.

[0323] c. In one example, when the dual codec tree structure type is applied, the indication of the IBC mode of the luma block may be inferred to be false.

[0324] d. In one example, when the dual codec tree structure type is applied, the indication of the IBC mode of the chroma block can be inferred to be false.

[0325] e. In one example, an indication of IBC mode may be inferred based on:

[0326] i. Messages signaled with SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0327] ii. Hash hit rate of IBC / inter mode in the previous coded picture / slice / strip / reconstructed region

[0328] iii. The block size of the current CTU and / or its adjacent CTUs

[0329] iv. Block shape of the current CTU and / or its neighboring CTUs

[0330] v. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0331] vi. Indication of color format (such as 4:2:0, 4:4:4)

[0332] vii. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0333] viii. Slice / slice group type and / or picture type

[0334] ix. Time domain layer ID

[0335] 11. Whether CCLM is enabled may depend on the codec tree structure type.

[0336] a. In one example, for a given codec tree structure type (eg, dual tree), signaling of an indication of CCLM mode and / or other syntax related to CCLM mode may be skipped and the indication and / or other syntax may be inferred.

[0337] b. In one example, when the dual codec tree structure type is applied, the indication of CCLM mode can be inferred to be false.

[0338] c. In one example, when a dual codec tree structure type is applied, the indication of CCLM mode can be inferred based on the following:

[0339] i. Messages signaled with SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0340] ii. Hash hit rate of IBC / inter mode in the previous coded picture / slice / strip / reconstructed region

[0341] iii. The block size of the current CTU and / or its adjacent CTUs

[0342] iv. Block shape of the current CTU and / or its neighboring CTUs

[0343] v. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0344] vi. Indication of color format (such as 4:2:0, 4:4:4)

[0345] vii. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0346] viii. Slice / slice group type and / or picture type

[0347] ix. Time domain layer ID

[0348] 12. Whether LMCS is enabled for chroma components may depend on the codec tree structure type.

[0349] a. In one example, for a given codec tree structure type (e.g., dual tree), the signaling of the indication of the LMCS mode of the chroma component and / or other syntax related to the LMCS mode can be skipped and the indication and / or other syntax can be inferred.

[0350] b. In one example, when the dual codec tree structure type is applied, the indication of the LMCS of the chroma component may be inferred to be false.

[0351] c. In one example, when a dual codec tree structure type is applied, the indication of the LMCS of the chroma component can be inferred based on the following:

[0352] i. Messages signaled with SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0353] ii. Hash hit rate of IBC / inter mode in the previous coded picture / slice / strip / reconstructed region

[0354] iii. The block size of the current CTU and / or its adjacent CTUs

[0355] iv. Block shape of the current CTU and / or its neighboring CTUs

[0356] v. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0357] vi. Indication of color format (such as 4:2:0, 4:4:4)

[0358] vii. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0359] viii. Slice / slice group type and / or picture type

[0360] ix. Time domain layer ID

[0361] 13. The codec tree structure may depend on whether IBC is used.

[0362] a. In one example, the dual-tree structure and the IBC method may not be enabled simultaneously at the sequence / picture / slice / brick / CTU / VPDU / 32x32 block / 64x32 block / 32x64 block levels.

[0363] b. Alternatively, further, in one example, if the IBC method is enabled, the dual tree structure may be disabled at the sequence / picture / slice / brick / CTU / VPDU / 32x32 block / 64x32 block / 32x64 block level.

[0364] c. In one example, when IBC is used in a region, the chroma codec tree structure can be consistent with the luma codec tree structure

[0365] i. In one example, the region can be a sequence / picture / slice / brick / CTU / VPDU / 32x32 block / 64x32 block / 32x64 block.

[0366] ii. In one example, when the collocated luminance block is divided into sub-blocks, the chrominance block may be divided into sub-blocks if the chrominance block is allowed to be divided.

[0367] iii. In one example, whether and how the chroma block is partitioned can be inferred from the codec structure of the collocated luminance block.

[0368] iv. In one example, when the chroma codec tree structure is inferred from the luma codec tree structure, encoding and decoding the signal of the chroma codec tree structure may be skipped.

[0369] v. In one example, a flag may be used to indicate whether the chroma codec structure can be inferred from the luma codec structure. The signaling of the flag may depend on:

[0370] 1. Messages signaled using SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0371] 2. Hash hit rate of IBC / inter-frame mode in the previous coded picture / slice / strip / reconstructed area

[0372] 3. Block size of the current CTU and / or its adjacent CTUs

[0373] 4. Block shape of the current CTU and / or its adjacent CTUs

[0374] 5. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0375] 6. Indication of color format (such as 4:2:0, 4:4:4)

[0376] 7. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0377] 8. Strip / slice group type and / or picture type

[0378] 9. Time domain layer ID

[0379] 14. Whether to enable palette codec mode may depend on the codec tree structure type.

[0380] a. In one example, for a given codec tree structure type (eg, dual-tree), signaling of the palette codec mode indication may be skipped and the indication may be inferred.

[0381] b. In one example, when the dual codec tree structure type is applied, the indication of palette codec mode may be inferred to be false.

[0382] c. In one example, when the dual codec tree structure type is applied, the indication of the palette codec mode for the luma block may be inferred to be false.

[0383] d. In one example, when the dual codec tree structure type is applied, the indication of palette codec mode for chroma blocks may be inferred to be false.

[0384] e. In one example, an indication of palette encoding / decoding mode may be inferred based on:

[0385] i. Messages signaled with SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0386] ii. Hash hit rate of IBC / inter mode in the previous coded picture / slice / strip / reconstructed region

[0387] iii. The block size of the current CTU and / or its adjacent CTUs

[0388] iv. Block shape of the current CTU and / or its neighboring CTUs

[0389] v. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0390] vi. Indication of color format (such as 4:2:0, 4:4:4)

[0391] vii. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0392] viii. Slice / slice group type and / or picture type

[0393] ix. Time domain layer ID

[0394] 15. The codec tree structure may depend on whether palette codec mode is used.

[0395] a. In one example, when the palette codec mode is used in a region, the chroma codec tree structure can be consistent with the luma codec tree structure

[0396] i. In one example, the region can be a sequence / picture / slice / brick / CTU / VPDU / 32x32 block / 64x32 block / 32x64 block

[0397] ii. In one example, when the collocated luma block is divided into sub-blocks, the chroma block may be divided into sub-blocks if it is allowed to be divided.

[0398] iii. In one example, whether and how the chroma block is partitioned can be inferred from the codec structure of the collocated luminance block.

[0399] iv. In one example, when the chroma codec tree structure is inferred from the luma codec tree structure, encoding and decoding the signal of the chroma codec tree structure may be skipped.

[0400] v. In one example, a flag may be used to indicate whether the chroma codec structure can be inferred from the luma codec structure. The signaling of the flag may depend on

[0401] 1. Messages signaled using SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0402] 2. Hash hit rate of IBC / inter-frame mode in the previous coded picture / slice / strip / reconstructed area

[0403] 3. Block size of the current CTU and / or its adjacent CTUs

[0404] 4. Block shape of the current CTU and / or its adjacent CTUs

[0405] 5. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0406] 6. Indication of color format (such as 4:2:0, 4:4:4)

[0407] 7. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0408] 8. Strip / slice group type and / or picture type

[0409] 9. Time domain layer ID

[0410] 16. The motion / block vectors of the sub-blocks / samples in the chroma IBC codec block can be derived from the first available IBC codec sub-region in the collocated luminance block.

[0411] a. In one example, a scan order for sub-regions within a collocated luminance block may be defined, such as a raster scan order.

[0412] b. In one example, a sub-region may be defined as a minimum decoding unit / minimum transform unit.

[0413] c. In one example, the motion / block vectors for all samples in chroma IBC mode can be derived based on the motion vector of the top leftmost sample in the collocated luma block coded in IBC or inter mode.

[0414] 17. Motion / block vectors can be signaled in chroma IBC mode.

[0415] a. In one example, the difference between the motion vector and the motion vector predictor may be signaled.

[0416] i. In one example, the motion vector prediction value can be derived based on the motion vector of the collocated luminance block, the adjacent luminance block of the collocated luminance block, and the adjacent chrominance block of the current chrominance block.

[0417] 1. In one example, the motion / block vector prediction value can be derived based on the motion vector of the top left corner sample in the collocated luma block.

[0418] 2. In one example, the motion / block vector prediction value can be derived based on the motion vector of the sample at the center position in the collocated luma block.

[0419] 3. In one example, the motion / block vector predictor can be derived based on the motion vector of the top leftmost sample in the collocated luma block coded in IBC or inter mode.

[0420] ii. In one example, a motion vector predictor associated with a sub-region of the luma component may be scaled before being used as a predictor.

[0421] iii. In one example, the block vector can be derived from the motion vector / block vector of the neighboring (adjacent or non-adjacent) chroma blocks.

[0422] b. In one example, a block vector candidate list may be constructed and the index of the list may be signaled.

[0423] i. In one example, the candidate list may include motion vectors / block vectors from the collocated luma block, the adjacent luma blocks of the collocated luma block, and the adjacent chroma blocks.

[0424] c. In one example, the indication of the AMVR flag can be inferred

[0425] i. In one example, in blocks encoded in chroma IBC mode, the indication of the AMVR flag may be inferred to be false (0).

[0426] ii. In one example, in blocks encoded in chroma IBC mode, the indication of motion vector differences can be inferred to integer precision

[0427] d. In one example, a separate HMVP table may be used on chroma IBC mode.

[0428] i. In one example, the sizes of the chroma HMVP table and the luma HMVP table may be different.

[0429] e. In one example, whether to signal blocks / motion vectors in chroma IBC mode can be based on:

[0430] i. Whether to encode and decode all sub-regions within the collocated luminance block in IBC mode.

[0431] 1. If yes, then the block vectors of the chroma blocks do not need to be signaled. Otherwise, the block vectors of the chroma blocks can be signaled.

[0432] ii. Whether all sub-regions within the collocated luminance block are coded and decoded in IBC mode, and whether all associated block vectors are valid.

[0433] 1. If yes, then the block vectors of the chroma blocks do not need to be signaled. Otherwise, the block vectors of the chroma blocks can be signaled.

[0434] iii. Messages signaled with SPS / VPS / PPS / picture header / slice header / slice group header / LCU line / LCU group / LCU / CU / video data unit.

[0435] iv. Hash hit rate of IBC / inter mode in the previous coded picture / slice / strip / reconstructed region

[0436] v. Block size of the current CTU and / or its neighboring CTUs

[0437] vi. Block shape of the current CTU and / or its adjacent CTUs

[0438] vii. Current quantization parameters of the current CTU and / or its neighboring CTUs

[0439] viii. Indication of color format (such as 4:2:0, 4:4:4)

[0440] ix. Codec tree structure type of the previous slice / slice / LCU / LCU row / LCU group / video data unit

[0441] x. Slice / slice group type and / or picture type

[0442] xi. Time domain layer ID

[0443] 18. Single tree and single / dual tree selection / signaling can be at the level below CTU / CTB

[0444] a. In one example, single tree and single / dual tree selection / signaling can be at VPDU level.

[0445] b. In one example, single tree and single / dual tree selection / signaling can be at a level where explicit split tree signaling starts.

[0446] The above examples may be combined in the context of the methods described below, such as methods 900 , 910 , 920 , 930 , and 940 , which may be implemented at a video decoder or a video encoder.

[0447] An exemplary method for video processing includes performing a conversion between a current video block and a bitstream representation of a video including the current video block, wherein the conversion selectively uses a transform skip mode for the conversion based on an indicator included in the bitstream representation, and wherein, using the transform skip mode, a residual of a prediction error of the current video block is represented in the bitstream representation without applying a transform.

[0448] In some embodiments, the indicator is a maximum allowed width and a maximum allowed height for transform skip mode.

[0449] In some embodiments, the maximum allowed width and the maximum allowed height are signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU) row, or an LCU group.

[0450] In some embodiments, the maximum allowed width and the maximum allowed height are signaled in different messages.

[0451] In some embodiments, the maximum allowed width and the maximum allowed height are signaled in a sequence parameter set (SPS) or a picture parameter set (PPS), and wherein the updated values ​​of the maximum allowed width and the maximum allowed height are signaled in a picture header, a slice header, a slice group header, a largest codec unit (LCU) row, or an LCU group.

[0452] Figure 9A A flow chart of another exemplary method for video processing is shown. Method 900 includes, at step 902, determining to use a transform skip mode for encoding and decoding a current video block.

[0453] The method 900 includes, at step 904 , performing, based on the determination, a conversion between the current video block and a bitstream representation of a video including the current video block.

[0454] In some embodiments, the current video block is divided into a plurality of coefficient groups, and the bitstream representation omits signaling of a codec block flag for at least one of the plurality of coefficient groups. In one example, the bitstream representation omits signaling of a codec block flag for each of the plurality of coefficient groups.

[0455] In some embodiments, a codec block flag omitted from signaling in a bitstream representation is inferred based on one or more of the following: (1) information signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group, and a codec unit (CU), (2) a position of at least one of a plurality of coefficient groups, (3) a block size of a current video block or at least one neighboring block of the current video block, (4) a block shape of the current video block or at least one neighboring block, (5) a block size of the current video block or at least one neighboring block, (10) a current quantization parameter (QP) of the current video block or of at least one neighboring block, (11) an indication of the color format of the current video block, (12) a single codec tree structure or a dual codec tree structure associated with the current video block, or (13) a slice type, slice group type, or picture type of the current video block.

[0456] In some embodiments, the current video block is divided into a plurality of coefficient groups, and the method 900 further includes the step of determining a coefficient scanning order for the plurality of coefficient groups. In one example, the coefficient scanning order is based on a message signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group, or a codec unit (CU).

[0457] In some embodiments, the plurality of coefficient groups or coefficient scanning order is based on an intra-prediction mode of the current video block. In one example, the coefficient scanning order is vertical, and wherein the intra-prediction mode is horizontally dominant. In another example, the coefficient scanning order is horizontal, and wherein the intra-prediction mode is horizontally dominant. For example, the index of the intra-prediction mode ranges from 2 to 33 or from 2 to 34.

[0458] In some embodiments, the plurality of coefficient groups or coefficient scanning order is based on an intra-prediction mode of the current video block. In one example, the coefficient scanning order is vertical, and wherein the intra-prediction mode is vertically dominant. In another example, the coefficient scanning order is horizontal, and wherein the intra-prediction mode is vertically dominant. For example, the index of the intra-prediction mode ranges from 34 to 66 or from 35 to 66.

[0459] In some embodiments, the context of the sign flag is based on neighboring information in a coefficient block associated with the current video block. In an example, the context of the sign flag is further based on the position of the coefficients of the coefficient block. In another example, the context of the sign flag is based on (x+y), min(x,y), or max(x,y), where x and y are the horizontal and vertical values ​​of the position of the coefficient, respectively.

[0460] Figure 9B A flow chart of yet another exemplary method for video processing is shown. Method 910 includes, at step 912, determining whether a chroma transform skip mode is applicable to a current video block.

[0461] The method 910 includes, at step 914 , performing, based on the determination, a conversion between the current video block and a bitstream representation of the video including the current video block.

[0462] In some embodiments, the determination is based on a message signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group, a codec unit (CU), or a video data unit.

[0463] In some embodiments, the determination is based on decode information from one or more representative video blocks that were decoded before performing the conversion, and wherein samples in each of the one or more representative video blocks and the current video block are based on common color information. In one example, the one or more representative video blocks include luma blocks or chroma blocks. In another example, the one or more representative video blocks include blocks within a collocated luma block.

[0464] Figure 9C A flow chart of yet another exemplary method for video processing is shown. Method 920 includes, at step 922, making a decision based on a condition regarding selectively applying a transform skip mode to the current video block during conversion between the current video block and a bitstream representation of a video including the current video block.

[0465] The method 920 includes, at step 924 , performing a conversion based on the determination.

[0466] In some embodiments, the condition is based on messages signaled by sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), picture header, slice header, slice group header, largest codec unit (LCU), LCU row, LCU group / codec unit (CU) / video data unit.

[0467] In some embodiments, the condition is based on one or more of: (1) a block size of the current video block or at least one neighboring block of the current video block, (2) a block shape of the current video block or at least one neighboring block, (3) a most probable mode of the current video block or at least one neighboring block, (4) a prediction mode of at least one neighboring block, (5) an intra-frame prediction mode of at least one neighboring block, (6) one or more motion vectors of at least one neighboring block, (7) an indication of a quantized residual block differential pulse codec modulation (QR-BDPCM) mode of at least one neighboring block, (8) a current quantization parameter (QP) of the current video block or at least one neighboring block, (9) an indication of a color format of the current video block, (10) a single codec tree structure or a dual codec tree structure associated with the current video block, (11) a slice type, a slice group type, or a picture type of the current video block, or (12) a temporal layer identifier (ID).

[0468] In some embodiments, the application of transform skip mode is performed, the prediction mode of the current video block is inter-block copy (IBC) mode, and the width or height of the current video block is compared to a threshold. In an example, the threshold is signaled in the bitstream representation. In another example, the threshold is 4, 8, 16, or 32.

[0469] In yet another example, the threshold is based on one or more of the following: (1) a message signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group, or a codec unit (CU), (2) a temporal layer identifier (ID), (3) a block size of the current video block or at least one neighboring block of the current video block, (4) a block shape of the current video block or at least one neighboring block, (5) a most probable mode of the current video block or at least one neighboring block, (6) prediction mode of at least one neighboring block, (7) an intra-frame prediction mode of at least one neighboring block, (8) one or more motion vectors of at least one neighboring block, (9) an indication of a quantized residual block differential pulse coding modulation (QR-BDPCM) mode of at least one neighboring block, (10) a current quantization parameter (QP) of the current video block or the at least one neighboring block, (11) an indication of a color format of the current video block, (12) a single codec tree structure or a dual codec tree structure associated with the current video block, or (13) a slice type, a slice group type, or a picture type of the current video block.

[0470] Figure 9D A flowchart of another exemplary method for video processing is shown. The method 930 includes, at step 932, making a decision regarding selective application of quantized residual block differential pulse codec modulation (QR-BDPCM) during conversion between a current video block and a bitstream representation of a video including the current video block based on an indication of a transform skip mode in the bitstream representation.

[0471] The method 930 includes, at step 934 , performing a conversion based on the determination.

[0472] In some embodiments, the indication of the transform skip mode is signaled at the transform unit (TU) level.

[0473] Figure 9E A flow chart of yet another exemplary method for video processing is shown. Method 940 includes, at step 942, making a decision regarding selective application of a single tree or a dual tree based on a condition during conversion between a current video block and a bitstream representation of a video including the current video block.

[0474] The method 940 includes, at step 944 , performing a conversion based on the determination.

[0475] In some embodiments, the condition is based on a message signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group, a codec unit (CU), or a video data unit.

[0476] In some embodiments, the condition is based on determining whether a slice, a slice, a largest codec unit (LCU), an LCU row, an LCU group, or a video data unit comprising the current video block is screen content. In an example, the determination is based on one or more of: (1) a message signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, an LCU, an LCU row, an LCU group, a codec unit (CU), or a video data unit, (2) a block size of the current video block or at least one neighboring block of the current video block, (3) a block shape of the current video block or at least one neighboring block, (4) a current quantization parameter (QP) of the current video block or at least one neighboring block, (5) an indication of a color format of the current video block, (6) a single codec tree structure or a dual codec tree structure associated with the current video block, (7) a slice type, a slice group type, or a picture type of the current video block, or (8) a temporal layer identifier (ID).

[0477] Figure 10 is a block diagram of a video processing device 1000. Device 1000 can be used to implement one or more of the methods described herein. Device 1000 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 1000 may include one or more processors 1002, one or more memories 1004, and video processing hardware 1006. (Multiple) processors 1002 can be configured to implement one or more methods described in this document (including but not limited to methods 900, 910, 920, 930, and 940). Memory (multiple memories) 1004 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 1006 can be used to implement some of the techniques described in this document in hardware circuits.

[0478] In some embodiments, the video encoding and decoding method can be used in Figure 10 The invention is implemented by means of the apparatus implemented on the described hardware platform.

[0479] In some embodiments, for example, as described in items 5 and 10 above, etc., a method of video processing includes: making a determination as to whether an intra block copy mode is applicable to conversion between a current video block of a video and a bitstream representation based on a type of a codec tree structure corresponding to the current video block; and performing conversion based on the determination.

[0480] In the above method, the bitstream representation excludes the indication of the intra block copy mode. In other words, the bitstream does not carry the display signaling of the IBC mode.

[0481] In the above method, the type of the codec tree structure is a dual codec tree structure, and it is determined that the intra block copy mode is not applicable.

[0482] Figure 11 1 is a block diagram illustrating an example video processing system 1100 in which the various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 1100. System 1100 may include an input 1102 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10 bit multi-component pixel values, or may be received in a compressed or encoded format. Input 1102 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0483] System 1100 may include a codec component 1104 that can implement the various codecs or encoding methods described in this document. The codec component 1104 can reduce the average bit rate of the video from the input 1102 to the output of the codec component 1104 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of the codec component 1104 can be stored or transmitted via a connected communication (as represented by component 1106). Component 1108 can use the bitstream (or codec) representation of the stored or communicated video received at the input 1102 to generate pixel values ​​or displayable video sent to the display interface 1110. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0484] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0485] Figure 121 is a flow chart of an example method for visual media encoding. The steps of the flow chart are discussed in conjunction with Example Embodiment 1 discussed in Section 4 of this document. At step 1202, the process determines a maximum allowed size for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, within which a current video block of the one or more video blocks is allowed to be encoded using a transform skip mode such that a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform. At step 1204, the process includes a syntax element in the bitstream representation indicating the maximum allowed size.

[0486] Figure 13 is a flowchart of an example method for visual media decoding. The steps of the flowchart are discussed in conjunction with Example Embodiment 1 discussed in Section 4 of this document. At step 1302, the process parses a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates a maximum allowed size within which a current video block of the one or more blocks of the video region is allowed to be encoded or decoded using a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform. At step 1304, the process generates a decoded video region from the bitstream representation by decoding the one or more video blocks according to the maximum allowed size.

[0487] Figure 14 is a flowchart of an example method for visual media processing. The steps of the flowchart are discussed in conjunction with Example Embodiment 2 discussed in Section 4 of this document. At step 1402, the process determines to encode and decode a current video block of visual media data using a transform skip mode. At step 1404, the process performs a conversion between the current video block and a bitstream representation of the visual media data based on the determination, wherein during the conversion, the current video block is divided into a plurality of coefficient groups and signaling of a codec block flag for at least one of the plurality of coefficient groups is excluded from the bitstream representation, wherein in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform.

[0488] Figure 15is a flowchart of an example method for visual media processing. The steps of the flowchart are discussed in conjunction with Example Embodiment 3 discussed in Section 4 of this document. At step 1502, the process determines to encode or decode a current video block of visual media data using a transform skip mode. At step 1504, the process performs a conversion between the current video block and a bitstream representation of the visual media data based on the determination, wherein during the conversion, the current video block is divided into a plurality of coefficient groups, wherein in the transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform, and further wherein during the conversion, a coefficient scanning order of the plurality of coefficient groups is determined based at least in part on an indication in the bitstream representation.

[0489] Figure 16 is a flow chart of an example method for visual media encoding. The steps of the flow chart are discussed in conjunction with Example Embodiment 4 discussed in Section 4 of this document. At step 1602, the process encodes a current video block in a video region of visual media data into a bitstream representation of the visual media data using a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform. At step 1604, the process selects a context of a sign flag for the current video block based on dividing the current video block into a plurality of coefficient groups and selecting a context of a sign flag for the current video block based on sign flags of one or more neighboring video blocks in the coefficient groups associated with the current video block.

[0490] Figure 17 is a flow chart of an example method for visual media decoding. The steps of the flow chart are discussed in conjunction with Example Embodiment 4 discussed in Section 4 of this document. At step 1702, the process parses a bitstream representation of visual media data including a video region including a current video block to identify a context for a sign flag used in a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform. At step 1704, the process generates a decoded video region from the bitstream representation such that the context for the sign flag is based on partitioning the current video block into a plurality of coefficient groups and based on sign flags of one or more neighboring video blocks in the coefficient groups associated with the current video block.

[0491] Figure 181 is a flow chart of an example method for visual media processing. At step 1802, the process determines a position of a current coefficient associated with the current video block of visual media data when the current video block is divided into a plurality of coefficient positions. At step 1804, the process derives a context for a sign flag of the current coefficient based at least on sign flags of one or more neighboring coefficients. At step 1806, the process generates a sign flag for the current coefficient based on the context, wherein the sign flag of the current coefficient is used in a transform skip mode in which the current video block is encoded or decoded without applying a transform.

[0492] Figure 19 is a flowchart of an example method for visual media encoding. The steps of the flowchart are discussed in conjunction with Example Embodiment 5 discussed in Section 4 of this document. At step 1902, to encode one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, the process determines, based on satisfying at least one rule, that a chroma transform skip mode is applicable to a current video block, wherein in chroma transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform. At step 1904, the process includes a syntax element in the bitstream representation indicating the chroma transform skip mode.

[0493] Figure 20 is a flow chart of an example method for visual media decoding. The steps of the flow chart are discussed in conjunction with Example Embodiment 5 discussed in Section 4 of this document. At step 2002, the process parses a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, the syntax element indicating that at least one rule associated with application of a chroma transform skip mode is satisfied, wherein in chroma transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform. At step 2004, the process generates a decoded video region from the bitstream representation by decoding the one or more video blocks.

[0494] Figure 212 is a flow chart of an example method for visual media encoding. The steps of the flow chart are discussed in conjunction with Example Embodiment 6 discussed in Section 4 of this document. At step 2102, to encode one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, the process makes a decision based on a condition regarding the selective application of a transform skip mode to a current video block. At step 2104, the process includes a syntax element in the bitstream representation indicating the condition, wherein in transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform.

[0495] Figure 22 is a flow chart of an example method for visual media decoding. The steps of the flow chart are discussed in conjunction with Example Embodiment 6 discussed in Section 4 of this document. At step 2202, the process parses a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates a condition related to the use of a transform skip mode in which a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation without applying a transform. At step 2204, the process generates a decoded video region from the bitstream representation by decoding the one or more video blocks according to the condition.

[0496] Figure 23 is a flowchart of an example method for visual media encoding. The steps of the flowchart are discussed in conjunction with Example Embodiment 8 discussed in Section 4 of this document. At step 2302, to encode one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, the process makes a decision regarding the selective application of a Quantized Residual Block Differential Pulse Coded Modulation (QR-BDPCM) technique based on an indication of a transform skip mode in the bitstream representation, wherein in the transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform, wherein in the QR-BDPCM technique, the residual of the prediction error in the horizontal direction and / or the vertical direction is quantized and entropy coded. At step 2304, an indication of the selective application of QR-BDPCM is included in the bitstream representation.

[0497] Figure 24is a flowchart of an example method for visual media decoding. The steps of the flowchart are discussed in conjunction with Example Embodiment 8 discussed in Section 4 of this document. At step 2402, the process parses a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates selective application of a Quantized Residual Block Differential Pulse Coded Modulation (QR-BDPCM) technique based on an indication of a transform skip mode in the bitstream representation, wherein in transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform, wherein in the QR-BDPCM technique, the residual of the prediction error in the horizontal and / or vertical directions is quantized and entropy coded. At step 2404, the process generates a decoded video region from the bitstream representation by decoding the one or more video blocks according to a maximum allowed size.

[0498] Figure 25 2 is a flow chart of an example method for visual media encoding. The steps of this flow chart are discussed in conjunction with Example Embodiment 9 discussed in Section 4 of this document. At step 2502, the process makes a conditional decision regarding the selective application of a single tree or a dual tree to encode one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data. At step 2504, the process includes a syntax element in the bitstream representation indicating the selective application of the single tree or the dual tree.

[0499] Figure 26 26 is a flowchart of an example method for visual media decoding. The steps of the flowchart are discussed in conjunction with Example Embodiment 9 discussed in Section 4 of this document. At step 2602, the process parses a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates selective application of a single tree or a dual tree based on a condition or inferred from a condition. At step 2604, the process generates a decoded video region from the bitstream representation by decoding the one or more video blocks according to the syntax element.

[0500] Some embodiments of this document are now presented in a clause-based format.

[0501] A1. A method for visual media encoding, comprising:

[0502] For encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, determining a maximum allowed size within which a current video block of the one or more video blocks is allowed to be encoded using a transform skip mode such that a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and

[0503] A syntax element indicating the maximum allowed size is included in the bitstream representation.

[0504] A2. A method for visual media decoding, comprising:

[0505] parsing a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates a maximum allowed size within which a current video block of the one or more blocks of the video region is allowed to be encoded and decoded using a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and

[0506] A decoded video region is generated from the bitstream representation by decoding one or more video blocks according to a maximum allowed size.

[0507] A3. A method as recited in any one or more of clauses A1-A2, wherein the maximum allowed size comprises a maximum allowed width and a maximum allowed height associated with a transform skip mode.

[0508] A4. A method according to clause A3, wherein the maximum allowed width and the maximum allowed height are signaled in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU) row or an LCU group.

[0509] A5. A method as recited in any one or more of clauses A1-A3, wherein the maximum allowed width and the maximum allowed height are signaled in different messages in the bitstream representation.

[0510] A6. A method according to any one or more of clauses A1-A2, wherein the initial values ​​of the maximum allowed width and the maximum allowed height are signaled in a sequence parameter set (SPS) or a picture parameter set (PPS), and further, wherein the updated values ​​of the maximum allowed width and the maximum allowed height are signaled in a picture header, a slice header, a slice group header, a largest codec unit (LCU) row or an LCU group.

[0511] A7. A method as recited in any one or more of clauses A2-A6, wherein generating the decoded video region comprises decoding one or more video blocks without using a transform skip mode.

[0512] A8. The method of any one or more of clauses A2-A6, wherein generating the decoded video region comprises decoding one or more video blocks based on use of a transform skip mode.

[0513] A9. A method for visual media processing, comprising:

[0514] Determining to use a transform skip mode to encode and decode a current video block of the visual media data; and

[0515] Based on the determination, a conversion between the current video block and a bitstream representation of the visual media data is performed, wherein during the conversion, the current video block is divided into a plurality of coefficient groups and signaling of a codec block flag for at least one of the plurality of coefficient groups is excluded in the bitstream representation, wherein in a transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform.

[0516] A10. The method of clause A9, wherein signaling of a codec block flag for each of the plurality of coefficient groups is excluded in the bitstream representation.

[0517] A11. A method according to any one or more of clauses A9-A10, wherein codec block flags of a plurality of coefficient groups excluded in the bitstream representation are inferred to be fixed values.

[0518] A12. A method according to clause A11, wherein the fixed value is 1.

[0519] A13. The method of any one or more of clauses A9-A12, further comprising:

[0520] A decision is made to selectively enable or disable signaling a codec block flag for a plurality of coefficient groups based on a prediction mode of a current video block.

[0521] A14. A method according to clause A13, wherein if the prediction mode is an intra prediction mode, signaling of codec block flags for the plurality of coefficient groups is excluded in the bitstream representation.

[0522] A15. A method according to any one or more of clauses A9-A14, wherein the codec block flag excluded in the signaling in the bitstream representation is inferred based on one or more of the following:

[0523] (1) Messages signaled by sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), picture header, slice header, slice group header, largest codec unit (LCU), LCU row, LCU group, or codec unit (CU),

[0524] (2) the position of at least one coefficient group among the plurality of coefficient groups,

[0525] (3) the block size of the current video block or at least one neighboring block of the current video block,

[0526] (4) the block shape of the current video block or at least one adjacent block,

[0527] (5) the most likely mode of the current video block or at least one neighboring block,

[0528] (6) prediction mode of at least one adjacent block,

[0529] (7) Intra-frame prediction mode of at least one neighboring block,

[0530] (8) one or more motion vectors of at least one adjacent block,

[0531] (9) an indication of the quantized residual block differential pulse coding modulation (QR-BDPCM) mode of at least one adjacent block,

[0532] (10) the current quantization parameter (QP) of the current video block or at least one adjacent block,

[0533] (11) an indication of the color format of the current video block,

[0534] (12) a single codec tree structure or a dual codec tree structure associated with the current video block, or

[0535] (13) The slice type, slice group type or picture type of the current video block.

[0536] A16. A method for visual media processing, further comprising:

[0537] Determining to use a transform skip mode to encode and decode a current video block of the visual media data; and

[0538] Based on the determination, a conversion between the current video block and a bitstream representation of the visual media data is performed, wherein during the conversion, the current video block is divided into a plurality of coefficient groups, wherein in a transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform, and wherein during the conversion, a coefficient scanning order of the plurality of coefficient groups is determined based at least in part on an indication in the bitstream representation.

[0539] A17. A method according to clause A16, wherein the coefficient scanning order is based on a message signaled by a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group or a codec unit (CU).

[0540] A18. The method of clause A17, wherein in transform skip mode, the plurality of coefficient groups or the coefficient scanning order is based on an intra-prediction mode of the current video block.

[0541] A19. The method of clause A18, wherein the coefficient scanning order is vertical, and wherein the intra prediction mode is horizontally dominant.

[0542] A20. The method of clause A18, wherein the coefficient scanning order is horizontal, and wherein the intra prediction mode is horizontally dominant.

[0543] A21. A method as described in any one or more of clauses A19-A20, wherein the index of the intra-frame prediction mode ranges from 2 to 33 or from 2 to 34.

[0544] A22. The method of clause A18, wherein the coefficient scanning order is vertical, and wherein the intra prediction mode is vertically dominant.

[0545] A23. The method of clause A18, wherein the coefficient scanning order is horizontal, and wherein the intra prediction mode is vertically dominant.

[0546] A24. A method as recited in any one or more of clauses A22-A23, wherein the index of the intra-prediction modes ranges from 34 to 66 or from 35 to 66.

[0547] A25. A method as recited in any one or more of clauses A9-A24, wherein converting comprises generating a bitstream representation from the current video block.

[0548] A65. A method as described in any one or more of clauses A9-A24, wherein converting includes generating pixel values ​​of the current video block from a bitstream representation.

[0549] C1. A method for visual media encoding, comprising:

[0550] To encode one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, determine that a chroma transform skip mode is applicable to a current video block based on satisfying at least one rule, wherein in the chroma transform skip mode, a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform;

[0551] A syntax element indicating the chroma transform skip mode is included in the bitstream representation.

[0552] C2. A method for visual media decoding, comprising:

[0553] parsing a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, the syntax element indicating that at least one rule associated with application of a chroma transform skip mode is satisfied, wherein in the chroma transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform; and

[0554] A decoded video region is generated from the bitstream representation by decoding one or more video blocks.

[0555] C3. A method according to any one or more of clauses C1-C2, wherein the determination is based on a message signaling a sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), picture header, slice header, slice group header, largest codec unit (LCU), LCU row, LCU group, codec unit (CU), or video data unit associated with a bitstream representation of the visual media data.

[0556] C4. A method as described in any one or more of clauses C2-C3, wherein the determination is based on decode information from one or more representative video blocks decoded before the conversion, and wherein the samples in each of the one or more representative video blocks and the current video block are based on a common color component.

[0557] C5. The method of clause C4, wherein the one or more representative video blocks comprise a luma block or a chroma block.

[0558] C6. The method of clause C4, wherein the one or more representative video blocks comprise blocks within a collocated luma block.

[0559] C7. The method of clause C4, wherein the current video block is a current chroma block, and wherein the one or more representative video blocks comprise neighboring chroma blocks of the current chroma block.

[0560] C8. The method of clause C4, wherein the current video block is a current chroma block, and wherein the one or more representative video blocks comprise blocks of corresponding luma samples covering a center chroma sample within the current chroma block.

[0561] C9. The method of clause C4, wherein the current video block is a current chroma block, and wherein the one or more representative video blocks include a block of corresponding luma samples covering a lower right chroma sample within the current chroma block.

[0562] D1. A method for visual media encoding, comprising:

[0563] For encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, making a determination based on a condition regarding selectively applying a transform skip mode to a current video block; and

[0564] A syntax element is included in the bitstream representation indicating a condition wherein, in a transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform.

[0565] D2. A method for visual media decoding, comprising:

[0566] parsing a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates a condition related to use of a transform skip mode in which a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation without applying a transform; and

[0567] A decoded video region is generated from a bitstream representation by decoding one or more video blocks according to a condition.

[0568] D3. A method according to clause D1, wherein the syntax element is signaled by a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU, a codec unit (CU), or a video data unit.

[0569] D4. A method according to any one or more of clauses D1-D3, wherein the condition is based on one or more of the following:

[0570] (1) the block size of the current video block or at least one neighboring block of the current video block,

[0571] (2) the block shape of the current video block or at least one adjacent block,

[0572] (3) the most likely mode of the current video block or at least one adjacent block,

[0573] (4) prediction mode of at least one adjacent block,

[0574] (5) Intra-frame prediction mode of at least one neighboring block,

[0575] (6) one or more motion vectors of at least one adjacent block,

[0576] (7) an indication of the quantized residual block differential pulse coding modulation (QR-BDPCM) mode of at least one adjacent block,

[0577] (8) the current quantization parameter (QP) of the current video block or at least one adjacent block,

[0578] (9) an indication of the color format of the current video block,

[0579] (10) a single codec tree structure or a dual codec tree structure associated with the current video block,

[0580] (11) The slice type, slice group type, or picture type of the current video block, or

[0581] (12) Time domain layer identification (ID).

[0582] D5. The method of clause D4, wherein transform skip mode is applied when the prediction mode of the current video block is inter-block copy (IBC) mode, and wherein a width or a height of the current video block satisfies a threshold.

[0583] D6. A method as recited in clause D5, wherein the threshold value is signaled in the bitstream representation.

[0584] D7. A method according to clause D6, wherein the value of the threshold is 4, 8, 16 or 32.

[0585] D8. A method according to any one or more of clauses D5-D7, wherein the threshold value is based on one or more of the following:

[0586] (1) Messages signaled by sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), picture header, slice header, slice group header, largest codec unit (LCU), LCU row, LCU group, or codec unit (CU),

[0587] (2) Time domain layer identification (ID),

[0588] (3) the block size of the current video block or at least one neighboring block of the current video block,

[0589] (4) the block shape of the current video block or at least one adjacent block,

[0590] (5) the most likely mode of the current video block or at least one neighboring block,

[0591] (6) prediction mode of at least one adjacent block,

[0592] (7) Intra-frame prediction mode of at least one neighboring block,

[0593] (8) one or more motion vectors of at least one adjacent block,

[0594] (9) an indication of the quantized residual block differential pulse coding modulation (QR-BDPCM) mode of at least one adjacent block,

[0595] (10) the current quantization parameter (QP) of the current video block or at least one adjacent block,

[0596] (11) an indication of the color format of the current video block,

[0597] (12) a single codec tree structure or a dual codec tree structure associated with the current video block, or

[0598] (13) The slice type, slice group type or picture type of the current video block.

[0599] E1. A method for visual media encoding, comprising:

[0600] for encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, making a decision regarding selective application of a Quantized Residual Block Differential Pulse Coded Modulation (QR-BDPCM) technique based on an indication of a transform skip mode in the bitstream representation, wherein in the transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform, wherein in the QR-BDPCM technique, the residual of the prediction error in a horizontal direction and / or a vertical direction is quantized and entropy coded; and

[0601] An indication of the selective application of the QR-BDPCM technique is included in the bitstream representation.

[0602] E2. A method for visual media decoding, comprising:

[0603] parsing a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates selective application of a Quantized Residual Block Differential Pulse Coded Modulation (QR-BDPCM) technique based on an indication of a transform skip mode in the bitstream representation, wherein in the transform skip mode, a residual of a prediction error between a current video block and a reference video block is represented in the bitstream representation of the visual media data without applying a transform, wherein in the QR-BDPCM technique, the residual of the prediction error in a horizontal direction and / or a vertical direction is quantized and entropy coded; and

[0604] A decoded video region is generated from the bitstream representation by decoding one or more video blocks according to a maximum allowed size.

[0605] E3. A method as recited in any one or more of clauses E1-E2, wherein the indication of the transform skip mode is signaled at a transform unit (TU) level associated with the current video block.

[0606] E4. A method according to any one or more of clauses E1-E3, wherein in the bitstream representation, a syntax element indicating application of the QR-BDPCM technique on the current video block is included after the indication of the transform skip mode.

[0607] E5. A method according to clause E4, wherein a first flag in the bitstream representation is used to indicate a transform skip mode and a second flag in the bitstream representation is used to indicate selective application of a QR-BDPCM technique.

[0608] E6. The method of clause E5, wherein the third flag is used to indicate application of the QR-BDPCM technique based on the horizontal prediction direction type, and the fourth flag is used to indicate application of the QR-BDPCM technique based on the vertical prediction direction type.

[0609] E7. A method according to clause E5, wherein the value of the second flag is inferred from the value of the first flag.

[0610] E8. A method as recited in clause E5, wherein the value of the first flag is not signaled but inferred from the value of the second flag.

[0611] E9. A method as recited in clause E8, wherein the value of the first flag is not signaled, and the value of the first flag is inferred to be true when the value of the second flag is true.

[0612] E10. The method of clause E7, wherein the first flag and the second flag are Boolean values, further comprising:

[0613] Upon determining that the value of the first flag is Boolean true, the value of the second flag is inferred to be Boolean true.

[0614] E11. The method of clause E7, wherein the first flag and the second flag are Boolean values, further comprising:

[0615] When the value of the first flag is determined to be Boolean false, the value of the second flag is inferred to be Boolean false.

[0616] F1. A method for visual media encoding, comprising:

[0617] For encoding one or more video blocks in a video region of visual media data into a bitstream representation of the visual media data, making a decision regarding selective application of a single tree or a dual tree based on a condition; and

[0618] A syntax element indicating the selective application of a single tree or a dual tree is included in the bitstream representation.

[0619] F2. A method for visual media decoding, comprising:

[0620] parsing a syntax element from a bitstream representation of visual media data comprising a video region, the video region comprising one or more video blocks, wherein the syntax element indicates selective application of a single tree or a dual tree based on or inferred from a condition; and

[0621] A decoded video region is generated from the bitstream representation by decoding one or more video blocks according to syntax elements.

[0622] F3. A method according to any one or more of clauses F1-F2, wherein the condition is based on a message signaled by a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, a slice group header, a largest codec unit (LCU), an LCU row, an LCU group, a codec unit or a video data unit.

[0623] F4. A method according to any one or more of clauses F1-F2, wherein the condition is based on determining whether a slice, a slice, a largest codec unit (LCU), an LCU row, an LCU group or a video data unit comprising the current video block is screen content.

[0624] F5. A method according to clause F3, wherein the condition includes one or more of the following:

[0625] (1) Messages signaled by a sequence parameter set (SPS), video parameter set (VPS), picture parameter set (PPS), picture header, slice header, slice group header, LCU, LCU row, LCU group, codec unit (CU), or video data unit,

[0626] (2) the block size of the current video block or at least one neighboring block of the current video block,

[0627] (3) the block shape of the current video block or at least one adjacent block,

[0628] (4) the current quantization parameter (QP) of the current video block or at least one adjacent block,

[0629] (5) an indication of the color format of the current video block,

[0630] (6) a single codec tree structure or a dual codec tree structure associated with the current video block,

[0631] (7) The slice type, slice group type, or picture type of the current video block, or

[0632] (8) Time domain layer ID.

[0633] F6. A video encoder apparatus comprising a processor configured to implement the method of any one or more of clauses A1-F5.

[0634] F7. A video decoder device comprising a processor configured to implement the method of any one or more of clauses A1-F5.

[0635] F8. A computer-readable medium having stored thereon code comprising processor-executable instructions for implementing the method of any one or more of clauses A1-F5.

[0636] B1. A method for visual media encoding, comprising:

[0637] encoding a current video block in a video region of the visual media data into a bitstream representation of the visual media data using a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and

[0638] A context for a sign flag of the current video block is selected based on partitioning the current video block into a plurality of coefficient groups, according to sign flags of one or more neighboring video blocks in the coefficient groups associated with the current video block.

[0639] B2. A method for visual media decoding, comprising:

[0640] parsing a bitstream representation of visual media data including a video region including a current video block to identify a context of a symbol flag used in a transform skip mode in which a residual of a prediction error between the current video block and a reference video block is represented in the bitstream representation without applying a transform; and

[0641] A decoded video region is generated from the bitstream representation such that a context of a sign flag is based on partitioning a current video block into a plurality of coefficient groups according to sign flags of one or more neighboring video blocks in the coefficient groups associated with the current video block.

[0642] B3. The method of any one or more of clauses B1-B2, wherein the one or more neighboring video blocks of the current video block include: a left neighbor and / or an upper neighbor and / or an upper-left neighbor and / or an upper-right neighbor.

[0643] B4. A method according to any one or more of clauses B1-B3, wherein the relationship between the context of the sign flag of the current video block and the sign flags of one or more adjacent video blocks is represented by:

[0644] C=(L+A), where C is the context ID of the context of the sign flag of the current video block, L is the sign flag of the left neighbor of the current video block, and A is the sign flag of the upper neighbor of the current video block.

[0645] B5. A method according to any one or more of clauses B1-B3, wherein the relationship between the context of the sign flag of the current video block and the sign flags of one or more adjacent video blocks is represented by:

[0646] C=(L+A*2), where C is the context ID of the context of the sign flag of the current video block, L is the sign flag of the left neighbor of the current video block, and A is the sign flag of the upper neighbor of the current video block.

[0647] B6. A method according to any one or more of clauses B1-B3, wherein the relationship between the context of the sign flag of the current video block and the sign flags of one or more adjacent video blocks is represented by:

[0648] C=(L*2+A), where C is the context ID of the context of the sign flag of the current video block, L is the sign flag of the left neighbor of the current video block, and A is the sign flag of the upper neighbor of the current video block.

[0649] B7. A method according to any one or more of clauses B1-B3, wherein the context of the sign flag of the current video block is based on the sign flags of one or more adjacent video blocks according to the following rules:

[0650] If the sign flags of a pair of adjacent video blocks are both negative, then the rule specifies that the first context of the sign flag is selected.

[0651] B8. A method for visual media processing, comprising:

[0652] determining a position of a current coefficient associated with the current video block of visual media data when the current video block is divided into a plurality of coefficient positions;

[0653] Derives a context for the sign of the current coefficient based at least on the sign of one or more adjacent coefficients; and

[0654] A sign flag of a current coefficient is generated based on the context, wherein the sign flag of the current coefficient is used in a transform skip mode in which a current video block is encoded and decoded without applying a transform.

[0655] B9. The method of clause B8, wherein the one or more neighboring coefficients of the current video block include: a left neighbor and / or an upper neighbor.

[0656] B10. A method according to clause B8, wherein the context of the sign flag of the current coefficient is also based on the position of the current coefficient.

[0657] B11. The method of clause B8, wherein the visual media processing comprises decoding the current video block from a bitstream representation.

[0658] B12. The method of clause B8, wherein the visual media processing comprises encoding the current video block into a bitstream representation.

[0659] B13. A method according to any one or more of clauses B1-B3 or B8-B12, wherein the context of the sign flag of the current video block is based on the sign flags of one or more adjacent video blocks according to the following rules:

[0660] If only one sign flag in a pair of adjacent video blocks is negative, then the rule specifies that the second context of the sign flag is selected.

[0661] B14. A method according to any one or more of clauses B1-B3 or B8-B12, wherein the context of the sign flag of the current video block is based on the sign flags of one or more adjacent video blocks according to the following rules:

[0662] If the sign flags of a pair of adjacent video blocks are both positive, then the rule specifies that the third context of the sign flag is selected.

[0663] B15. A method as described in any one or more of clauses B1-B3 or B8-B12, wherein the context of the sign flag of the current video block is further based on the position of the coefficient group in the plurality of coefficient groups.

[0664] B16. A method according to clause B10, wherein the context of the sign flag is based on at least one of the following: (x+y), min(x,y) or max(x,y), where x and y are the horizontal and vertical values ​​of the position of the coefficient group, respectively.

[0665] B17. A video encoder apparatus comprising a processor configured to implement the method of any one or more of clauses B1-B16.

[0666] B18. A video decoder device comprising a processor configured to implement the method of any one or more of clauses B1-B16.

[0667] B19. A computer-readable medium having stored thereon code comprising processor-executable instructions for implementing the method of any one or more of clauses B1-B19.

[0668] In this document, the terms "video processing" or "visual media processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits co-located or distributed across different locations in the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on error residual values ​​from transforms and codecs, and may also be encoded using bits in headers and other fields in the bitstream.

[0669] From the foregoing, it will be appreciated that specific embodiments of the presently disclosed technology have been described herein for illustrative purposes, but that various modifications may be made without departing from the scope of the present invention. Accordingly, the presently disclosed technology is not to be limited, except as in the appended claims.

[0670] The subject matter and implementations of the functional operations described in this patent document can be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more thereof. The implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by a data processing device or for controlling the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material combination that implements a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing unit" or "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus can include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof.

[0671] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0672] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0673] By way of example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. Examples include semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0674] This specification and drawings are to be regarded as exemplary only, where exemplary means example. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, the use of "or" is intended to include "and / or" unless the context clearly indicates otherwise.

[0675] Although this patent document contains many details, these should not be construed as limitations on any subject matter or the scope of what is claimed, but rather as descriptions of features specific to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, alone or in any suitable subcombination. Furthermore, although the features described above may be described as functioning in certain combinations, or even initially claimed to be protected as such, in some cases one or more features in the combination may be deleted from the claimed combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0676] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that these operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0677] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

[0678] 5. Additional Embodiments of the Disclosed Technology

[0679] 5.1 Binary Tree Related Examples

[0680] Changes at the top of the draft provided by JVET-N1001-v7 are highlighted in bold, underlined, and italic text. Uppercase bold text appearing below and elsewhere in this document indicates text that may be deleted from the VVC standard.

[0681] 5.1.1 Example #1

[0682]

[0683] 5.1.2 Example #2

[0684] 7.3.2.3 Sequence Singular Set RBSP Syntax

[0685]

[0686] 5.1.3 Example #3

[0687] 7.3.5 Strip Header Syntax

[0688] 7.3.5.1 General Strip Header Syntax

[0689]

[0690] sps_asdt_flag equal to 1 specifies that the proposed method can be enabled on the current sequence. It is equal to 0 and specifies that the proposed method cannot be enabled on the current sequence. When sps_ASDT_flag is not present, it is inferred to be equal to 0.

[0691] slice_asdt_flag equal to 1 specifies that the proposed method can be enabled on the current slice. It is equal to 0 and specifies that the proposed method cannot be enabled on the current slice. When slice_asdt_flag is not present, it is inferred to be equal to 0.

[0692] 7.3.7.2 Codec Tree Unit Syntax

[0693]

[0694] ctu_dual_tree_intra_flag equal to 1 specifies that the current CTU adopts a dual-tree codec structure. It is equal to 0 and specifies that the current CTU adopts a single-tree codec structure. When ctu_dual_tree_intra_flag is not present, it is inferred to be equal to 0.

[0695] 5.1.4 Example #4

[0696] 7.3.5 Strip Header Syntax

[0697] 7.3.5.1 General Strip Header Syntax

[0698]

[0699] sps_asdt_flag equal to 1 specifies that the proposed method can be enabled on the current sequence. It is equal to 0 and specifies that the proposed method cannot be enabled on the current sequence. When sps_ASDT_flag is not present, it is inferred to be equal to 0.

[0700] slice_asdt_flag equal to 1 specifies that the proposed method can be enabled on the current slice. It is equal to 0 and specifies that the proposed method cannot be enabled on the current slice. When slice_asdt_flag is not present, it is inferred to be equal to 0.

[0701] 7.3.7.2 Codec Tree Unit Syntax

[0702]

[0703] ctu_dual_tree_intra_flag equal to 1 specifies that the current CTU adopts a dual-tree codec structure. It is equal to 0 and specifies that the current CTU adopts a single-tree codec structure. When ctu_dual_tree_intra_flag is not present, it is inferred to be equal to 0.

[0704] 5.2 IBC-Related Examples

[0705] 7.3.7.5 Codec unit syntax

[0706]

[0707] 8.6 Decoding Process of Codec Units Encoding and Decoding in IBC Prediction Mode

[0708] 8.6.1 General decoding process for codecs encoding and decoding in IBC prediction mode

[0709] The inputs to this process are:

[0710] - Luma position (xCb, yCb), specifies the upper left corner sample of the current codec block relative to the upper left corner luminance sample of the current picture,

[0711] -The variable cbWidth specifies the width of the luminance sample of the current codec block,

[0712] -The variable cbHeight specifies the height of the luminance sample of the current codec block,

[0713] - The variable treeType specifies whether a single tree or a dual tree is used, and if a dual tree is used, whether the current tree corresponds to the luma or chroma components.

[0714] The output of this process is the modified reconstructed picture before loop filtering.

[0715] The derivation process for the quantization parameters specified in clause 8.7.1 is called with the luma position (xCb, yCb), the width cbWidth and height cbHeight of the luma samples of the current codec block, and the variable treeType as input.

[0716] The decoding process of the codec unit in IBC prediction mode consists of the following ordered steps:

[0717] 1. The motion vector components of the current codec unit are derived as follows:

[0718] 1. If TREETYPE is equal to SINGLE_TREE or DUAL_TREE_LUMA, then the following procedure applies:

[0719] - Call the derivation process of motion vector components specified in clause 8.6.2.1 with the luma codec block position (xCb, yCb), luma codec block width cbWidth and luma codec block height cbHeight as input, and output is the luma motion vector mvL[0][0].

[0720] - When treeType is equal to SINGLE_TREE, the derivation process of the chroma motion vector in clause 8.6.2.5 is called with the luma motion vector mvL[0][0] as input, and the output is the chroma motion vector mvC[0][0].

[0721] - The number of luma codec sub-blocks numSbX in the horizontal direction and the number of luma codec sub-blocks numSbY in the vertical direction are both set equal to 1.

[0722] 1. Otherwise, if TREETYPE is equal to DUAL_TREE_CHROMA, then the following process applies:

[0723] -The number of luma codec sub-blocks NUMSBX in the horizontal direction and the number of luma codec sub-blocks NUMSBY in the vertical direction are derived as follows:

[0724] NUMSBX=(CBWIDTH>>2) (8-871)

[0725] NUMSBY=(CBHEIGHT>>2) (8-872)

[0726] - For XSBIDX=0..NUMSBX-1, YSBIDX=0..NUMSBY-1, the chroma motion vector MVC[XSCBIDX][YSBIDX] is derived as follows:

[0727] - Luma motion vector MVL[XSBIDX][YSBIDX] is derived as follows:

[0728] - The position of the collocated luma codec unit (XCUY, YCUY) is derived as follows:

[0729] XCUY=XCB+XSBIDX*4 (8-873)

[0730] YCUY=YCB+YSBIDX*4 (8-874)

[0731] - If CUPREDMODE[XCUY][YCUY] is equal to MODE_INTRA, then the following applies:

[0732] MVL[XSBIDX][YSBIDX][0]=0 (8-875)

[0733] MVL[XSBIDX][YSBIDX][1]=0 (8-876)

[0734] PREDFLAGL0[XSBIDX][YSBIDX]=0 (8-877)

[0735] PREDFLAGL1[XSBIDX][YSBIDX]=0 (8-878)

[0736] Otherwise (CUPREDMODE[XCUY][YCUY] is equal to MODE_IBC), the following applies:

[0737] MVL[XSBIDX][YSBIDX][0]=MVL0[XCUY][YCUY][0] (8-879)

[0738] MVL[XSBIDX][YSBIDX][1]=MVL0[XCUY][YCUY][1] (8-880)

[0739] PREDFLAGL0[XSBIDX][YSBIDX]=1 (8-881)

[0740] PREDFLAGL1[XSBIDX][YSBIDX]=0 (8-882)

[0741] - Invoke the derivation process for chroma motion vectors in clause 8.6.2.5 with MVL[XSBIDX][YSBIDX] as input and output MVC[XSBIDX][YSBIDX].

[0742] - Bitstream conformance requires that the chroma motion vector MVC [XSCBIDX][YSBIDX] shall obey the following constraints:

[0743] - When the derivation procedure for block availability specified in clause 6.4.X [ED.(BB): Neighboring Block Availability Check Procedure TBD] is called with the current chroma position (XCURR, YCURR) set equal to (XCB / SUBWIDTHC, YCB / SUBHEIGHTC) and the neighboring chroma positions (XCB / SUBWIDTHC + (MVC[XSBIDX][YSBIDX][0]>>5), YCB / SUBHEIGHTC + (MVC[XSBIDX][YSBIDX][1]>>5)) as input, the output shall be equal to true.

[0744] - When the derivation procedure for block availability specified in clause 6.4.X [ED.(BB): Neighboring Block Availability Check Procedure TBD] is called with the current chroma position (XCURR, YCURR) set equal to (XCB / SUBWIDTHC, YCB / SUBHEIGHTC) and the neighboring chroma position (XCB / SUBWIDTHC + (MVC[XSBIDX][YSBIDX][0]>>5) + CBWIDTH / SUBWIDTHC-1, YCB / SUBHEIGHTC + (MVC[XSBIDX][YSBIDX][1]>>5) + CBHEIGHT / SUBHEIGHTC-1), the output shall be equal to true.

[0745] - One or both of the following conditions should be true:

[0746] -(MVC[XSBIDX][YSBIDX][0]>>5)+XSBIDX*2+2 is less than or equal to 0.

[0747] -(MVC[XSBIDX][YSBIDX][1]>>5)+YSBIDX*2+2 is less than or equal to 0.

[0748] 2. The predicted sample points of the current codec unit are derived as follows:

[0749] If TREETYPE is SINGLE_TREE or DUAL_TREE_LUMA, the prediction samples for the current codec unit are derived as follows:

[0750] Invoke the decoding process of the ibc block specified in clause 8.6.3.1 with as input the luma codec block position (xCb, yCb), the luma codec block width cbWidth and the luma codec block height cbHeight, the number of luma codec sub-blocks in the horizontal direction numSbX and the number of luma codec sub-blocks in the vertical direction numSbY, the luma motion vector mvL[xSbIdx][ySbIdx] (xSbIdx = 0..numSbX–1 and ySbIdx = 0..numSbY–1), and the variable cIdx set equal to 0, and output is an array predSamples of (cbWidth)x(cbHeight) that is the predicted luma samples L The ibc prediction sample (predSample).

[0751] Otherwise, if treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the prediction samples for the current codec unit are derived as follows:

[0752] Invoke the decoding process of the ibc block specified in clause 8.6.3.1 with as input the luma codec block position (xCb, yCb), the luma codec block width cbWidth and the luma codec block height cbHeight, the number of luma codec sub-blocks in the horizontal direction numSbX and the number of luma codec sub-blocks in the vertical direction numSbY, the chroma motion vector mvC[xSbIdx][ySbIdx] (xSbIdx = 0..numSbX–1 and ySbIdx = 0..numSbY–1), and the variable cIdx set equal to 1, and output is an array predSamples of (cbWidth / 2)x(cbHeight / 2) that is the predicted chroma samples for the chroma component Cb Cb The ibc prediction sample (predSample).

[0753] Invoke the decoding process of the ibc block specified in clause 8.6.3.1 with as input the luma codec block position (xCb, yCb), the luma codec block width cbWidth and the luma codec block height cbHeight, the number of luma codec sub-blocks in the horizontal direction numSbX and the number of luma codec sub-blocks in the vertical direction numSbY, the chroma motion vector mvC[xSbIdx][ySbIdx] (xSbIdx = 0..numSbX–1 and ySbIdx = 0..numSbY–1), and the variable cIdx set equal to 2, and output is an array predSamples of (cbWidth / 2)x(cbHeight / 2) that is the predicted chroma samples of the chroma component Cr Cr The ibc prediction sample (predSample).

[0754] 3. The variables NumSbX[xCb][yCb] and NumSbY[xCb][yCb] are set equal to numSbX and numSbY respectively.

[0755] 4. The residual sample points of the current codec unit are derived as follows:

[0756] - when TREETYPE is equal to SINGLE_TREE or TREETYPE is equal to DUAL_TREE_LUMA, the decoding process for the residual signal of a codec block coded in inter prediction mode as specified in clause 8.5.8 is called with as input the position (xTb0, yTb0) set equal to the luma position (xCb, yCb), the width nTbW set equal to the luma codec block width cbWidth, the height nTbH set equal to the luma codec block height cbHeight, and the variable cIdxset equal to 0, and the output is the array resSamplesL .

[0757] - When treeType is equal to SINGLE_TREE or TREETYPE is equal to DUAL_TREE_CHROMA, the decoding process for the residual signal of a codec block coded in inter prediction mode as specified in clause 8.5.8 is called with as input the position (xTb0, yTb0) set equal to the chroma position (xCb / 2, yCb / 2), the width nTbW set equal to the chroma codec block width cbWidth / 2, the height nTbH set equal to the chroma codec block height cbHeight / 2, and the variable cIdxset equal to 1, and the output is the array resSamples Cb .

[0758] - When treeType is equal to SINGLE_TREE or TREETYPE is equal to DUAL_TREE_CHROMA, the decoding process for the residual signal of a codec block coded in inter prediction mode as specified in clause 8.5.8 is called with as input the position (xTb0, yTb0) set equal to the chroma position (xCb / 2, yCb / 2), the width nTbW set equal to the chroma codec block width cbWidth / 2, the height nTbH set equal to the chroma codec block height cbHeight / 2, and the variable cIdxset equal to 2, and the output is the array resSamples Cr .

[0759] 5. The reconstructed sample points of the current codec unit are derived as follows:

[0760] - When TREETYPE is equal to SINGLE_TREE or TREETYPE is equal to DUAL_TREE_LUMA, with the block position (xB, yB) set equal to (xCb, yCb), the block width bWidth set equal to cbWidth, the block height bHeight set equal to cbHeight, the variable cIdx set equal to 0, and the variable predSamples set equal to L The (cbWidth) x (cbHeight) array of predSample is set equal to resSamples L The (cbWidth)x(cbHeight) array resSample is used as input to invoke the picture reconstruction process for color components as specified in clause 8.7.5, and the output is the modified reconstructed picture before loop filtering.

[0761] - When treeType is equal to SINGLE_TREE or TREETYPE is equal to DUAL_TREE_CHROMA, with the block position (xB, yB) set equal to (xCb / 2, yCb / 2), the block width bWidth set equal to cbWidth / 2, the block height bheight set equal to cbHeight / 2, the variable cIdx set equal to 1, and the variable preSamples set equal to 0. Cb The (cbWidth / 2)x(cbHeight / 2) array predSample is set and set equal to resSamples Cb The (cbWidth / 2)x(cbHeight / 2) array resSample is used as input to invoke the picture reconstruction process for color components as specified in Clause 8.7.5, and the output is the modified reconstructed picture before loop filtering.

[0762] - When treeType is equal to SINGLE_TREE or TREETYPE is equal to DUAL_TREE_CHROMA, with the block position (xB, yB) set equal to (xCb / 2, yCb / 2), the block width bWidth set equal to cbWidth / 2, the block height bheight set equal to cbHeight / 2, the variable cIdx set equal to 2, and the variable predSamples set equal to Cr The (cbWidth / 2)x(cbHeight / 2) array predSample is set and set equal to resSamples Cr The (cbWidth / 2)x(cbHeight / 2) array resSample is used as input to invoke the picture reconstruction process for color components as specified in Clause 8.7.5, and the output is the modified reconstructed picture before loop filtering.

Claims

1. A method for processing video data, comprising: utilizing a transform skip mode for converting between a current video block in a video region of video data and a bitstream of the video data for the current video block, wherein the current video block is divided into a plurality of coefficient groups, and wherein in the transform skip mode, a transform is skipped for a prediction residual between the current video block and a reference video block; selecting a context for a sign flag of a current coefficient of the current video block based on one or more sign flags of one or more neighboring coefficients in the current video block; as well as performing said conversion based on said selection, The one or more adjacent coefficients in the current video block are: a left adjacent coefficient and an upper adjacent coefficient, The context of the sign flag of the current coefficient is selected based on the sign flags of the left adjacent coefficient and the upper adjacent coefficient according to the following rules: If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both negative, then the rule specifies that the first context is selected, If the sign flag of one of the left neighbor coefficient and the above neighbor coefficient is negative and the sign flag of the other is positive, then the rule specifies that a second context is selected, If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both positive, then the rule specifies that a third context is selected, wherein a first syntax element indicating whether the transform skip mode is applied to the current video block is conditionally included in the bitstream, wherein whether the first syntax element is included in the bitstream is based on allowed intra prediction modes in a quantized residual block differential pulse coding modulation codec block and use of quantized residual block differential pulse coding modulation, and when the indication of the quantized residual block differential pulse coding modulation mode is 1, it is inferred that the transform skip mode is enabled, Wherein, when the indication of the differential pulse codec modulation mode of the quantized residual block is 1, it is inferred that the transform skip mode is enabled, and is applied based on the block size of the current video block and the single / dual codec tree structure of the current video block.

2. The method according to claim 1, wherein A second syntax element indicating a maximum allowed block size for the transform skip mode is included in the bitstream.

3. The method according to claim 1, wherein The converting includes encoding the current video block into the bitstream.

4. The method according to claim 1, wherein The converting includes decoding the current video block from the bitstream.

5. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: utilizing a transform skip mode for converting between a current video block in a video region of video data and a bitstream of the video data for the current video block, wherein the current video block is divided into a plurality of coefficient groups, and wherein in the transform skip mode, a transform is skipped for a prediction residual between the current video block and a reference video block; selecting a context for a sign flag of a current coefficient of the current video block based on one or more sign flags of one or more neighboring coefficients in the current video block; as well as performing said conversion based on said selection, The one or more adjacent coefficients in the current video block are: a left adjacent coefficient and an upper adjacent coefficient, The context of the sign flag of the current coefficient is selected based on the sign flags of the left adjacent coefficient and the upper adjacent coefficient according to the following rules: If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both negative, then the rule specifies that the first context is selected, If the sign flag of one of the left neighbor coefficient and the above neighbor coefficient is negative and the sign flag of the other is positive, then the rule specifies that a second context is selected, If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both positive, then the rule specifies that a third context is selected, wherein a first syntax element indicating whether the transform skip mode is applied to the current video block is conditionally included in the bitstream, wherein whether the first syntax element is included in the bitstream is based on allowed intra prediction modes in a quantized residual block differential pulse coding modulation codec block and use of quantized residual block differential pulse coding modulation, and when the indication of the quantized residual block differential pulse coding modulation mode is 1, it is inferred that the transform skip mode is enabled, Wherein, when the indication of the differential pulse codec modulation mode of the quantized residual block is 1, it is inferred that the transform skip mode is enabled, and is applied based on the block size of the current video block and the single / dual codec tree structure of the current video block.

6. A non-transitory computer-readable storage medium storing instructions that cause a processor to: utilizing a transform skip mode for converting between a current video block in a video region of video data and a bitstream of the video data for the current video block, wherein the current video block is divided into a plurality of coefficient groups, and wherein in the transform skip mode, a transform is skipped for a prediction residual between the current video block and a reference video block; selecting a context for a sign flag of a current coefficient of the current video block based on one or more sign flags of one or more neighboring coefficients in the current video block; as well as performing said conversion based on said selection, The one or more adjacent coefficients in the current video block are: a left adjacent coefficient and an upper adjacent coefficient, The context of the sign flag of the current coefficient is selected based on the sign flags of the left adjacent coefficient and the upper adjacent coefficient according to the following rules: If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both negative, then the rule specifies that the first context is selected, If the sign flag of one of the left neighbor coefficient and the above neighbor coefficient is negative and the sign flag of the other is positive, then the rule specifies that a second context is selected, If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both positive, then the rule specifies that a third context is selected, wherein a first syntax element indicating whether the transform skip mode is applied to the current video block is conditionally included in the bitstream, wherein whether the first syntax element is included in the bitstream is based on allowed intra prediction modes in a quantized residual block differential pulse coding modulation codec block and use of quantized residual block differential pulse coding modulation, and when the indication of the quantized residual block differential pulse coding modulation mode is 1, it is inferred that the transform skip mode is enabled, Wherein, when the indication of the differential pulse codec modulation mode of the quantized residual block is 1, it is inferred that the transform skip mode is enabled, and is applied based on the block size of the current video block and the single / dual codec tree structure of the current video block.

7. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein The method comprises: using a transform skip mode for a current video block in a video region of video data, wherein the current video block is divided into a plurality of coefficient groups, and wherein in the transform skip mode, a transform is skipped on a prediction residual between the current video block and a reference video block; Selecting a context for a sign flag of a current coefficient of the current video block based on one or more sign flags of one or more neighboring coefficients in the current video block; and generating a bitstream of the video data based on the selection, The one or more adjacent coefficients in the current video block are: a left adjacent coefficient and an upper adjacent coefficient, The context of the sign flag of the current coefficient is selected based on the sign flags of the left adjacent coefficient and the upper adjacent coefficient according to the following rules: If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both negative, then the rule specifies that the first context is selected, If the sign flag of one of the left neighbor coefficient and the above neighbor coefficient is negative and the sign flag of the other is positive, then the rule specifies that a second context is selected, If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both positive, then the rule specifies that a third context is selected, wherein a first syntax element indicating whether the transform skip mode is applied to the current video block is conditionally included in the bitstream, wherein whether the first syntax element is included in the bitstream is based on allowed intra prediction modes in a quantized residual block differential pulse coding modulation codec block and use of quantized residual block differential pulse coding modulation, and when the indication of the quantized residual block differential pulse coding modulation mode is 1, it is inferred that the transform skip mode is enabled, Wherein, when the indication of the differential pulse codec modulation mode of the quantized residual block is 1, it is inferred that the transform skip mode is enabled, and is applied based on the block size of the current video block and the single / dual codec tree structure of the current video block.

8. A method for storing a bitstream of a video, comprising: using a transform skip mode for a current video block of a video region of video data, wherein the current video block is divided into a plurality of coefficient groups, and wherein in the transform skip mode, a transform is skipped for a prediction residual between the current video block and a reference video block; selecting a context for a sign flag of a current coefficient of the current video block based on one or more sign flags of one or more neighboring coefficients in the current video block; generating a bitstream of the video data based on the selection; as well as storing the generated bitstream on a non-transitory computer-readable recording medium, The one or more adjacent coefficients in the current video block are: a left adjacent coefficient and an upper adjacent coefficient, The context of the sign flag of the current coefficient is selected based on the sign flags of the left adjacent coefficient and the upper adjacent coefficient according to the following rules: If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both negative, then the rule specifies that the first context is selected, If the sign flag of one of the left neighbor coefficient and the above neighbor coefficient is negative and the sign flag of the other is positive, then the rule specifies that a second context is selected, If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both positive, then the rule specifies that a third context is selected, wherein a first syntax element indicating whether the transform skip mode is applied to the current video block is conditionally included in the bitstream, wherein whether the first syntax element is included in the bitstream is based on allowed intra prediction modes in a quantized residual block differential pulse coding modulation codec block and use of quantized residual block differential pulse coding modulation, and when the indication of the quantized residual block differential pulse coding modulation mode is 1, it is inferred that the transform skip mode is enabled, Wherein, when the indication of the differential pulse codec modulation mode of the quantized residual block is 1, it is inferred that the transform skip mode is enabled, and is applied based on the block size of the current video block and the single / dual codec tree structure of the current video block.

9. A method for visual media processing, comprising: determining a position of a current coefficient associated with the current video block of visual media data when the current video block is divided into a plurality of coefficient positions; Derives a context for the sign of the current coefficient based at least on the sign of one or more adjacent coefficients; as well as generating a sign flag of a current coefficient based on the context, wherein the sign flag of the current coefficient is used in a transform skip mode in which the current video block is encoded and decoded without applying a transform, The one or more adjacent coefficients are: a left adjacent coefficient and an upper adjacent coefficient, The context of the sign flag of the current coefficient is selected based on the sign flags of the left adjacent coefficient and the upper adjacent coefficient according to the following rules: If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both negative, then the rule specifies that the first context is selected, If the sign flag of one of the left neighbor coefficient and the above neighbor coefficient is negative and the sign flag of the other is positive, then the rule specifies that a second context is selected, If the sign flags of the left neighbor coefficient and the upper neighbor coefficient are both positive, then the rule specifies that a third context is selected, wherein a first syntax element indicating whether the transform skip mode is applied to the current video block is conditionally included in a bitstream, wherein whether the first syntax element is included in the bitstream is based on allowed intra prediction modes in a quantized residual block differential pulse coding modulation codec block and use of quantized residual block differential pulse coding modulation, and when the indication of the quantized residual block differential pulse coding modulation mode is 1, it is inferred that the transform skip mode is enabled, Wherein, when the indication of the differential pulse codec modulation mode of the quantized residual block is 1, it is inferred that the transform skip mode is enabled, and is applied based on the block size of the current video block and the single / dual codec tree structure of the current video block.

10. The method of claim 9, wherein visual media processing comprises decoding the current video block from a bitstream representation.

11. The method of claim 9, wherein visual media processing comprises encoding the current video block into a bitstream representation.

12. A video encoder apparatus comprising a processor configured to implement the method of any one of claims 2-3, 8-9 and 11.

13. A video decoder device comprising a processor configured to implement the method of any one of claims 2, 4, and 9-10.

14. A computer-readable medium having code stored thereon, the code comprising processor-executable instructions for implementing the method of any one of claims 2-4 and 8-11.

Citation Information

Patent Citations

  • Coding sign information of video data

    CN108353167A

  • Determining contexts for coding transform coefficient data in video coding

    US20130182758A1