Palette coding mode

By using dictionary-based encoding/decoding modes and virtual buffer technology, the encoding/decoding process of video blocks is optimized, solving the problem of low efficiency in screen content encoding/decoding in existing technologies and achieving more efficient video quality improvement.

CN114747217BActive Publication Date: 2025-10-24DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080082778.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-30
Filing Date
2020-11-30
Publication Date
2025-10-24
Estimated Expiration
2040-11-30

AI Technical Summary

Technical Problem

Existing video codec standards struggle to effectively utilize repeating patterns within screen content, resulting in low encoding and decoding efficiency.

Method used

A dictionary-based encoding/decoding mode is adopted, which optimizes the encoding/decoding process of video blocks by using a palette mode and virtual buffer technology, including the encoding/decoding of palette indices and the determination of prediction blocks, thereby improving encoding/decoding efficiency.

Benefits of technology

It significantly improves the efficiency of screen content encoding and decoding, reduces redundancy, and enhances video quality and encoding/decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114747217B_ABST
    Figure CN114747217B_ABST
Patent Text Reader

Abstract

A method of video processing is described. The method includes performing a conversion between a current block of a video and a bitstream of the video, wherein the current video block is coded using a palette mode in which the current block is represented using a palette of representative sample values, and wherein the conversion includes selectively saving and loading palette prediction values of the palette in the palette mode based on a use of local binary trees in the conversion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 122267, filed on November 30, 2019, under applicable patent laws and / or rules of the Paris Convention. The entire disclosure of the above application is incorporated by reference as part of the disclosure of this application. Technical Field

[0003] This document covers video coding and decoding technologies, systems, and devices. Background Art

[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] Devices, systems, and methods related to digital video coding and decoding are described, including a dictionary-based coding mode for screen content coding and decoding. The described methods are applicable to existing video coding standards (e.g., High Efficiency Video Codec (HEVC) and / or Versatile Video Codec (VVC)) and future video coding standards or video codecs.

[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of video and a bitstream representation of the video, wherein the current block is encoded in a dictionary-based codec mode using one or more dictionaries, and wherein the conversion is based on the one or more dictionaries.

[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining, for conversion between a video including a video block and a bitstream representation of the video, based on a rule and one or more codec characteristics of the video block, to use one or more dictionaries for the video block; and performing the conversion based on the determination, wherein the codec characteristics include a size of the video block and / or information in the bitstream representation.

[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing, the method comprising: determining, for a current block of a video to which a dictionary-based codec mode is applied, a prediction block for the current block based on one or more entries of a dictionary; and performing conversion between the current block and a bitstream representation of the video based on the determination.

[0009] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using a dictionary-based coding mode using a dictionary, and wherein one or more entries of the dictionary are included in the bitstream representation.

[0010] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using a dictionary-based coding mode using a dictionary, and wherein one or more entries of the dictionary are included in the bitstream representation.

[0011] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using a dictionary-based coding mode using a dictionary, and wherein one or more entries of the dictionary are included in the bitstream representation.

[0012] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using a palette mode, and wherein the bitstream representation includes a syntax element representing an escape point for each of a set of coefficients of the current block.

[0013] In yet another representative aspect, the above method is implemented in the form of processor-executable code, and stored in a computer-readable program medium.

[0014] In yet another representative aspect, a device configured or operable to perform the above method is disclosed. The device can include a processor programmed to implement the method.

[0015] In yet another representative aspect, a video decoder apparatus can implement the method as described herein.

[0016] The above and other aspects and features of the disclosed technology are more fully described in the accompanying drawings, specification, and claims. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 An example of intra block copy is shown.

[0018] Figure 2 An example of five spatial neighboring candidates is shown.

[0019] Figure 3 An example of a block coded in palette mode is shown.

[0020] Figure 4 An example of signaling palette entries using palette predictor is shown.

[0021] Figure 5 An example of horizontal and vertical traversal scan is shown.

[0022] Figure 6 An example of coding palette index is shown.

[0023] Figure 7A and Figure 7B An example of a minimum chroma inter prediction unit (SCIPU) is shown.

[0024] Figure 8 An example of the problem of repeated palette (PLT) entries in the case of local dual tree is shown.

[0025] Figure 9 An example of the coding process of dictionary based coding mode is shown.

[0026] Figure 10 An example of the update process of dictionary based coding mode is shown.

[0027] Figure 11 An example of a template of a block is shown.

[0028] Figure 12 A flowchart of an example of a video processing method.

[0029] Figure 13 A flowchart of another example of a video processing method.

[0030] Figure 14A A block diagram of an example of a video processing apparatus.

[0031] Figure 14B A block diagram of an example video processing system in which the disclosed technology can be implemented.

[0032] Figure 15 A block diagram illustrating an example video coding system.

[0033] Figure 16 A block diagram showing an encoder according to some embodiments of the disclosed technology.

[0034] Figure 17 A block diagram showing a decoder according to some embodiments of the disclosed technology.

[0035] Figures 18A to 18CA flowchart showing an example method of video processing based on some implementations of the disclosed technology. DETAILED DESCRIPTION

[0036] This document provides various techniques that a decoder of an image or video bitstream can use to improve the quality of decompressed or decoded digital video or images. For brevity, the term “video” is used herein to include sequences of pictures (traditionally referred to as video) and individual images. Also, a video encoder can implement these techniques during the encoding process as well in order to reconstruct decoded frames for further encoding.

[0037] The section headings used herein are for organizational purposes only and are not to be construed as limiting the disclosed embodiments and techniques in any way. As such, the examples from one section can be combined with examples from another section.

[0038] 1. SUMMARY

[0039] This document relates to video coding technology. In particular, it relates to dictionary-based coding modes for screen content coding. It is applicable to existing video coding standards such as HEVC, or the upcoming standard (Versatile Video Coding). It is also applicable to future video coding standards or video codecs.

[0040] 2. EXAMPLE EMBODIMENTS OF VIDEO CODING

[0041] Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure wherein temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). In April 2018, the Joint Video Team (JVT) of VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was formed to work on the VVC standard with the goal of 50% bitrate reduction compared to HEVC.

[0042] 2.1 INTRA BLOCK COPY

[0043] Intra-frame block copy (IBC), also known as current picture referencing, has been adopted in HEVC Screen Content Codec Extension (HEVC-SCC) and the current VVC Test Model (VTM-4.0). IBC extends the concept of motion compensation from inter-frame codecs to intra-frame codecs. Figure 1 As shown in the figure, when IBC is applied, the current block is predicted by the reference block in the same picture. The samples in the reference block must have been reconstructed before the current block is encoded or decoded. Although IBC is not very efficient for sequences captured by most cameras, it shows significant codec gains for screen content. The reason is that there are a large number of repeated patterns in screen content pictures, such as icons and text characters. IBC can effectively eliminate the redundancy between these repeated patterns. In HEVC-SCC, if the inter-frame coding and decoding codec unit (CU) selects the current picture as its reference picture, it can apply IBC. In this case, MV is renamed block vector (BV), and BV always has integer pixel precision. In order to be compatible with the main profile HEVC, the current picture is marked as a "long-term" reference picture in the decoded picture buffer (DPB). It should be noted that, similarly, in the multi-view / 3D video coding standard, inter-view reference pictures are also marked as "long-term" reference pictures.

[0044] Following the BV to find its reference block, a prediction can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixels from the original signal. Transformation and quantization can then be applied as in other codec modes.

[0045] However, when the reference block is outside the picture, overlaps with the current block, is outside the reconstructed area, or is outside the valid area subject to certain constraints, some or all pixel values ​​are undefined. There are basically two solutions to deal with this problem. One is to disallow this situation, such as in bitstream conformance. The other is to apply padding to those undefined pixel values. The following substudies describe these solutions in detail.

[0046] 2.1.1 HEVC IBC Screen Content Codec Extension

[0047] In the screen content codec extension of HEVC, when a block uses the current picture as a reference, it should ensure that the entire reference block is within the available reconstruction area, as shown in the following normative text:

[0048] The variables offsetX and offsetY are derived as follows:

[0049] offsetX=(ChromaArrayType==0)? 0:(mvCLX[0]&0x7?2:0) (8-106)

[0050] offsetY = ( ChromaArrayType == 0 )? 0 : ( mvCLX[ 1 ] & 0x7? 2 : 0 ) (8-107)

[0051] The requirement of bitstream conformance is that, when the reference picture is the current picture, the luma motion vector mvLX shall obey the following constraints:

[0052] 1. When the derivation process for z-scan order block availability specified in clause 6.4.1 is invoked with ( xCurr, yCurr ) set equal to ( xCb, yCb ) and the neighboring luma location ( xNbY, yNbY ) set equal to ( xPb + ( mvLX[ 0 ] » 2 ) - offsetX, yPb + ( mvLX[ 1 ] » 2 ) - offsetY ) as inputs, the output shall be equal to true.

[0053] 2. When the derivation process for z-scan order block availability specified in clause 6.4.1 is invoked with ( xCurr, yCurr ) set equal to ( xCb, yCb ) and the neighboring luma location ( xNbY, yNbY ) set equal to ( xPb + ( mvLX[ 0 ] » 2 ) + nPbW - 1 + offsetX, yPb + ( mvLX[ 1 ] » 2 ) + nPbH - 1 + offsetY ) as inputs, the output shall be equal to true.

[0054] 3. One or both of the following conditions shall be true:

[0055] - The value of ( mvLX[ 0 ] » 2 ) + nPbW + xB1 + offsetX is less than or equal to 0.

[0056] - The value of ( mvLX[ 1 ] » 2 ) + nPbH + yB1 + offsetY is less than or equal to 0.

[0057] 4. The following condition shall be true:

[0058] ( xPb + ( mvLX[ 0 ] » 2 ) + nPbSw - 1 + offsetX ) / CtbSizeY - xCurr / CtbSizeY <= yCurr / CtbSizeY - ( yPb + ( mvLX[ 1 ] » 2 ) + nPbSh - 1 + offsetY ) / CtbSizeY (8-108)

[0059] Therefore, the case where the reference block overlaps with the current block or the reference block is outside the picture will not happen. There is no need to pad the reference or prediction block.

[0060] 2.1.2 IBC in VVC Test Model

[0061] In the current VVC test model, i.e. VTM-4.0 design, the whole reference block should be together with the current coding tree unit (CTU) and not overlap with the current block. Therefore, no padding of the reference or prediction block is needed. The IBC flag is coded as the prediction mode of the current CU. Therefore, there are three prediction modes, MODE_INTRA, MODE_INTER and MODE_IBC, in total for each CU.

[0062] 2.1.2.1 IBC Merge mode

[0063] In IBC Merge mode, an index pointing to an entry in the IBC Merge candidate list is parsed from the bitstream. The construction of the IBC Merge list can be summarized according to the following sequence of steps:

[0064] • Step 1: Derivation of the spatial domain candidates

[0065] • Step 2: Insertion of HMVP candidates

[0066] • Step 3: Insertion of pairwise average candidate values

[0067] In the derivation of the spatial domain Merge candidates, up to four Merge candidates are selected from the candidates located at the positions shown in the figure. The order of derivation is A1, B1, B0, A0 and B2. Position B2 is only considered if any of the positions A1, B1, B0, A0 is not available (e.g. because it belongs to another slice or tile) or not coded with IBC mode. After the insertion of the candidate at position A1, the insertion of the remaining candidates is subject to a redundancy check which ensures that candidates with the same motion information are excluded from the list, thus improving coding efficiency.

[0068] After the insertion of the spatial domain candidates, IBC candidates from the HMVP table can be inserted if the IBC Merge list size is still smaller than the maximum IBC Merge list size. When inserting the HMVP candidates, a redundancy check is performed.

[0069] Finally, pairwise average candidates are inserted into the IBC Merge list.

[0070] A Merge candidate is called invalid Merge candidate when the reference block identified by the Merge candidate is outside the picture, or overlaps with the current block, or is outside the reconstructed area, or is outside the valid area restricted by certain constraints.

[0071] Note that invalid Merge candidates can be inserted into the IBC Merge list.

[0072] JVET-N0843 was adopted into VVC. In JVET-N0843. The BV predictors for Merge mode and AMVP mode in IBC will share a common predictor list, which includes the following elements:

[0073] o 2 spatial neighboring positions (e.g.A1, B1 in Figure 2

[0074] o 5 HMVP entries

[0075] o Default to zero vector

[0076] For Merge mode, the first up to 6 entries of the list will be used; for AMVP mode, the first 2 entries of the list will be used. And the list meets the shared Merge list region requirement (same list shared within SMR).

[0077] In addition to the above BV predictor candidate list, JVET-N0843 proposes to simplify the pruning operation between HMVP candidates and existing Merge candidates (A1, B1). In the simplification, there will be up to 2 pruning operations, as it only compares the first HMVP candidate with the spatial Merge candidate.

[0078] In the latest VVC and VTM5, it is proposed to explicitly use syntax constraints to disable the 128x128 IBC mode on top of the current bitstream constraints in previous VTM and VVC versions, which makes the presence of IBC flag dependent on CU size < 128x128.

[0079] 2.1.2.2 IBC AMVP mode

[0080] In IBC AMVP mode, an AMVP index pointing to an entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized in the following steps:

[0081] Step 1: Derivation of spatial candidates

[0082] A0, A1 are checked until a valid candidate is found.

[0083] B0, B1, B2 are checked until a valid candidate is found.

[0084] Step 2: Insertion of HMVP candidates

[0085] Step 3: Insertion of zero candidate

[0086] After the insertion of spatial candidates, IBC candidates from the HMVP table can be inserted if the IBC AMVP list size is still smaller than the maximum IBC AMVP list size.

[0087] Finally, the zero candidate is inserted into the IBC AMVP list.

[0088] 2.1.2.3 IBC virtual buffer

[0089] The concept of virtual buffer is introduced to help describe the reference region of IBC prediction mode, and is adopted in the VVC draft. For a CTU with size ctbSize, we denote wIbcBuf = 128*128 / ctbSize, and define a virtual IBC buffer IbcBuf with width wIbcBuf and height ctbSize. Thus,

[0090] For a CTU with size 128x128, the size of ibcBuf is also 128x128.

[0091] For a CTU with size 64x64, the size of ibcBuf is 256x64.

[0092] For a CTU with size 32x32, the size of ibcBuf is 512x32.

[0093] Note that the VPDU width and height are min(ctbSize, 64). We denote Wv = min(ctbSize, 64).

[0094] The virtual IBC buffer ibcBuf is maintained as follows.

[0095] (1) At the beginning of decoding each CTU row, flush the entire ibcBuf with value (-1).

[0096] (2) At the beginning of decoding a VPDU (xVPDU, yVPDU) relative to the top-left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU % wIbcBuf, …, xVPDU % wIbcBuf + Wv - 1; y = yVPDU % ctbSize, …, yVPDU % ctbSize + Wv - 1.

[0097] (3) After decoding a CU relative to the top-left corner of the picture contains (x, y), set

[0098] ibcBuf[x % wIbcBuf][y % ctbSize] = recSample[x][y]

[0099] Thus the bitstream constraint can be simply described as

[0100] The requirement of bitstream conformance is that, for bv, ibcBuf[(x+bv[0]) % wIbcBuf][(y+bv[1]) % ctbSize] shall not be equal to -1.

[0101] With the concept of IBC reference buffer, it also simplifies the text of the decoding process by avoiding the reference inter- frame interpolation and motion compensation process, including sub-block process.

[0102] In addition, the text of the general decoding process of a coding unit coded with IBC prediction mode in JVET-P2001 is shown as follows.

[0103] The inputs of this process are:

[0104] 5. luma position (xCb, yCb) specifying the top-left sample of the current coding block relative to the top-left luma sample of the current picture,

[0105] 6. variable cbWidth specifying the width of the current coding block in luma samples,

[0106] 7. variable cbHeight specifying the height of the current coding block in luma samples,

[0107] 8. variable treeType specifying whether a single tree or a dual tree is used, and if a dual tree is used, it specifies whether the current tree corresponds to the luma component or the chroma component.

[0108] The output of this process is the modified reconstructed picture before loop filtering.

[0109] The quantization parameter derivation process specified in clause 8.7.1 is invoked with the luma position (xCb, yCb), the width cbWidth of the current coding block in luma samples, and the height cbHeight of the current coding block in luma samples, and the variable treeType as inputs.

[0110] The variable IsGt4by4 is derived as follows:

[0111] IsGt4by4 = (cbWidth * cbHeight > 16)? TRUE : FALSE (8-893)

[0112] The decoding process of a coding unit coded with IBC prediction mode includes the following ordered steps:

[0113] 1. The block vector components of the current coding unit are derived as follows:

[0114] 9. Invoke the derivation process of the block vector components as specified in clause 8.6.2.1 with luma coded block position (xCb, yCb), luma coded block width cbWidth and luma coded block height cbHeight as input and with luma block vector bvL as output.

[0115] 10. When treeType is equal to SINGLE_TREE, invoke the derivation process of the chroma block vector as specified in clause 8.6.2.5 with luma block vector bvL as input and with chroma block vector bvC as output.

[0116] 2. The prediction samples of the current coding unit are derived as follows:

[0117] 11. Invoke the decoding process of an IBC block as specified in clause 8.6.3.1 with luma coded block position (xCb, yCb), luma coded block width cbWidth and luma coded block height cbHeight, luma block vector bvL, variable cldx set to 0 as input and with the IBC prediction samples (predSamples) of a (cbWidth) x (cbHeight) array predSamplesL of predicted luma samples as output.

[0118] 12. When treeType is equal to SINGLE_TREE, the prediction samples of the current coding unit are derived as follows:

[0119] i. Invoke the decoding process of an IBC block as specified in clause 8.6.3.1 with luma coded block position (xCb, yCb), luma coded block width cbWidth and luma coded block height cbHeight, chroma block vector bvC and variable cldx set to equal 1 as input and with the IBC prediction samples (predSamples) of a (cbWidth / SubWidthC) x (cbHeight / SubHeightC) array predSamplesCb of predicted chroma samples of chroma component Cb as output.

[0120] ii. The decoding process for IBC blocks as specified in clause 8.6.3.1 is invoked with luma coded block position (xCb, yCb), luma coded block width cbWidth and luma coded block height cbHeight, chroma block vector bvC and variable cldx set equal to 2 as inputs and the IBC predicted samples (predSamples) as an array of (cbWidth / SubWidthC) x (cbHeight / SubHeightC) predicted luma samples for luma component Y and an array of (cbWidth / SubWidthC) x (cbHeight / SubHeightC) predicted chroma samples for chroma component Cr (predSamplesCr) as outputs.

[0121] 3. The residual samples of the current coding unit are derived as follows:

[0122] 13. The decoding process for residual signal of a coded block coded with inter prediction mode as specified in clause 8.5.8 is invoked with position (xTbO, yTbO) set equal to luma position (xCb, yCb), width nTbW set equal to luma coded block width cbWidth, height nTbH set equal to luma coded block height cbHeight, and variable cldx set equal to 0 as inputs and array resSamplesL as output.

[0123] 14. When treeType is equal to SINGLE_TREE, the decoding process for residual signal of a coded block coded with inter prediction mode as specified in clause 8.5.8 is invoked with position (xTbO, yTbO) set equal to chroma position (xCb / SubWidthC, yCb / SubHeightC), width nTbW set equal to chroma coded block width cbWidth / SubWidthC, height nTbH set equal to chroma coded block height cbHeight / SubHeightC, and variable cldx set equal to 1 as inputs and array resSamplesCb as output.

[0124] 15. When treeType is equal to SINGLE_TREE, the decoding process for residual signal of a coded block coded with inter prediction mode as specified in clause 8.5.8 is invoked with position (xTbO, yTbO) set equal to chroma position (xCb / SubWidthC, yCb / SubHeightC), width nTbW set equal to chroma coded block width cbWidth / SubWidthC, height nTbH set equal to chroma coded block height cbHeight / SubHeightC, and variable cldx set equal to 2 as inputs and array resSamplesCr as output.

[0125] 4. The reconstructed samples of the current coding unit are derived as follows:

[0126] 16. The picture reconstruction process for the color component specified in clause 8.7.5 is invoked with the block position (xB, yB) set equal to (xCb, yCb), the block width bWidth set equal to cbWidth, the block height bHeight set equal to cbHeight, the variable treeType, the variable cldx set equal to 0, the (cbWidth) x (cbHeight) array predSamplesL set equal to predSamplesL, and the (cbWidth) x (cbHeight) array resSamples set equal to resSamplesL as inputs, and the output being the modified reconstructed picture before loop filtering.

[0127] 17. When treeType is equal to SINGLE_TREE, the picture reconstruction process for the color component specified in clause 8.7.5 is invoked with the block position (xB, yB) set equal to (xCb / SubWidthC, yCb / SubHeightC), the block width bWidth set equal to cbWidth / SubWidthC, the block height bHeight set equal to cbHeight / SubHeightC, the variable treeType, the variable cldx set equal to 1, the (cbWidth / SubWidthC) x (cbHeight / SubHeightC) array predSamples set equal to predSamplesCb, and the (cbWidth / SubWidthC) x (cbHeight / SubHeightC) array resSamples set equal to resSamplesCb as inputs, and the output being the modified reconstructed picture before loop filtering.

[0128] 18. When treeType is equal to SINGLE_TREE, invoke the picture reconstruction process for the color component specified in clause 8.7.5 with the block position (xB, yB) set to equal (xCb / SubWidthC, yCb / SubHeightC), the block width bWidth set to equal cbWidth / SubWidthC, the block height bHeight set to equal cbHeight / SubHeightC, the variable treeType, the variable cldx set to equal 2, the (cbWidth / SubWidthC) x (cbHeight / SubHeightC) array predSamples set to equal predSamplesCr, and the (cbWidth / SubWidthC) x (cbHeight / SubHeightC) array resSamples set to equal resSamplesCr as inputs, and the output is the reconstructed picture modified before loop filtering.

[0129] 2.2 Palette mode in HEVC screen content coding extension (HEVC-SCC)

[0130] 2.2.1 Concept of palette mode

[0131] The basic idea behind palette mode is that pixels in a CU are represented by a small set of representative color values. This set is called palette. And it is also possible to indicate samples outside the palette by signaling an escape symbol followed by a (possibly quantized) component value. This type of pixel is called escape pixel. Palette mode is illustrated in Figure 3 As illustrated in Figure 3 For each pixel with three color components (luma and two chroma components), an index to the palette is established and the block can be reconstructed based on the values established in the palette.

[0132] Palette mode in HEVC-SSC

[0133] For the coding of palette entries, a palette predictor is maintained. The maximum size of the palette as well as the palette predictor is signaled in the SPS. In HEVC-SCC, a palette_predictor_initializer_present_flag is introduced in the PPS. When this flag is equal to 1, an entry that initializes the palette predictor is signaled in the bitstream. The palette predictor is initialized at the beginning of each CTU row, each slice and each tile. Depending on the value of the palette_predictor_initializer_present_flag, the palette predictor is either reset to 0 or initialized using the palette predictor initializer entry signaled in the PPS. In HEVC-SCC, a palette predictor initializer with size 0 is enabled to allow explicit disabling of the palette predictor initialization at the PPS level.

[0134] For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette or not. This is illustrated in Figure 4 The reuse flag is sent using run-length coding with zero. Thereafter, the number of new palette entries is signaled using a zeroth order exponential Golomb code. Finally, the component values of the new palette entries are signaled.

[0135] The palette indices are coded using a horizontal and vertical traversing scan, as illustrated in Figure 5 The scan order is explicitly signaled in the bitstream using the palette_transpose_flag. For the rest of this subsection, it is assumed that the scan is horizontal.

[0136] The palette indices are coded using two main palette sample modes: 'INDEX' and 'COPY_ABOVE'. As mentioned before, the escape symbol is also signaled as 'INDEX' mode and is assigned an index equal to the maximum palette size. The mode is signaled using a flag, except for the top row or when the previous mode was 'COPY_ABOVE'. In 'COPY_ABOVE' mode, the palette index of the sample above is copied. In 'INDEX' mode, the palette index is explicitly signaled. For both 'INDEX' and 'COPY_ABOVE' modes, a run value is signaled that specifies the number of subsequent samples that are also coded using the same mode. When the escape symbol is part of a run in 'INDEX' or 'COPY_ABOVE' mode, the escape component value is signaled for each escape symbol. The coding of the palette indices is illustrated in Figure 6

[0137] ​This syntax order is accomplished in the following way. First, the number of index values for the CU is signaled. This is followed by signaling the actual index values for the entire CU using truncated binarization. Both the number of indices and the index values are coded in bypass mode. This groups the bypass bins related to the indices together. Then, the palette sample mode (if necessary) and the run are signaled in a run-by-run fashion. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and coded in bypass mode.

[0138] The additional syntax element last_run_type_flag is signaled after the index values are signaled. This syntax element, in combination with the number of indices, eliminates the need to signal the run value corresponding to the last run in the block.

[0139] In HEVC-SCC, the palette mode also supports 4:2:2, 4:2:0 and monochrome chroma formats. The signaling of palette entries and palette indices is almost identical for all chroma formats. In the case of non-monochrome formats, each palette entry includes 3 components. For monochrome formats, each palette entry includes one component. For sub-sampled chroma directions, the chroma samples are associated with the luma sample indices that are divisible by 2. After the palette indices are reconstructed for a CU, only the first component of the palette entry is used if the sample has only a single component associated with it. The only difference in signaling is for the escape component values. For each escape sample, the number of escape component values signaled can be different depending on the number of components associated with the sample.

[0140] Furthermore, there is an index adjustment procedure in palette index coding. When a palette index is signaled, the left or above neighboring index should be different from the current index. Therefore, the range of the current palette index can be reduced by 1 by removing one possibility. After this, the index is signaled using truncated binarization (TB) binarization.

[0141] The text related to this section is shown below, where CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the predicted index.

[0142] The variable PaletteIndexMap[ xC ][ yC ] specifies the palette index, which is the index into the array represented by CurrentPaletteEntries. The array indices xC, yC specify the position of the sample relative to the top-left luma sample of the picture ( xC, yC ). The value of PaletteIndexMap[ xC ][ yC ] shall be in the range of 0 to MaxPaletteIndex, inclusive.

[0143] The derivation of the variable adjustedRefPaletteIndex is as follows:

[0144]

[0145]

[0146] The variable CurrPaletteIndex is derived as follows when CopyAboveIndicesFlag[ xC ][ yC ] is equal to 0:

[0147]

[0148] if ( CurrPaletteIndex >= adjustedRefPaletteIndex )

[0149] CurrPaletteIndex++

[0150] 2.2.3 Palette mode in VVC

[0151] 2.2.3.1 Palette in dual tree

[0152] In VVC, dual tree coding structure is used for coding intra slices, so that the luma component and the two chroma components can have different palettes and palette indices. In addition, the two chroma components share the same palette and palette index.

[0153] 2.2.3.2 Palette as a separate mode

[0154] In JVET-N0258 and current VTM, the prediction mode of a coding unit can be MODE_INTRA, MODE_INTER, MODE_IBC and MODE_PLT. The binarization of the prediction mode changes accordingly.

[0155] When IBC is turned off, on I slices, the first, one bin is used to indicate whether the current prediction mode is MODE_PLT. When on P / B slices, the first bin is used to indicate whether the current prediction mode is MODE_INTRA. If not, one additional bin is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTER.

[0156] ​When IBC is on, on I-slices, the first bin is used to indicate whether the current prediction mode is MODE_IBC or not. If not, the second bin is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. When on P / B-slices, the first bin is used to indicate whether the current prediction mode is MODE_INTRA or not. If it is an intra mode, the second bin is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. If not, the second bin is used to indicate whether the current prediction mode is MODE_IBC or MODE_INTER.

[0157] Furthermore, JVET-P0516 is adopted. It is proposed to signal MODE_PLT when the prediction mode is MODE_INTRA under all conditions. The proposed modification introduces changes only under the condition that only inter, intra, and PLT modes are allowed.

[0158] 2.2.3.3 Line-based CG palette mode

[0159] VVC adopted a line-based CG palette mode. In this method, each CU of palette mode is divided into multiple segments of m samples (m = 16 in this test) based on the traversing scan pattern. The coding order of palette runs in each segment is as follows: for each pixel, a context-coded bin run copy flag = 0 is signaled to indicate whether the pixel has the same mode as the previous pixel, i.e., whether both the previously scanned pixel and the current pixel are of run type COPY ABOVE, or both the previous scanned pixel and the current pixel are of run type INDEX and the same index value. Otherwise, run copy flag = 1 is signaled. If the pixel and the previous pixel are of different modes, a context-coded bin copy above palette indices flag is signaled to indicate the run type of the pixel, i.e., INDEX or COPY ABOVE. As in the palette mode in VTM 6.0, if a sample is at the first row (horizontal traversing scan) or the first column (vertical traversing scan), the decoder does not have to resolve the run type because the INDEX mode is used by default. In addition, if the previously resolved run type is COPY ABOVE, the decoder does not have to resolve the run type. After the palette run coding of the pixels in one segment, the index values (for INDEX mode) and the quantized escape colors are bypass-coded and grouped separately from the context-coded bins for coding / parse to improve the throughput within each line CG. Because the index values are now coded / parsed after the run coding, instead of being processed before the palette run coding as in VTM, the encoder does not have to signal the number of index values num_palette_indices_minusl and the last run type copy above indices for final run flag.

[0160] The text of the line-based CG mode provided for palette mode in JVET-P0077 is shown below.

[0161]

[0162]

[0163]

[0164]

[0165] Palette coding semantics

[0166] In the following semantics, array indices x0, y0 specify the position of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture (x0, y0). Array indices xC, yC specify the position of the sample relative to the top-left luma sample of the picture (xC, yC). Array index startComp specifies the first colour component of the current palette list. startComp equal to 0 indicates the Y component; startComp equal to 1 indicates the Cb component; startComp equal to 2 indicates the Cr component. numComps specifies the number of colour components in the current palette list.

[0167] The predictor palette includes palette entries from a previous coding unit for predicting entries in the current palette.

[0168] The variable PredictorPaletteSize[startComp] specifies the size of the predictor palette for the first colour component of the current palette list startComp. PredictorPaletteSize is derived according to the specification in clause 8.4.5.3.

[0169] The variable PalettePredictorEntryReuseFlags[i] equal to 1 specifies that the i-th entry in the predictor palette is reused in the current palette. PalettePredictorEntryReuseFlags[i] equal to 0 specifies that the i-th entry in the predictor palette is not an entry in the current palette. All elements of the array PalettePredictorEntryReuseFlags[i] are initialized to 0.

[0170] palette_predictor_run is used to determine the number of zeros that face the non-zero entries in the array PalettePredictorEntryReuseFlags.

[0171] A requirement for bitstream conformance is that the value of palette_predictor_run shall be in the range of 0 to (PredictorPaletteSize - predictorEntryIdx), inclusive, where predictorEntryIdx corresponds to the current position in the array PalettePredictorEntryReuseFlags. The variable NumPredictedPaletteEntries specifies the number of entries in the current palette that are reused from the predictor palette. The value of NumPredictedPaletteEntries shall be in the range of 0 to palette_max_size, inclusive.

[0172] num_signalled_palette_entries specifies the number of entries in the current palette that are explicitly signalled for the first colour component of the current palette list startComp.

[0173] When num_signalled_palette_entries is not present, it is inferred to be equal to 0.

[0174] The variable CurrentPaletteSize[ startComp ] specifies the size of the current palette for the first colour component of the current palette list startComp and is derived as follows:

[0175] CurrentPaletteSize[ startComp ] = NumPredictedPaletteEntries + num_signalled_palette_entries (7-155)

[0176] The value of CurrentPaletteSize[ startComp ] shall be in the range of 0 to palette_max_size, inclusive.

[0177] new_palette_entries[ cIdx ][ i ] specifies the value of the i-th signalled palette entry for colour component cIdx.

[0178] The variable PredictorPaletteEntries[ cIdx ][ i ] specifies the i-th element in the predictor value palette for colour component cIdx.

[0179] The variable CurrentPaletteEntries[ cIdx ][ i ] specifies the i-th element in the current palette for colour component cIdx and is derived as follows:

[0180]

[0181]

[0182] palette_escape_val_present_flag equal to 1 specifies that the current coding unit contains at least one escape coded sample. escape_val_present_flag equal to 0 specifies that there are no escape coded samples in the current coding unit. When not present, the value of palette_escape_val_present_flag is inferred to be equal to 1.

[0183] The variable MaxPaletteIndex specifies the maximum possible value of palette indices for the current coding unit. The value of MaxPaletteIndex is set equal to CurrentPaletteSize[ startComp ] - 1 + palette_escape_val_present_flag.

[0184] palette_idx_idc is an indication of the index of the palette list CurrentPaletteEntries. The value of palette_idx_idc shall be in the range of 0 to MaxPaletteIndex, inclusive, for the first index in the block, and in the range of 0 to (MaxPaletteIndex - 1), inclusive, for the remaining indices in the block.

[0185] When palette_idx_idc is not present, it is inferred to be equal to 0.

[0186] palette_transpose_flag equal to 1 specifies that a vertical traversal scan is applied to scan the indices of samples in the current coding unit. palette_transpose_flag equal to 0 specifies that a horizontal traversal scan is applied to scan the indices of samples in the current coding unit. When not present, the value of palette_transpose_flag is inferred to be equal to 0.

[0187] The array TraverseScanOrder specifies the scan order array for palette coding. If palette_transpose_flag is equal to 0, TraverseScanOrder is assigned the horizontal scan order HorTravScanOrder, and if palette_transpose_flag is equal to 1, TraverseScanOrder is assigned the vertical scan order VerTravScanOrder.

[0188] If copy_above_palette_indices_flag is equal to 0, run_copy_flag equal to 1 specifies that the palette run type is the same as the run type of the previous scan position and the palette run index is the same as the index of the previous position. Otherwise, run_copy_flag is equal to 0

[0189] copy_above_palette_indices_flag equal to 1 specifies that the palette index is equal to the palette index in the same position in the above line (if using horizontal traversal scan) or in the same position in the left column (if using vertical traversal scan). copy_above_palette_indices_flag equal to 0 specifies that the indication of the palette index of the sample is coded in the bitstream or inferred.

[0190] The variable CopyAboveIndicesFlag[ xC ][ yC ] equal to 1 specifies that the palette index is copied from the palette index in the above line (horizontal scan) or in the left column (vertical scan). CopyAboveIndicesFlag[ xC ][ yC ] equal to 0 specifies that the palette index is explicitly coded in the bitstream or inferred. The array indices xC, yC specify the position of the sample relative to the top-left luma sample of the picture ( xC, yC ).

[0191] The variable PaletteIndexMap[ xC ][ yC ] specifies the palette index, which is the index into the array represented by CurrentPaletteEntries. The array indices xC, yC specify the position of the sample relative to the top-left luma sample of the picture ( xC, yC ). The value of PaletteIndexMap[ xC ][ yC ] shall be in the range of 0 to MaxPaletteIndex, inclusive.

[0192] The variable adjustedRefPaletteIndex is derived as follows:

[0193]

[0194] When CopyAboveIndicesFlag[ xC ][ yC ] is equal to 0, the variable CurrPaletteIndex is derived as follows:

[0195] if( CurrPaletteIndex >= adjustedRefPaletteIndex )

[0196] CurrPaletteIndex++ (7-158)

[0197] palette_escape_val specifies the quantized escape-coded sample value for the component.

[0198] The variable PaletteEscapeVal[ cldx ][ xC ][ yC ] specifies the escape value of the sample where PalettelndexMap[ xC ][ yC ] is equal to MaxPalettelndex and palette escape val present flag is equal to 1. The array index cldx specifies the color component. The array indices xC, yC specify the position ( xC, yC ) of the sample relative to the top-left luma sample of the picture.

[0199] The requirement of bitstream conformance is that PaletteEscapeVal[ cldx ][ xC ][ yC ] shall be in the range of 0 to ( 1 « ( BitDepthY + 1 ) ) - 1, inclusive, for cldx equal to 0, and in the range of 0 to ( 1 « ( BitDepthC + 1 ) ) - 1, inclusive, for cldx not equal to 0.

[0200] 2.3 Local dual tree in VVC

[0201] In typical hardware video encoders and decoders, the processing throughput drops when a picture has more small intra blocks due to sample processing data dependency between neighboring intra blocks. The prediction value generation of an intra block requires samples from the top and left boundary reconstruction of neighboring blocks. Therefore, intra prediction has to be processed sequentially block by block.

[0202] In HEVC, the minimum intra CU is 8x8 luma samples. The luma component of a minimum intra CU can be further partitioned into four 4x4 luma intra prediction units (PUs), but the chroma components of a minimum intra CU cannot be further partitioned. Therefore, the hardware processing throughput is worst when processing 4x4 chroma intra blocks or 4x4 luma intra blocks.

[0203] In VTM5.0, the minimum intra CU is 4x4 luma samples in a single coding tree due to the fact that chroma partition always follows luma, and therefore the minimum chroma intra CB is 2x2. Therefore, the minimum chroma intra CB in a single coding tree is 2x2 in VTM5.0. The worst case hardware processing throughput of VVC decoding is only 1 / 4 of that of HEVC decoding. Furthermore, after adopting tools including cross-component linear model (CCLM), 4-tap interpolation filter, position-dependent intra prediction combination (PDPC), and combined inter-intra prediction (CIIP), the reconstruction process of chroma intra CB becomes much more complex than in HEVC. It is challenging to achieve high processing throughput in a hardware decoder. In this chapter, we will propose a method to improve the worst case hardware processing throughput.

[0204] The goal of the method is to prohibit chroma intra CBs smaller than 16 chroma samples by constraining the partitioning of chroma intra CBs.

[0205] In a single coding tree, a SCIPU is defined as a coding tree node whose chroma block size is greater than or equal to TH chroma samples and has at least one sub-luma block smaller than 4TH luma samples, where TH is set to 16 in this contribution. It is required that in each SCIPU, all CBs are either inter or all CBs are non-inter, i.e., intra or IBC. In the case of non-inter SCIPU, it is further required that the chroma of the non-inter SCIPU should not be further partitioned, and the luma of the SCIPU is allowed to be further partitioned. In this way, the minimum chroma intra CB size is 16 chroma samples, and 2x2, 2x4 and 4x2 chroma CBs are removed. Furthermore, in the case of non-inter SCIPU, no chroma scaling is applied. Moreover, when the luma block is further partitioned while the chroma block is not, a local dual tree coding structure is constructed.

[0206] Figure 7A and Figure 7B Two SCIPU examples are shown. In Figure 7A , one chroma CB of 8x4 chroma samples and three luma CBs (4x8, 8x8, 4x8 luma CBs) form one SCIPU, because a ternary tree (TT) split from 8x4 chroma samples will result in a chroma CB smaller than 16 chroma samples. In Figure 7B , one chroma CB of 4x4 chroma samples (left side of 8x4 chroma samples) and three luma CBs (8x4, 4x4, 4x4 luma CBs) form one SCIPU, while the other chroma CB of 4x4 samples (right side of 8x4 chroma samples) and two luma CBs (8x4, 8x4 luma CBs) form one SCIPU, because a binary tree (BT) split from 4x4 chroma samples will result in a chroma CB smaller than 16 chroma samples.

[0207] In the proposed method, if the current slice is an I slice or the current SCIPU has 4x4 luma partitioning in it after further partitioning once (because inter 4x4 is not allowed in VVC), the type of the SCIPU is inferred to be non-inter; otherwise, the type (inter or non-inter) of the SCIPU is indicated by one signaled flag before parsing the CUs in the SCIPU.

[0208] By applying the above method, the worst case hardware processing throughput occurs when processing 4x4, 2x8 or 8x2 chroma blocks instead of 2x2 chroma blocks. The worst case hardware processing throughput is the same as HEVC, and is 4 times the hardware processing throughput in VTM5.0.

[0209] 3. Examples of problems solved by embodiments

[0210] 1. To reduce the bandwidth cost, IBC can only refer to the blocks above the CTU in the left side coding tree unit (CTU) as the prediction block in VVC, which limits the performance of IBC mode for screen content coding.

[0211] 2. The binarization depends on the escape flag signaled at the CU level, and when the escape flag is true, one index usually costs more bits. However, some CGs can not have escape samples, so some bits can be saved, which is not considered in the current design.

[0212] 3. Local dual tree and PLT cannot be applied at the same time, because when coding from single tree region to dual tree region, some palette entries can be repeated. An example is shown as Figure 8 .

[0213] 4. The number of entries in the palette prediction value and the maximum allowed number of palette entries are fixed, which can lose the flexibility of controlling the efficiency and throughput of the palette mode.

[0214] 4. Examples of embodiments

[0215] The following detailed description should be considered in the context of the examples of the general concepts. These descriptions should not be interpreted in a narrow way. In addition, these descriptions can be combined in any way.

[0216] DCM

[0217] 1. The DCM can be considered as a new prediction mode in addition to the existing prediction modes (e.g., intra / inter / IBC prediction modes).

[0218] a. In one instance, the DCM can be considered as part of a selected existing prediction mode (e.g., IBC prediction mode).

[0219] i. Alternatively, in addition to when utilizing the selected existing prediction mode, an indication of the use of the DCM can be further signaled.

[0220] ii. In one example, the indication of the use of the DCM can be context coded or bypass coded.

[0221] b. Alternatively, in one example, the indication of the use of the DCM can be signaled / parsed as a separate prediction mode (e.g., MODE_DCM).

[0222] c. In one example, the indication of the use of the DCM can be signaled / parsed under certain conditions.

[0223] i. An indication of the DCM usage can be signaled / parsed under a conditional check on the block dimension.

[0224] 1. In one example, it can be signaled only if the block size is smaller or equal to MxN.

[0225] a) Or, in one example, it can be signaled only if the block width is smaller or equal to M and / or the block height is smaller or equal to N.

[0226] ii. An indication of the DCM usage can be signaled / parsed under a conditional check on the block position.

[0227] 1. In one example, it can be signaled only if the vertical and / or horizontal coordinates of the top-left sample in the block are not equal to 0, e.g., relative to the slice / tile / tile group containing the block.

[0228] 2. In one example, it can be signaled only for blocks not contained in a first CTU, e.g., the slice / tile / tile group containing the block.

[0229] iii. An indication of the DCM usage can be signaled / parsed under a conditional check on previously coded information.

[0230] 1. In one example, it can be signaled only for luma blocks when a dual tree coding structure is applied.

[0231] 2. Multiple dictionaries can be utilized and how a dictionary is selected for a block can depend on coded information, e.g., according to the block dimension or a value signaled in the bitstream.

[0232] a. In one example, multiple dictionaries can be utilized for blocks with the same coded information, e.g., the same block width and height.

[0233] i. Or, in addition, an index of the multiple dictionaries can be signaled / parsed in the bitstream.

[0234] b. In one example, only one dictionary can be utilized for blocks with the same coded information, e.g., the same block width and height, and a different dictionary can be utilized for blocks with different coded information.

[0235] i. In one example, a first dictionary is utilized for blocks with K*L dimensions and a second dictionary is utilized for blocks with M*N dimensions, where M! = K and / or N! = L.

[0236] 3. A dictionary utilized in the DCM can contain one or more entries.

[0237] a. The maximum number of entries in the dictionary can be predefined or signaled in the bitstream.

[0238] b. In one example, an entry of the dictionary can contain multiple samples / pixels, e.g., K*L samples / pixels.

[0239] c. In one example, an entry of the dictionary can be a block in the reconstructed region.

[0240] i. Alternatively, in addition, an entry of the dictionary can be a luma block in the reconstructed region.

[0241] 3. When the dictionary-based coding mode (DCM) is applied, the prediction block of the current block can be generated according to one or more entries of one or more dictionaries.

[0242] a. In one example, an entry in the dictionary can be used as the prediction block of the current block.

[0243] i. In one example, the index of this entry can be signaled to the decoder.

[0244] 1. Alternatively, in one example, the index of this entry can be inferred as N.

[0245] ii. In one example, how to select the best entry can be determined by minimizing a certain cost.

[0246] 1. In one example, the cost can represent the rate-distortion cost between the entry and the current block.

[0247] 2. Alternatively, in one example, the cost can represent the distortion between the entry and the current block, such as SAD, SATD, SSE, or MSE.

[0248] b. In one example, the prediction block of the current block can be completely dependent on the entries in the dictionary.

[0249] i. Alternatively, the prediction block of the current block can rely on the reconstructed region in the dictionary and the current picture that is not included in the dictionary.

[0250] c. In one example, the procedure of DCM can be as shown in Figure 9 . In Figure 9 , the dictionary has N entries, and B i is the i-th entry in the dictionary. The current block is denoted by C. The encoder can first check each entry B i (0<=i<=N-1), and determine the best prediction block of C under certain criteria, which is denoted by B K in Figure 9 . After that, C and B KThe corresponding residual blocks between the two can be transformed, quantized and / or entropy coded.

[0251] 4. It is proposed to reset the dictionary before coding a slice / tile / picture and then update the dictionary after coding the video unit.

[0252] a. In one example, the video unit can be a block / CU / CTU / CTU row.

[0253] b. In one example, all overlapping blocks with different block sizes contained in the video unit can be used to update the dictionary.

[0254] i. Alternatively, in one example, only blocks with size equal to MxN can be used to update the dictionary.

[0255] c. In one example, a hash function can be used to exclude similar / same blocks from the dictionary.

[0256] i. In one example, the hash function can be a CRC function with N bits.

[0257] ii. In one example, in a raster order, if two blocks have the same hash value, only the latter block can be included in the dictionary.

[0258] 1. Alternatively, in one example, if two blocks have the same hash value, both blocks can be included in the dictionary.

[0259] iii. In one example, the update process of the DCM can be as shown in Figure 10 . In Figure 10 , the dictionary has N entries, B i is the i-th entry in the dictionary. The current block is denoted by C. Let x be the hash value of C. When C is used to update the dictionary, x is first derived, and then the entry at the x-th position can be filled / replaced by C. In Figure 10 , x is equal to 1, so in this example, entry B1 can be replaced by C.

[0260] iv. In one example, each entry in the dictionary can be a list of blocks storing blocks with the same hash value.

[0261] 1. In one example, the list can be updated with a first-in-first-out (FIFO) policy.

[0262] 2. In one example, the list size can be equal to m

[0263] a. In one example, M can be set equal to 1.

[0264] 5. The entries in the dictionary can first be sorted before being used to derive the prediction / reconstruction of the current block.

[0265] a. Propose to sort the entries in the dictionary based on the distortion between the template of each entry and the template of the current block.

[0266] b. In one example, as shown in Figure 11 , if the current block is S1 x S2, the template can represent an M x N region other than the region of the current block, where M > S1 and N > S2.

[0267] c. In one example, the distortion in the above example can represent the distortion between two templates, such as SAD, SATD, SSE, or MSE.

[0268] d. In one example, the entry can include not only the reconstructed region of the block, but also the template of the block.

[0269] e. In one example, the dictionary can be sorted in ascending / descending order based on the template distortion cost.

[0270] f. In one example, after sorting the dictionary in descending order based on the template distortion cost, only the top K entries can be applied when coding the current block with DCM.

[0271] i. In one example, assuming the dictionary size is N, the index range can be reduced from [0, N-1] to [0, K-1].

[0272] 6. For the block coded with DCM, the selected entry(ies) index(es) can be explicitly or implicitly signaled in the bitstream.

[0273] a. In one example, fixed length coding / Exp-Golomb / truncated unary / truncated binary can be used to binarize the entry index.

[0274] b. The binarization of the signaled index in DCM can depend on the number of possible entries used in DCM.

[0275] i. In one example, the binarization can be fixed length with M bits.

[0276] 1. In one example, M can be set equal to floor(log2(N)).

[0277] a. In one example, log2(N) is a function that obtains the logarithm of N with base 2.

[0278] b. In one example, floor(N) is a function that obtains the nearest integer of the upper bound of N.

[0279] ii. In one example, the binarization can be truncated unary, where cMax equals N.

[0280] iii. In the above example, N can be the number of all entries in the dictionary.

[0281] 1. Alternatively, N can be the number of available entries in the dictionary, such as K in bullet 5.f.

[0282] Escape flag in line-based CG mode of palette mode

[0283] 7. These can be indicated for each CG whether they are escape samples or not, and the escape flags for a CG can be coded together.

[0284] a. In one example, a syntax element, e.g., palette_escape_val_present_flag, can be signaled at the CG level.

[0285] i. In one example, a flag can be first signaled to indicate whether the escape flag for all CGs is false. If there is any escape flag equal to true, the escape flag for each CG can be further signaled.

[0286] ii. In one example, the index of the escape flag equal to true can be signaled.

[0287] iii. In one example, the index of the escape flag equal to false can be signaled.

[0288] b. Alternatively, a block-based flag can be sent to indicate whether there is an escape sample or not.

[0289] i. In one example, when the block-based flag indicates there is no escape sample, the CG-level escape sample present flag can be skipped.

[0290] c. The values of the escape flags for all CGs in a block can be concatenated and coded together.

[0291] i. In one example, each bit of the variable E represents the escape sample present flag for a CG.

[0292] ii. In one example, a fixed-length binarization can be used to code E.

[0293] 1. Alternatively, in one example, a truncated unary binarization can be used to code E.

[0294] 2. Alternatively, in one example, an exp-Golomb binarization with order k can be used to code E.

[0295] iii. In one example, E can be bypass coded or context coded.

[0296] Palette predictor correlation

[0297] 8. It is proposed to save and load palette predictor values based on the usage of local dual tree.

[0298] a. In one example, when the single tree is switched to the local dual tree, i.e. after decoding the last block in the single tree or before decoding the first block in the local dual tree, the palette predictor values can be saved into a temporary buffer.

[0299] i. In one example, the palette predictor values can be reset before being used for decoding the first block in the local dual tree.

[0300] b. In one example, when the local dual tree is switched to the single tree, i.e. after decoding the last block in the local dual tree or before decoding the first block in the single tree, the palette predictor values can be loaded from the temporary buffer.

[0301] 9. It is proposed to signal the maximum allowed palette size in DPS / SPS / VPS / PPS / APS / Picture header / Slice header / Tile group header / Maximum coding unit (LCU) / Coding unit (CU) / LCU row / group of LCUs / TU / PU block / video coding unit.

[0302] 10. It is proposed to signal the maximum allowed size of palette predictor values in DPS / SPS / VPS / PPS / APS / Picture header / Slice header / Tile group header / Maximum coding unit (LCU) / Coding unit (CU) / LCU row / group of LCUs / TU / PU block / video coding unit.

[0303] Technical solution applicable to all the above items

[0304] 11. M, N, K and / or L in the above examples can be integers.

[0305] a. In one example, both M and N can be equal to 4.

[0306] 1. In one example, N can be a predefined constant value for all QPs.

[0307] 2. In one example, N can be signaled to the decoder.

[0308] 3. In one example, N can be based on

[0309] a. video content (e.g. screen content or natural content)

[0310] b. messages signaled in DPS / SPS / VPS / PPS / APS / picture header / slice header / tile group header / largest coding unit (LCU) / coding unit (CU) / LCU row / group of LCUs / TU / PU block / video coding unit

[0311] c. location of CU / PU / TU / block / video coding unit

[0312] d. block size of current block and / or its neighboring blocks

[0313] e. block shape of current block and / or its neighboring blocks

[0314] f. quantization parameter of current block

[0315] g. indication of color format (e.g. 4:2:0, 4:4:4, RGB or YUV)

[0316] h. coding tree structure (e.g. dual tree or single tree)

[0317] i. slice / tile group type and / or picture type

[0318] j. color component (e.g. can be applied to luma component and / or chroma component only)

[0319] k. temporal layer ID

[0320] 12. whether and / or how to apply the above methods can be based on:

[0321] a. video content (e.g. screen content or natural content)

[0322] i. in one example, the above methods can be applied to screen content only.

[0323] b. messages signaled in DPS / SPS / VPS / PPS / APS / picture header / slice header / tile group header / largest coding unit (LCU) / coding unit (CU) / LCU row / group of LCUs / TU / PU block / video coding unit

[0324] c. location of CU / PU / TU / block / video coding unit

[0325] d. block size of current block and / or its neighboring blocks

[0326] e. block shape of current block and / or its neighboring blocks

[0327] f. quantization parameter of current block

[0328] g. indication of color format (e.g. 4:2:0, 4:4:4, RGB or YUV)

[0329] h. Coded tree structure (e.g., dual tree or single tree)

[0330] i. Slice / tile group type and / or picture type

[0331] j. Color component (e.g., can be applied to luma component only and / or chroma component)

[0332] k. Temporal layer ID

[0333] l. Configuration file / level / tier of the standard

[0334] The examples described above can be incorporated in the context of the methods described below, e.g., methods 1200 and 1300, which can be implemented at a video decoder or a video encoder.

[0335] Figure 12 A flowchart of an example method 1200 of video processing is shown. The method 1200 includes, at operation 1210, performing a conversion between a current block of a video and a bitstream representation of the video, the conversion using a dictionary-based coding mode that maintains or updates one or more dictionaries during an encoding and / or decoding process of the current block, and the conversion being based on the one or more dictionaries.

[0336] Figure 13 A flowchart of an example method 1300 of video processing is shown. The method 1300 includes, at operation 1310, performing a conversion between a current block of a video that is coded using a palette mode and a bitstream representation of the video, the conversion including saving and loading palette prediction values of a palette in the palette mode based on a use of local dual tree in the conversion.

[0337] Figure 14A is a block diagram of a video processing device 1400. The apparatus 1400 can be used to implement one or more methods described herein. The apparatus 1400 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1400 can include one or more processors 1402, one or more memories 1404, and video processing hardware 1406. The processor(s) 1402 can be configured to implement one or more methods described in the present document. The memory (memories) 1404 can be used for storing data and code used during the operation of the present methods and techniques. The video processing hardware 1406 can be used to implement, in hardware circuitry, some of the techniques described in the present document.

[0338] Figure 14Bis a block diagram illustrating an example video processing system 2100 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 2100. The system 2100 can include an input 2102 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be in a compressed or encoded format. The input 2102 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., as well as wireless interfaces, e.g., Wi-Fi or cellular interfaces.

[0339] The system 2100 can include a codec component 2104, which can implement various encoding or coding methods described in this document. The codec component 2104 can reduce the average bitrate of video from the input 2102 to the output of the codec component 2104 to produce an encoded representation of the video. Thus, the codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 2104 can be stored, or transmitted via a connected communication, as represented by component 2106. The component 2108 can use a stored or transmitted bitstream (or encoded) representation of the video received at the input 2102 to generate pixel values or displayable video sent to a display interface 2110. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Moreover, while certain video processing operations are referred to as “codec” operations or tools, it should be understood that the codec tools or operations are used at an encoder, and corresponding decoding tools or operations, which reverse the results of the encoding, will be performed by a decoder.

[0340] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices, such as mobile telephones, laptop computers, smartphones, or other devices that are capable of performing digital data processing and / or video display.

[0341] Figure 15 is a block diagram illustrating an example video coding system 100 in which techniques of this disclosure can be utilized.

[0342] As shown in Figure 15 The video coding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which can be referred to as a video decoding device.

[0343] Source device 110 can include a video source 112, a video encoder 114, and an input / output (VO) interface 116.

[0344] Video source 112 can include a source, such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. VO interface 116 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via VO interface 116 through network 130a. Video data can also be stored onto storage medium / server 130b for access by destination device 120.

[0345] Destination device 120 can include VO interface 126, video decoder 124, and display device 122.

[0346] VO interface 126 can include a receiver and / or a modem. VO interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.

[0347] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other current and / or further standards.

[0348] Figure 16 is a block diagram illustrating an example of a video encoder 200 that can be Figure 15 video encoder 114 in system 100.

[0349] Video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 16 In examples, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0350] The functional components of video encoder 200 can include partitioning unit 201, prediction unit 202 which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.

[0351] In other examples, video encoder 200 can include more, less, or different functional components. In an example, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0352] Furthermore, some components, such as motion estimation unit 204 and motion compensation unit 205, can be highly integrated but are represented separately for illustrative purposes. Figure 16

[0353] Partitioning unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.

[0354] Mode select unit 203 can select one of the coding modes (intra or inter), for example, based on error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode select unit 203 can select a combination of intra and inter prediction (CIIP) mode, in which the prediction is based on both inter prediction signals and intra prediction signals. In the case of inter prediction, mode select unit 203 can also select the resolution of the motion vectors for the block (e.g., sub-pixel or integer pixel precision).

[0355] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples for pictures from buffer 213 other than the picture in which the current video block is associated.

[0356] Motion estimation unit 204 and motion compensation unit 205 can perform different operations for a current video block, for example, depending on whether the current video block is in an I slice, P slice, or B slice.

[0357] ​In some examples, the motion estimation unit 204 can perform uni-prediction for the current video block, and the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in List 0 or List 1. The motion estimation unit 204 can then generate a reference index indicating the reference picture in List 0 or List 1 that includes the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0358] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in List 0 and can also search for another reference video block for the current video block in a reference picture in List 1. The motion estimation unit 204 can then generate a reference index indicating the reference pictures in List 0 and List 1 that include the reference video blocks and a motion vector indicating a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.

[0359] In some examples, the motion estimation unit 204 can output a set of all motion information for a decoding process of a decoder.

[0360] In some examples, the motion estimation unit 204 can not output a set of all motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.

[0361] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.

[0362] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0363] As described above, video encoder 200 can predictively signal motion vectors. Two examples of prediction signaling techniques that can be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0364] Intra prediction unit 206 can perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0365] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block for the current video block from the current video block. The residual data for the current video block can include a residual video block that corresponds to different sample components of samples in the current video block.

[0366] In other examples, the current video block can not have residual data for the current video block, such as in skip mode, and residual generation unit 207 can not perform the subtraction operation.

[0367] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0368] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0369] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to a transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.

[0370] After reconstruction unit 212 reconstructs a video block, loop filtering operations can be performed to reduce video block artifacts in the video block.

[0371] Entropy coding unit 214 can receive data from other functional components of video encoder 200. When entropy coding unit 214 receives data, entropy coding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0372] Figure 17 is a block diagram illustrating an example of a video decoder 300 that can be Figure 15 the video decoder 114 in the system 100 shown.

[0373] The video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 17 example, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0374] In Figure 17 example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200 (e.g., Figure 16 ).

[0375] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy coded video data, and from the entropy decoded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge modes.

[0376] The motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax elements.

[0377] The motion compensation unit 302 can use an interpolation filter as used by the video encoder 20 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.

[0378] The motion compensation unit 302 can use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information to decode the encoded video sequence.

[0379] The intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form a predicted block from spatially neighboring blocks. The inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0380] The reconstruction unit 306 can add the residual block to the corresponding predicted block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation.

[0381] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, an encoder will use or implement the tool or mode in the processing of blocks of video, but can not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from blocks of video to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, a decoder will process a bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to blocks of video will be performed using the video processing tool or mode enabled based on the decision or determination.

[0382] In this document, the term “video processing” can refer to video encoding, video decoding, video compression, or video decompression. For example, during a conversion from a pixel representation of a video to a corresponding bitstream representation, video compression algorithms can be applied, and vice versa. As defined by the syntax, a bitstream representation of a current video block can correspond, for example, to bits that are co-located or interspersed at different locations within the bitstream. For example, a macroblock can be encoded according to transformed and coded error residual values, and also using bits in headers and other fields in the bitstream.

[0383] It should be understood that by allowing the use of the technology disclosed in this document, the disclosed methods and techniques will benefit video encoder and / or decoder embodiments incorporated in video processing devices, such as smartphones, laptops, desktops, and similar devices.

[0384] Some embodiments can be described using the following clause-based format. The first set of clauses shows example embodiments of the technology discussed in the previous sections.

[0385] A1. A method of video processing. Comprising performing a conversion between a current block of a video and a bitstream representation of the video, wherein the conversion uses a dictionary-based coding mode that maintains or updates one or more dictionaries during an encoding and / or decoding process of the current block, and wherein the conversion is based on the one or more dictionaries.

[0386] A2. The method of clause Al, wherein the dictionary-based coding mode is different from an inter prediction mode, an intra prediction mode, and an intra block copy (IBC) prediction mode.

[0387] A3. The method of clause Al, wherein the dictionary-based coding mode is considered as part of an existing prediction mode used to code the current block.

[0388] A4. The method of clause A3, wherein the existing prediction mode is an intra block copy (IBC) prediction mode.

[0389] A5. The method of clause A3 or A4, wherein the bitstream representation includes a first indication of a use of the existing prediction mode and a second indication of a use of the dictionary-based coding mode.

[0390] A6. The method of clause A5, wherein the second indication is context coded or bypass coded.

[0391] A7. The method of clause Al or A2, further comprising selective inclusion of a decision to make an indication of a use of the dictionary-based coding mode.

[0392] A8. The method of clause A7, wherein the indication is MODE_DCM.

[0393] A9. The method of clause A7 or A8, wherein the indication is included as a result of a determination that a size of the current block is less than or equal to MxN, wherein M and N are positive integers.

[0394] A10. The method of clause A7 or A8, wherein the indication is included as a result of a determination that a width of the current block is less than or equal to M and / or a height of the current block is less than or equal to N, wherein M and N are positive integers.

[0395] A11. The method of clause A7 or A8, wherein the indication is included as a result of a determination that a vertical coordinate and / or a horizontal coordinate of a top-left sample of the current block is not equal to zero.

[0396] A12. The method of clause A7 or A8, wherein the indication is included as a result of a determination that the current block is not contained in a first coding tree unit (CTU).

[0397] A13. The method of clause A7 or A8, wherein the indication is included as a result of a determination that the current block is a luma block that includes a dual tree coding structure.

[0398] A14. The method of clause A1, wherein one or more dictionaries for the current block are selected based on a size of the current block and / or information in the bitstream representation.

[0399] A15. The method of clause A14, wherein a first dictionary of the one or more dictionaries is used for blocks having a size of KxL, wherein a second dictionary different from the first dictionary is used for blocks having a size of MxN, and wherein M≠K and / or N≠L

[0400] A16. The method of any of clauses A1 to A15, wherein each of the one or more dictionaries comprises one or more entries.

[0401] A17. The method of clause A16, wherein a maximum number of entries in at least one of the one or more dictionaries is predefined or signaled in the bitstream representation.

[0402] A18. The method of clause A16, wherein at least one of the one or more entries comprises a plurality of samples or pixels.

[0403] A19. The method of clause A16, wherein at least one of the one or more entries comprises a block from a reconstructed region.

[0404] A20. The method of clause A1, wherein the one or more dictionaries are reset prior to coding a slice, tile, or picture that includes the current block, and wherein the one or more dictionaries are updated after coding a video unit.

[0405] A21. The method of clause A20, wherein the video unit is the current block, or a coding unit (CU), coding tree unit (CTU), or CTU row associated with the current block.

[0406] A22. The method of clause A20, wherein at least one of the one or more dictionaries is updated based on overlapping blocks having different block sizes contained in the video unit.

[0407] A23. The method of clause A1, further comprising, prior to the converting, sorting entries in the one or more dictionaries.

[0408] A24. The method of clause A23, wherein the sorting is based on a distortion measure between a template of the entries in the one or more dictionaries and a template of the current block.

[0409] A25. The method of clause A24, wherein the distortion measure is based on a sum of absolute difference (SAD) calculation, a sum of absolute transformed difference (SATD) calculation, a sum of squared error (SSE) calculation, or a mean squared error (MSE) calculation.

[0410] A26. The method of clause A23, wherein the ordering is in descending order, and wherein only the top K entries are used in the conversion.

[0411] A27. The method of clause A1, wherein an index of one or more entries of one or more dictionaries is implicitly or explicitly signaled in a bitstream representation.

[0412] A28. The method of clause A27, wherein the index is binarized using a truncated unary coding, a truncated binary coding, a fixed length coding, or an exponential Golomb coding.

[0413] A29. The method of clause A27, wherein the index is binarized based on a number of one or more entries in one of the one or more dictionaries.

[0414] A30. A method of video processing, comprising: performing a conversion between a current block of a video that is coded using a palette mode and a bitstream representation of the video, wherein the conversion includes saving and loading palette prediction values of a palette in the palette mode based on a use of local double tree in the conversion.

[0415] A31. The method of clause A30, wherein the palette prediction values are saved to a temporary buffer upon determining that a single tree is switched to a local double tree.

[0416] A32. The method of clause A31, wherein the palette prediction values are reset before being used to decode a first block in the local double tree.

[0417] A33. The method of clause A30, wherein the palette prediction values are loaded from the temporary buffer upon determining that the local double tree is switched to the single tree.

[0418] A34. The method of clause A30, wherein a maximum allowed size of the palette is signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block.

[0419] A35. The method of clause A30, wherein a maximum allowed size of the palette prediction value is signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a line of LCUs, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block.

[0420] A36. The method of any of clauses A1 to A35, wherein performing the conversion is further based on at least one of: (a) screen content or natural content associated with the current block, (b) a message signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a line of LCUs, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block, (c) a location of at least one of the CU, the PU, the TU, the current block, or the video coding unit, (d) a height or a width of the current block or a neighboring block, (e) a shape of the current block or a neighboring block, (f) a quantization parameter (QP) of the current block, (g) an indication of a color format of the current block, (h) a coding tree structure applied to the current block, (i) a slice or tile group type or a picture type of a slice, tile, or picture that includes the current block, respectively, (j) a color component of a color representation of the video, (k) a temporal layer identification (ID), or (1) a profile, a level, or a tier of a standard associated with the conversion.

[0421] A37. The method of any of clauses A1 to A36, wherein the conversion generates the current block from a bitstream representation.

[0422] A38. The method of any of clauses A1 to A36, wherein the conversion generates a bitstream representation from the current block.

[0423] A39. An apparatus in a video system comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of any of clauses A1 to A38.

[0424] A40. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for implementing the method of any of clauses A1 to A38.

[0425] The second set of clauses show example embodiments of the techniques discussed in the previous sections (e.g., embodiments 1 to 7).

[0426] 1. A method of video processing (e.g., method 1810 as shown in FIG. 18), comprising performing 1812 a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using one or more dictionaries in a dictionary-based coding mode, and wherein the conversion is based on the one or more dictionaries. Figure 18A

[0427] 2. The method of clause 1, wherein the dictionary-based coding mode is signaled in the bitstream representation separately from an inter prediction mode, an intra prediction mode, and an intra block copy (IBC) prediction mode.

[0428] 3. The method of clause 1, wherein the dictionary-based coding mode is considered as part of an existing prediction mode for the video.

[0429] 4. The method of clause 3, the existing prediction mode comprising an IBC mode.

[0430] 5. The method of clause 3 or 4, wherein the bitstream representation includes a first indication of use of the existing prediction mode and a second indication of use of the dictionary-based coding mode, the second indication being signaled based on the first indication.

[0431] 6. The method of clause 5, wherein the second indication is context coded or bypass coded.

[0432] 7. The method of clause 1 or 2, wherein the bitstream representation includes an indication of use of the dictionary-based coding mode.

[0433] 8. The method of clause 7, wherein the dictionary-based coding mode is considered as a mode separate from an existing prediction mode for the video.

[0434] 9. The method of clause 1 or 2, wherein the bitstream representation includes the indication of use of the dictionary-based coding mode based on dimensions of the current block.

[0435] 10. The method of clause 9, wherein the indication is included in the bitstream representation due to a size of the current block being less than or equal to MxN, wherein M and N are positive integers.

[0436] 11. The method of clause 9, wherein the indication is included in the bitstream representation due to a width of the current block being less than or equal to M and / or a height of the current block being less than or equal to N, wherein M and N are positive integers.

[0437] 12. The method of clause 1 or 2, wherein the bitstream representation includes the indication of use of the dictionary-based coding mode based on a position of the current block.

[0438] ​13. The method of clause 12, wherein the indication is included in the bitstream representation due to a vertical coordinate and / or a horizontal coordinate of a top-left sample of the current block not being equal to zero.

[0439] 14. The method of clause 12, wherein the indication is included in the bitstream representation due to the current block not being contained in a first coding tree unit (CTU).

[0440] 15. The method of clause 1 or 2, wherein the bitstream representation includes an indication of use of a dictionary-based coding mode based on previously coded information.

[0441] 16. The method of clause 15, wherein the indication is included in the bitstream representation due to the current block being a luma block when a dual tree partitioning structure is applied.

[0442] 17. A method of video processing (e.g., the method 1820 as shown in Figure 18B FIG. 18), comprising, for a conversion between a video comprising a video block and a bitstream representation of the video, determining 1822, according to a rule, one or more dictionaries to use for the video block based on one or more coding properties of the video block, and performing 1824 the conversion based on the determination, wherein the coding properties comprise a size of the video block and / or information in the bitstream representation.

[0443] 18. The method of clause 17, wherein the rule specifies to use a plurality of dictionaries for the video block and another video block having the same coding properties.

[0444] 19. The method of clause 18, wherein the bitstream representation includes an index of the plurality of dictionaries.

[0445] 20. The method of clause 17, wherein the rule specifies a first dictionary of the one or more dictionaries to use for a video block having a size of KxL, wherein a second dictionary different from the first dictionary is to use for another video block having a size of MxN, and wherein M≠K and / or N≠L

[0446] 21. The method of any of the preceding clauses, wherein each of the one or more dictionaries comprises one or more entries.

[0447] 22. The method of clause 21, wherein a maximum number of entries in at least one of the one or more dictionaries is predefined or signaled in the bitstream representation.

[0448] 23. The method of clause 21, wherein at least one of the one or more entries comprises a plurality of samples or pixels.

[0449] 24. The method of clause 21, wherein at least one of the one or more entries comprises a block from a reconstructed region.

[0450] 25. A method of video processing (e.g., as shown in FIG. 18), comprising determining 1832, for a current block of a video for which a dictionary-based coding mode is applied, a prediction block for the current block based on one or more entries of a dictionary; and performing 1834 a conversion between the current block and a bitstream representation of the video based on the determining. Figure 18C

[0451] 26. The method of clause 25, wherein the prediction block is determined using an entry of the dictionary.

[0452] 27. The method of clause 26, wherein an index of the entry is included in the bitstream representation.

[0453] 28. The method of clause 26, wherein an index of the entry is inferred to be N, where N is an integer.

[0454] 29. The method of clause 25, wherein how the entry is selected is determined by minimizing a particular cost.

[0455] 30. The method of clause 29, wherein the particular cost corresponds to a rate-distortion cost or a rate-distortion characteristic between the entry and the current block.

[0456] 31. The method of clause 29, wherein the particular cost corresponds to a distortion between the entry and the current block.

[0457] 32. The method of clause 25, wherein the prediction block is determined based on only a plurality of entries of the dictionary.

[0458] 33. The method of clause 25, wherein the prediction block is determined based on a plurality of entries of the dictionary and a reconstructed region in a current picture that includes the current block.

[0459] 34. The method of clause 25, wherein a corresponding residual block between the current block and the prediction block is transformed, quantized, and / or entropy coded.

[0460] 35. A method of video processing, comprising performing a conversion between a current video unit of a current video region of a video and a bitstream representation of the video, wherein the current video unit is coded in a dictionary-based coding mode using a dictionary, and wherein the dictionary is reset prior to coding the current video region and updated after coding the current video unit.

[0461] 36. The method of clause 35, wherein the current video unit is a video block, a coding unit (CU), a coding tree unit (CTU), or a row of CTUs, and the current video region is a slice, a tile, or a picture.

[0462] 37. The method of clause 35, wherein the dictionary is updated based on overlapping video blocks of different sizes contained in the current video region. ​

[0463] 38. The method of clause 35, wherein the dictionary is updated based on overlapping video blocks of size equal to MxN, where M and N are positive integers.

[0464] 39. The method of clause 35, wherein a hash function is used to exclude similar or identical video blocks from the dictionary.

[0465] 40. The method of clause 39, wherein the hash function is a CRC (Cyclic Redundancy Check) function with N bits, where N is a positive integer.

[0466] 41. The method of clause 39, wherein only the later of two video blocks having the same hash value according to a raster order is included in the dictionary.

[0467] 42. The method of clause 39, wherein both of the two video blocks having the same hash value are included in the dictionary.

[0468] 43. The method of clause 39, wherein the current video unit corresponds to a video block, and wherein a hash value of the current video unit is used to update the dictionary.

[0469] 44. The method of clause 35, wherein each entry in the dictionary is a list of blocks storing blocks having the same hash value.

[0470] 45. A method of video processing, comprising: performing a conversion between a current block of a video and a bitstream representation of the video, wherein a dictionary is used to code the current block in a dictionary-based coding mode, wherein a prediction block or a reconstructed block of the current block is derived by sorting entries in the dictionary.

[0471] 46. The method of clause 45, wherein the sorting is based on a distortion measure between a template of the entries in the dictionary and a template of the current block.

[0472] 47. The method of clause 46, wherein, for a current block having size S1xS2, the template is represented as an MxN region other than a region corresponding to size S1xS2, where S1, S2, M and N are positive integers, M > S1, and M > S2.

[0473] 48. The method of clause 46, wherein the distortion measure is based on a sum of absolute difference (SAD) calculation, a sum of absolute transformed difference (SATD) calculation, a sum of squared error (SSE) calculation, or a mean squared error (MSE) calculation.

[0474] 49. The method of clause 46, wherein the entries include the current block in the reconstructed region and the template of the current block.

[0475] 50. The method of clause 45, wherein the sorting is in ascending or descending order based on a template distortion cost.

[0476] 51. The method of clause 45, wherein the ordering is in descending order, and wherein only the top K entries are used in the conversion.

[0477] 52. A method of video processing, comprising: performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using a dictionary-based coding mode using a dictionary, and wherein an index of one or more entries of the dictionary is included in the bitstream representation.

[0478] 53. The method of clause 52, wherein the index is binarized using a truncated unary coding, a truncated binary coding, a fixed length coding, or an exponential-Golomb coding.

[0479] 54. The method of clause 52, wherein the index is binarized based on a number of one or more entries in the dictionary.

[0480] 55. The method of clause 54, wherein the binarization of the index has a fixed length of M bits.

[0481] 56. The method of clause 55, wherein M is equal to floor(log2(N)), where log2(N) is a function that obtains a base-2 logarithm of N, and floor(N) is a function that obtains the nearest integer that is an upper bound of N.

[0482] 57. The method of clause 54, wherein the index is binarized using a truncated binary coding or a truncated unary coding, where cMax is equal to N.

[0483] 58. The method of clause 56 or 57, wherein N is a number of all entries or available entries in the dictionary.

[0484] 59. A method of video processing, performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current block is coded using a palette mode, and wherein the bitstream representation includes a syntax element that represents an escape flag for each of a set of coefficients of the current block.

[0485] 60. The method of clause 59, wherein the syntax element is signaled at a set of coefficients level.

[0486] 61. The method of clause 59 or 60, wherein the syntax element includes a flag that indicates whether an escape flag for all sets of coefficients is false.

[0487] 62. The method of clause 59 or 60, wherein an index of a syntax element that is equal to true is signaled.

[0488] 63. The method of clause 59 or 60, wherein an index of a syntax element that is equal to false is signaled.

[0489] 64. The method of clause 59, wherein another syntax element is indicated at the video block level and the syntax element is skipped in case the other syntax element indicates that there is no escape sample for the current block.

[0490] 65. The method of clause 59, wherein values of the syntax elements of all coefficient groups are concatenated and coded together.

[0491] 66. The method of clause 65, wherein a variable (E) is used to represent the syntax elements of the coefficient group of the current block.

[0492] 67. The method of clause 66, wherein the variable (E) is binarized using fixed length coding, truncated unary coding or exponential-Golomb coding.

[0493] 68. The method of clause 66, wherein the variable (E) is bypass coded or context coded.

[0494] 69. The method of any of the preceding clauses, wherein M and N are equal to 4.

[0495] 70. The method of any of the preceding clauses, wherein N corresponds to a predefined constant value of a quantization parameter.

[0496] 71. The method of any of the preceding clauses, wherein N is included in a bitstream representation of the video.

[0497] 72. The method of any of the preceding clauses, wherein N is based on at least one of:

[0498] (a) screen content or natural content associated with the current block,

[0499] (b) a message signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block or a video coding unit associated with the current block,

[0500] (c) a position of at least one of the CU, the PU, the TU, the current block, or the video coding unit,

[0501] (d) a height or a width of the current block or a neighboring block,

[0502] (e) a shape of the current block or a neighboring block,

[0503] (f) a quantization parameter (QP) of the current block,

[0504] (g) an indication of a color format of the current block,

[0505] (h) a coding tree structure applied to the current block,

[0506] (i) a slice, tile, or group type of a slice or tile that respectively includes the current block, or a picture type,

[0507] (j) a color component of a color representation of the video, or

[0508] (k) a temporal layer identification (ID).

[0509] 73. The method of any of clauses 1-72, wherein the conversion comprises encoding the video into a bitstream representation.

[0510] 74. The method of any of clauses 1-72, wherein the conversion comprises decoding the video from a bitstream representation.

[0511] 75. A video processing apparatus comprising a processor configured to implement a method of any one or more of clauses 1-74.

[0512] 76. A computer-readable medium storing program code that, when executed, causes a processor to implement a method of any one or more of clauses 1-75.

[0513] 77. A computer-readable medium storing an encoded representation or bitstream representation generated according to any of the above methods.

[0514] The third set of clauses show example embodiments of techniques discussed in the previous sections (e.g., embodiments 8-12).

[0515] 1. A method of video processing, comprising: performing a conversion between a current block of a video and a bitstream representation of the video, wherein the current video block is coded using a palette mode in which a palette of representative sample values is used to represent the current block; wherein the conversion includes selectively saving and loading palette prediction values of the palette in the palette mode based on use of local double tree in the conversion.

[0516] 2. The method of clause 1, wherein the palette prediction values are saved to a temporary buffer upon determining that a single tree is switched to a local double tree.

[0517] 3. The method of clause 2, wherein the palette prediction values are reset prior to being used to decode a first block in the local double tree.

[0518] 4. The method of clause 1, wherein the palette prediction values are loaded from the temporary buffer upon determining that the local double tree is switched to the single tree.

[0519] 5. The method of clause 1, wherein the maximum allowed size of the palette is signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a line of LCUs, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block.

[0520] 6. The method of clause 1, wherein the maximum allowed size of the palette prediction value is signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a line of LCUs, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block.

[0521] 7. The method of any of clauses 1-6, wherein the manner in which the conversion is performed is further based on at least one of:

[0522] (a) screen content or natural content associated with the current block,

[0523] (b) a message signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a line of LCUs, a group of LCUs, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block,

[0524] (c) a location of at least one of the CU, the PU, the TU, the current block, or the video coding unit,

[0525] (d) a height or a width of the current block or a neighboring block,

[0526] (e) a shape of the current block or a neighboring block,

[0527] (f) a quantization parameter (QP) of the current block,

[0528] (g) an indication of a color format of the current block,

[0529] (h) a coding tree structure applied to the current block,

[0530] (i) a slice or tile group type or a picture type of a slice, a tile, or a picture that respectively includes the current block,

[0531] (j) a color component of a color representation of the video,

[0532] (k) a temporal layer identification (ID), or

[0533] (l) a profile, level, or tier of a standard associated with the conversion.

[0534] 8. The method of any of clauses 1 to 7, wherein the conversion comprises encoding the current block into a bitstream representation.

[0535] 9. The method of any of clauses 1 to 7, wherein the conversion comprises decoding the current block from a bitstream representation.

[0536] 10. A video processing apparatus comprising a processor configured to implement a method of any one or more of clauses 1 to 9.

[0537] 11. A computer readable medium storing program code that, when executed, causes a processor to implement a method of any one or more of clauses 1 to 9.

[0538] 12. A computer readable medium storing an encoded representation or bitstream representation generated according to any of the above methods.

[0539] The disclosed and other clauses, solutions, examples, embodiments, modules and functional operations described in this document can be realized in digital electronic circuitry, or in a computer software, firmware, or hardware, including the structural equivalents of the disclosure disclosed herein, or in combinations of one or more of them. The disclosed and other embodiments can be realized as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of one or more of them, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.

[0540] A computer program, which can also be referred to or referred to as a program, software, a software application, an app, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.

[0541] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and / or devices can be implemented as special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0542] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0543] Although this patent document contains many details, these details should not be construed as limitations on the scope of any subject matter or of what is claimed, but rather as descriptions of features unique to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, alone or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed as such, one or more features from a claimed combination may in some cases be deleted from that combination, and a claimed combination may be directed to subcombinations or variations of subcombinations.

[0544] Similarly, while operations may be described in a particular order in the drawings, this should not be understood as requiring that these operations be performed in the particular order or order shown, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0545] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method of video processing, comprising: performing a conversion between a current block of a video and a bitstream of the video, wherein the current block is coded using a palette mode in which the current block is represented using a palette of representative sample values; wherein the conversion includes selectively saving and loading palette predictor values of the palette in the palette mode based on a use of a local dual tree in the conversion, wherein the use of the local dual tree includes a single tree switched to a local dual tree or a local dual tree switched to a single tree; wherein the palette predictor values are saved to a temporary buffer upon a determination that the single tree is switched to the local dual tree; wherein the palette predictor values are loaded from the temporary buffer upon a determination that the local dual tree is switched to the single tree.

2. The method of claim 1, wherein the palette predictor values are reset before being used to decode a first block in the local dual tree.

3. The method of claim 1, wherein a maximum allowed size of the palette is signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a LCU group, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block.

4. The method of claim 1, wherein a maximum allowed size of the palette predictor values is signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a LCU group, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block.

5. The method of claim 1, wherein a manner in which the conversion is performed is further based on at least one of: (a) screen content or natural content associated with the current block, (b) a message signaled in a decoder parameter set (DPS), a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a LCU group, a transform unit (TU), or a prediction unit (PU) block, or a video coding unit associated with the current block, (c) a position of at least one of a CU, a PU, a TU, the current block, or a video coding unit, (d) a height or a width of the current block or a neighboring block, (e) a shape of the current block or a neighboring block, (f) a quantization parameter (QP) of the current block, (g) an indication of a color format of the current block, (h) a coding tree structure applied to the current block, (i) a slice, tile or group type of a slice or tile group, or a picture type of a slice, tile or picture that includes the current block, respectively, (j) a color component of a color representation of the video, (k) a temporal layer identification (ID), or (l) a profile, level or tier of a standard associated with the conversion.

6. The method of any of claims 1 to 5, wherein the conversion comprises encoding the current block into the bitstream.

7. The method of any of claims 1 to 5, wherein the conversion comprises decoding the current block from the bitstream.

8. A video processing apparatus comprising a processor configured to implement a method recited in any of claims 1 to 7.

9. A computer readable medium storing program code that, when executed, causes a processor to implement a method recited in any of claims 1 to 7.

10. A computer readable medium for storing a bitstream generated by a method recited in any of claims 1 to 7, when executed by a video processing device.

Citation Information

Patent Citations

  • Palette mode coding for video coding

    US20160227239A1

  • Methods and apparatus for palette coding

    US20190281311A1