Cluster-based palette patterns for video encoding and decoding

By deriving the palette clustering step size in video encoding and decoding, and utilizing the rate and distortion cost of escaped encoding and decoding samples, the problem of the palette pattern clustering step size depending on the quantization step size in the prior art is solved, thus optimizing the encoding and decoding efficiency under the dual-tree structure and improving the encoding and decoding quality.

CN115211118BActive Publication Date: 2026-03-13DOUYIN VISION CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In existing video encoding and decoding technologies, the clustering step size design of the palette mode relies on the quantization step size, and the clustering step size of the luminance and chrominance channels cannot be effectively distinguished when the dual-tree is on and off, resulting in poor encoding and decoding efficiency.

Method used

By deriving the palette clustering step size, and utilizing the estimated rate and distortion cost of the escaped encoding/decoding samples, a suitable clustering step size is determined to adapt to different color components and segmentation tree types, thereby optimizing the encoding/decoding process.

Benefits of technology

It improves the efficiency and quality of video encoding and decoding, especially in the dual-tree structure, where the clustering step size of the luminance and chrominance channels is more precisely adjusted, thus enhancing encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115211118B_ABST
    Figure CN115211118B_ABST
Patent Text Reader

Abstract

Video encoding and decoding apparatus, methods, and systems are described. An example method for encoding and decoding video data includes: determining one or more clustering steps for a palette pattern representation of one or more component blocks of the current video block, for a conversion between a current video block and the video bitstream; and performing the conversion using the one or more clustering steps. The one or more clustering steps are derived from encoding / decoding characteristics according to rules.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority and interest in PCT application PCT / CN2019 / 130368, filed December 31, 2019, in accordance with the patent law and / or rules applicable under the Paris Convention. For all purposes required by law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth on the internet and other digital communication networks. As the number of networked user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to perform video encoding or decoding using improved palette encoding / decoding modes.

[0006] In one example aspect, a video processing method is disclosed. The method includes: determining one or more clustering steps for a palette pattern representation of one or more component blocks of the current video block, for a conversion between a current video block and a bitstream of the video; and performing the conversion using the one or more clustering steps; wherein the one or more clustering steps are derived from encoding / decoding characteristics according to rules.

[0007] In another example, a video processing method is disclosed. The method includes: determining one or more clustering steps for a palette mode representation of one or more component blocks of the current video block, for a conversion between a current video block and a bitstream of the video; and performing a conversion based on the one or more clustering steps; wherein the one or more clustering steps are determined according to an index of one or more lists of quantization steps or quantization parameters of the current video block.

[0008] In another example, a video processing method is disclosed. The method includes: performing a conversion between a current video block and a bitstream of the video; determining, based on encoding / decoding conditions, whether to use a predictor palette for predicting a palette for the conversion; and performing the conversion based on that determination.

[0009] In yet another example, a video encoder apparatus is disclosed. The video encoder apparatus includes a processor configured to implement the methods described above.

[0010] In yet another example, a video decoder apparatus is disclosed. The video decoder apparatus includes a processor configured to implement the methods described above.

[0011] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0012] In yet another example, a non-transitory computer-readable recording medium is disclosed that stores a bitstream of video generated by the above-described method performed by a video processing apparatus.

[0013] These features, along with others, will be described throughout this document. Attached Figure Description

[0014] Figure 1 An example of a block encoded and decoded in palette mode is shown.

[0015] Figure 2 An example of using a predictor palette to signal palette entries is shown.

[0016] Figure 3 Examples of horizontal and vertical traversal scans are shown.

[0017] Figure 4 An example of encoding and decoding a palette index is shown.

[0018] Figure 5 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0019] Figure 6 This is a block diagram of an example hardware platform used for video processing.

[0020] Figure 7 This is a block diagram illustrating an example video codec system.

[0021] Figure 8 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0022] Figure 9 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0023] Figure 10 This is a flowchart of an example method for video processing.

[0024] Figure 11 This is a flowchart of an example method for video data encoding and decoding.

[0025] Figure 12 This is a flowchart of an example method for video data encoding and decoding.

[0026] Figure 13 This is a flowchart of an example method for video data encoding and decoding. Detailed Implementation

[0027] The use of chapter headings in this document is for ease of understanding and not to limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.

[0028] 1. Preliminary Discussion

[0029] This document relates to video codec technology. Specifically, it concerns palette mode codec. It can be applied to existing video codec standards (such as HEVC) or upcoming standards (General Video Codec). It can also be applied to future video codec standards or video codecs.

[0030] 2. Introduction to Video Encoding and Decoding

[0031] Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC[1]. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. In order to explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM)[1]. In April 2018, the Joint Video Experts Group (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.

[0032] The latest version of the VVC draft (i.e., Universal Video Codec (Draft 7)) can be found at the following URL:

[0033] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 16_Geneva / wg11 / JVET-P2001-v14.zip

[0034] The latest reference software for VVC (called VTM) can be found at the following website:

[0035] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-7.0

[0036] 2.1 Palette Mode in HEVC Screen Content Codec Extension (HEVC-SCC)

[0037] 2.1.1 The concept of palette mode

[0038] The basic idea behind palette mode is that pixels in a CU are represented by a small set of representative color values. This set is called a palette. Samples outside the palette can be indicated by signaling followed by an escape symbol with (potentially quantized) component values. These pixels are called escape pixels. Palette mode is as follows: Figure 1 As shown. Figure 1 As shown, for each pixel with three juxtaposed components (luminance component and two chrominance components), an index of the color palette is established, and the block can be reconstructed based on the values ​​established in the color palette.

[0039] 2.1.2 Encoding and Decoding of Palette Entries

[0040] For the palette codec block, the following key aspects are introduced:

[0041] 1) Construct the current palette based on the predictor palette and new entries (if any) for signaling notification of the current palette.

[0042] 2) Divide the current samples / pixels into two categories: one (Category 1) includes samples / pixels in the current color palette, and the other (Category 2) includes samples / pixels outside the current color palette.

[0043] a. For samples / pixels in the second category, quantization is applied to the sample / pixel (at the encoder), and signaling informs the quantized value; and dequantization is applied (at the decoder).

[0044] 2.1.2.1 Predictor Palette

[0045] For the encoding and decoding of palette entries, maintain the predictor palette and update it after decoding the palette encoding and decoding blocks.

[0046] 2.1.2.1.1 Predictor Palette Initialization

[0047] The predictor palette is initialized at the beginning of each strip and each slice.

[0048] The maximum size of the palette and predictor palette is signaled in the SPS. In HEVC-SCC, the `palette_predictor_initializer_present_flag` is introduced in the PPS. When this flag is 1, the entries used to initialize the predictor palette are signaled in the bitstream.

[0049] Depending on the value of `palette_predictor_initializer_present_flag`, the predictor palette size is either reset to 0 or initialized using a predictor palette initializer entry signaled in the PPS. In HEVC-SCC, a predictor palette initializer of size 0 is enabled to allow explicit disabling of predictor palette initialization at the PPS level.

[0050] The corresponding syntax, semantics, and decoding process are defined as follows:

[0051] 7.3.2.2.3 Sequence Parameter Set Screen Content Encoding and Decoding Extended Syntax

[0052]

[0053]

[0054] A palette_mode_enabled_flag value of 1 indicates that the palette mode decoding process can be used for intra-frame blocks. A palette_mode_enabled_flag value of 0 indicates that the palette mode decoding process is not applied. When it does not exist, the value of palette_mode_enabled_flag is inferred to be 0.

[0055] palette_max_size specifies the maximum allowed palette size. When it does not exist, the value of palette_max_size is inferred to be 0.

[0056] `delta_palette_max_predictor_size` specifies the difference between the maximum permissible palette predictor size and the maximum permissible palette size. When it does not exist, the value of `delta_palette_max_predictor_size` is inferred to be 0. The variable `PaletteMaxPredictorSize` is derived as follows: `PaletteMaxPredictorSize = palette_max_size + delta_palette_max_predictor_size`

[0057] (0-1)

[0058] The requirement for bitstream consistency is that when palette_max_size equals 0, the value of delta_palette_max_predictor_size should be equal to 0.

[0059] A value of 1 for `sps_palette_predictor_initializer_present_flag` indicates that the sequence palette predictor is initialized using `sps_palette_predictor_initializers` as specified in the standard. A value of 0 for `sps_palette_predictor_initializer_flag` indicates that entries in the sequence palette predictor are initialized to 0. When it does not exist, the value of `sps_palette_predictor_initializer_flag` is inferred to be 0.

[0060] The requirement for bitstream consistency is that when palette_max_size equals 0, the value of sps_palette_predictor_initializer_present_flag should be equal to 0.

[0061] sps_num_palette_predictor_initializer_minus1 increments by 1 to specify the number of entries in the sequence palette predictor initializer.

[0062] The requirement for bitstream consistency is that the value of sps_num_palette_predictor_initializer_minus1 plus 1 should be less than or equal to PaletteMaxPredictorSize.

[0063] `sps_palette_predictor_initializers[comp][i]` specifies the value of the `comp`th component of the `i`th palette entry in the SPS, which is used to initialize the array `PredictorPaletteEntries`. For values ​​of `i` in the range of 0 to `sps_num_palette_predictor_initializer_minus1`, inclusive, the value of `sps_palette_predictor_initializer_minus1` should be in the range of 0 to (1 < 1). <BitDepth YThe range is 1-1, including 0 and (1 < 1). <BitDepth Y )–1, and the values ​​of sps_palette_predictor_initializers[1][i] and sps_palette_predictor_initializers[2][i] should be between 0 and (1< <BitDepth C The range is 1-1, including 0 and (1 < 1). <BitDepth C )–1.

[0064] 7.3.2.3.3 Image Parameter Set Screen Content Encoding / Decoding Extended Syntax

[0065]

[0066] A value of 1 for `pps_palette_predictor_initializer_present_flag` indicates that the palette predictor initializer used for images referenced by PPS is derived based on the palette predictor initializer specified by PPS. A value of 0 for `pps_palette_predictor_initializer_present_flag` indicates that the palette predictor initializer used for images referenced by PPS is inferred to be equal to those specified by the active SPS. When it does not exist, the value of `pps_palette_predictor_initializer_present_flag` is inferred to be 0.

[0067] The requirement for bitstream consistency is that when palette_max_size is equal to 0 or palette_mode_enabled_flag is equal to 0, the value of pps_palette_predictor_initializer_present_flag should be equal to 0.

[0068] pps_num_palette_predictor_initializer specifies the number of entries in the picture palette predictor initializer.

[0069] The requirement for bitstream consistency is that the value of pps_num_palette_predictor_initializer should be less than or equal to PaletteMaxPredictorSize.

[0070] The palette predictor variables are initialized as follows:

[0071] – If the codec tree unit is the first codec tree unit in the chip, then the following applies:

[0072] – The initialization procedure for the palette predictor variables is invoked as specified in Clause 9.3.2.3.

[0073] Otherwise, if entropy_coding_sync_enabled_flag equals 1, and CtbAddrInRs%PicWidthInCtbsY equals 0 or TileId[CtbAddrInTs] is not equal to TileId[CtbAddrRsToTs[CtbAddrInRs-1]], then the following applies:

[0074] – Spatial neighboring block T( Figure 9-2 The position (xNbT, yNbT) of the top-left luminance sample of the current codec block is derived using the position (x0, y0) of the top-left luminance sample of the current codec block, as shown below:

[0075] (xNbT,yNbT)=(x0+CtbSizeY,y0-CtbSizeY) (2-2)

[0076] – Call the availability derivation procedure for blocks in z-scan order as specified in Section 6.4.1, where the position (xCurr, yCurr) is set to be equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to be equal to (xNbT, yNbT) as input, and the output is assigned to availableFlagT.

[0077] – The synchronization procedures for context variables, Rice parameter initialization state, and palette predictor variables are called as follows:

[0078] – If availableFlagT equals 1, then take TableStateIdxWpp, TableMpsValWpp, TableStatCoeffWpp, PredictorPaletteSizeWpp, and TablePredictorPaletteEntriesWpp as input and invoke the synchronization procedures specified in Section 9.3.2.5 for context variables, Rice parameter initialization states, and palette predictor variables.

[0079] Otherwise, the following applies:

[0080] – The initialization procedure for the palette predictor variables is invoked as specified in Clause 9.3.2.3.

[0081] Otherwise, if CtbAddrInRs equals slice_segment_address and dependent_slice_segment_flag equals 1, then take TableStateIdxDs, TableMpsValDs, TableStatCoeffDs, PredictorPaletteSizeDs, and TablePredictorPaletteEntriesDs as input, and invoke the synchronization procedure specified in Section 9.3.2.5 for initializing the state of context variables and Rice parameters.

[0082] Otherwise, the following applies:

[0083] – The initialization procedure for the palette predictor variables is invoked as specified in Clause 9.3.2.3.

[0084] 9.3.2.3 Initialization process for palette predictor entries

[0085] The output of this process is the initialized palette predictor variables PredictorPaletteSize and PredictorPaletteEntries.

[0086] The variable numComps is derived as follows:

[0087] numComps=(ChromaArrayType==0)? 1:3 (2-3)

[0088] – If pps_palette_predictor_initializer_present_flag equals 1, then the following applies:

[0089] –PredictorPaletteSize is set to equal to pps_num_palette_predictor_initializer.

[0090] – The array PredictorPaletteEntries is derived as follows:

[0091]

[0092] – Otherwise (pps_palette_predictor_initializer_present_flag equals 0), then if sps_palette_predictor_initializer_present_flag equals 1, the following applies:

[0093] –PredictorPaletteSize is set to equal to sps_num_palette_predictor_initializer_minus1 plus 1.

[0094] – The array PredictorPaletteEntries is derived as follows:

[0095]

[0096] Otherwise (pps_palette_predictor_initializer_present_flag equals 0 and sps_palette_predictor_initializer_present_flag equals 0), then PredictorPaletteSize is set to equal to 0.

[0097] 2.1.2.1.2 Use of the Predictor Palette

[0098] For each entry in the palette predictor, a signaling notification reuses a flag to indicate whether it is part of the current palette. This is as follows: Figure 2 As shown, a zero-run-length encoding / decoding is used to send the reuse flag. Subsequently, a zero-order exponential Golomb (EG) code (i.e., EG-0) is used to signal the number of new palette entries. Finally, the component values ​​of the new palette entries are signaled.

[0099] 2.1.2.2 Update of the predictor palette

[0100] The update of the predictor palette is performed using the following steps:

[0101] 1. Before decoding the current block, there exists a predictor palette represented by PltPred0.

[0102] 2. Construct the current palette table by first inserting the entries from PltPred0, and then inserting new entries for the current palette.

[0103] 3. Construct PltPred1:

[0104] a. First, add the entries from the current palette table (which may include those from PltPred0).

[0105] b. If not full, add unreferenced entries from PltPred0 based on the ascending entry index.

[0106] 2.1.3 Palette Index Encoding and Decoding

[0107] The palette index is encoded and decoded using horizontal and vertical traversal scans, such as... Figure 3 As shown. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag. For the remainder of this section, it is assumed that the scan is horizontal.

[0108] The palette index is encoded and decoded using two palette sample modes: "COPY_LEFT" and "COPY_ABOVE". In "COPY_LEFT" mode, the palette index is assigned to the decoded index. In "COPY_ABOVE" mode, the palette index of the sample in the upper row is copied. For both "COPY_LEFT" and "COPY_ABOVE" modes, a signaling run value is used to specify the number of subsequent samples that are also encoded and decoded using the same mode.

[0109] In palette mode, the index value of the escaped sample is the number of palette entries. Furthermore, when the escaped symbol is part of a run in "COPY_LEFT" or "COPY_ABOVE" mode, the escaped component value is signaled for each escaped symbol. The encoding and decoding of the palette index is as follows: Figure 4 As shown.

[0110] The syntax is performed as follows: First, the number of index values ​​for the CU is signaled. Then, truncated binary encoding / decoding is used to signal the actual index values ​​for the entire CU. Both the index number and the index values ​​are encoded / decoded in bypass mode. This groups the bypass binary bits (bins) associated with the indexes together. Then, the palette sample mode (if necessary) and run length are signaled in an interleaved manner. Finally, the component escape values ​​corresponding to the escape samples of the entire CU are grouped together and encoded / decoded in bypass mode. The binaryization of the escape samples is done using EG encoding / decoding of order three, i.e., EG-3.

[0111] The additional syntax element `last_run_type_flag` is signaled after the index value. This syntax element, combined with the index number, eliminates the need for signaling to notify the run value corresponding to the last run in the block.

[0112] In HEVC-SCC, palette mode is also enabled for 4:2:2, 4:2:0, and monochrome chroma formats. For all chroma formats, the signaling for palette entries and palette indices is almost identical. In non-monochrome formats, each palette entry consists of 3 components. In monochrome formats, each palette entry consists of a single component. For the chroma direction of double sampling, chroma samples are associated with a luminance sample index divisible by 2. After reconstructing the CU's palette index, if a sample has only a single associated component, only the first component of the palette entry is used. The only difference in the signaling is for the escaped component values. For each escaped sample, the number of escaped component values ​​signaled may vary depending on the number of components associated with that sample.

[0113] Furthermore, there is an index adjustment process in the palette index encoding and decoding. When signaling informs the palette index, the left-adjacent index or the upper-adjacent index should be different from the current index. Therefore, by eliminating one possibility, the range of the current palette index is reduced by 1. Afterwards, the index is signaled using truncated binary (TB) binaryization.

[0114] The relevant text for this section is shown below, where CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the predicted index.

[0115] The variable `PaletteIndexMap[xC][yC]` specifies the palette index, which is the index of an array represented by `CurrentPaletteEntries`. The array indices `xC` and `yC` specify the position (xC, yC) of the sample relative to the top-left luminance sample of the image. The value of `PaletteIndexMap[xC][yC]` should be in the range of 0 to `MaxPaletteIndex`, inclusive.

[0116] The variable adjustedRefPaletteIndex is derived as follows:

[0117]

[0118] When CopyAboveIndicesFlag[xC][yC] equals 0, the variable CurrPaletteIndex is deduced as follows:

[0119] if(CurrPaletteIndex>=adjustedRefPaletteIndex)

[0120] CurrPaletteIndex++

[0121] 2.1.3.1 Decoding process of the palette codec block

[0122] 1) Read the prediction information to mark which entries in the predictor palette will be reused;

[0123] (palette_predictor_run)

[0124] 2) Read the new palette entry for the current block.

[0125] a.num_signalled_palette_entries

[0126] b.new_palette_entries

[0127] 3) Construct CurrentPaletteEntries based on a) and b)

[0128] 4) Read the escape symbol presence flag: palette_escape_val_present_flag to deduce MaxPaletteIndex.

[0129] 5) Encode and decode the samples that are not encoded / decoded using copy mode / run-length mode.

[0130] a.num_palette_indices_minus1

[0131] b. For each sample that is not encoded / decoded using copy mode / run-length mode, encode / decode palette_idx_idc in the current plt table.

[0132] 2.2 Palette Mode in VVC

[0133] 2.2.1 Palette in Two Trees

[0134] In VVC, the dual-tree encoding / decoding structure is used for encoding and decoding intra-frame stripes, so the luma component and the two chroma components can have different palettes and palette indices. Furthermore, the two chroma components share the same palette and palette index.

[0135] 2.2.2 Palette as a Separate Mode

[0136] In JVET-N0258[2] and the current VTM, the prediction mode of the codec unit can be MODE_INTRA, MODE_INTER, MODE_IBC and MODE_PLT. The binary representation of the prediction mode changes accordingly.

[0137] When IBC is disabled, on the I-chip, the first bit is used to indicate whether the current prediction mode is MODE_PLT. On the P / B-chip, the first bit is used to indicate whether the current prediction mode is MODE_INTRA. If not, an additional bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTER.

[0138] When IBC is enabled, on the I-chip, the first bit is used to indicate whether the current prediction mode is MODE_IBC. If not, the second bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. On the P / B-chip, the first bit is used to indicate whether the current prediction mode is MODE_INTRA. If it is an intra-frame mode, the second bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. If not, the second bit is used to indicate whether the current prediction mode is MODE_IBC or MODE_INTER.

[0139] The relevant text in JVET-O2001-vE is shown below.

[0140] Encoder / decoder unit syntax

[0141]

[0142]

[0143] 2.2.3 Encoder Algorithm for Palette Mode Encoding and Decoding

[0144] Using the palette prediction pattern, the dominant color is clustered into a palette and encoded / decoded using an entry index and color value. Escaped colors are quantized, binaryized, and signaled. Palette generation is non-standardized but critical to encoding / decoding performance. Similar colors can be grouped into an entry if the absolute difference between the sample color and the entry color is less than a preset step size. In VTM-7.0[1], the palette clustering step size is represented by a lookup table indexed by the basic QP.

[0145] Table 2-1 shows the original g_paletteQuant, which represents the corresponding clustering step size for a given QP in VTM-7.0 for different QPs.

[0146] Table 2-1: g_paletteQuant in VTM-7.0

[0147] QP 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 g_paletteQuant 0 0 0 0 1 1 1 2 2 2 3 3 3 4 4 4 5 5 QP 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 g_paletteQuant 5 6 6 7 7 8 9 9 10 11 12 13 14 15 16 17 19 21 QP 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 g_paletteQuant 22 24 23 25 26 28 29 31 32 34 36 37 39 41 42 45

[0148] 3 Examples of technical problems solved by the solutions provided in this article

[0149] The current design of the palette mode codec has the following problems:

[0150] 1. Existing designs for palette clustering step sizes depend on quantization step sizes. Note the correlation between clustering step sizes and escape encoding / decoding methods (i.e., how escaped samples are encoded / decoded). How to utilize this information is currently unknown.

[0151] 2. For both dual-tree enabled and dual-tree disabled scenarios, the palette clustering step size for all three channels is always the same. Note that when dual-tree enabled, the palette clustering processes for the luminance and chroma channels are separate, and the relationship between rate and distortion differs for both dual-tree enabled and dual-tree disabled scenarios. How to utilize this information is currently unknown.

[0152] 4 Example Technologies and Embodiments

[0153] To address the aforementioned issues, the palette clustering step size is derived from the estimated rate and distortion cost of escaping the encoding / decoding samples. The palette clustering step size (for the i-th color component, it is determined by pltQstep) is... i The distance between the three-channel color samples (S0, S1, S2) and the cluster centroids (C0, C1, C2) is typically set as the maximum permissible distance between the color samples and the cluster centroids. In one example, the distance between the three-channel color samples (S0, S1, S2) and the cluster centroids (C0, C1, C2) can be described as follows.

[0154] D0 = |S0 - C0|

[0155] D1 = |S1 - C1|

[0156] D2=|S2-C2|

[0157] Samples will be assigned to the nearest centroid. However, if the minimum distance of any color component is still greater than pltQstep... i If so, a new entry will be created.

[0158] The list below should be considered as examples used to explain the overall concepts. These items should not be interpreted in a narrow way. Furthermore, these items can be combined in any way.

[0159] 1. By pltQstep i The determination of the represented palette clustering step size may depend on the color components and / or the segmentation tree type (e.g., dual-tree or single-tree) and / or how many color components are associated with an escape value and / or the strip or picture type (e.g., intra-frame or inter-frame strip / picture).

[0160] a. For dual-tree enabled or dual-tree disabled scenarios, different color components / channels i can share the same or different pltQsteps. i .

[0161] b. In one example, color components represented by C0 (e.g., Y in YCbCr, G in RGB), C1 (e.g., Cb in YCbCr, B in RGB), and C2 (e.g., Cr in YCbCr, R in RGB) can use the same clustering step size pltQstep. i (i = 0, 1, or 2).

[0162] c. In one example, components C1 and C2 can use the same clustering step size pltQstep. i (i = 1 or 2), while C0 utilizes pltQstep. i (i = 1 or 2) Same pltQstep2.

[0163] d. In one example, components C0, C1, and C2 can use three different clustering steps.

[0164] e. Alternatively, the above bullet points can also be applied to both double-tree and single-tree cases.

[0165] 2. Palette clustering step size pltQstep i It can be determined by the binary conversion method of the escaped encoding / decoding samples.

[0166] a. In one example, for fixed-length codec escape samples, the corresponding clustering step size may differ from that of truncated binary codec escape samples.

[0167] 3. For the i-th color component, by pltQstep i The represented palette clustering step size can depend on the rate-distortion (RD) cost of the outgoing codec samples.

[0168] a. In one example, the palette clustering step size can be obtained using the following equation:

[0169]

[0170] i. In one example, A, B, or C equals 1.

[0171] ii. In one example, B equals 0 and C equals 2.

[0172] iii. In one example, A equals 1 / 4, B equals 0, and C equals 2.

[0173] iv. In one example, A equals 1, B equals 0, and C equals 2.

[0174] v. In one example, A, B, or C equals 2. K or -2 K , where K is an integer value, such as a value in the range [-M, N], where M and N are not less than 0.

[0175] b. In one example, the RD cost of the escaped codec sample of the i-th color component can be estimated by the following equation:

[0176]

[0177] Where D i esc λ represents the estimated distortion of the escaped encoding / decoding samples, and λ i R represents the Lagrange parameter. i esc It is the estimated codec bits of the escaped codec sample.

[0178] i. In one example, variables X and Y are two variables that can represent weighting factors, and Z is the offset value.

[0179] 1) In one example, X or Y or Z equals 1.

[0180] 2) In one example, X or Y or Z equals 0.

[0181] 3) In one example, X and Y are equal to 1, and Z is equal to 0.

[0182] 4) In one example, X equals 0, Y equals 1, and Z equals 0.

[0183] 5) In one example, X, Y, or Z equals 2. K or -2 K , where K is an integer value, such as a value in the range [-M, N], where M and N are not less than 0.

[0184] c. In one example, by D i esc The distortion of the quantized escaped codec samples can be estimated and derived using the following equation:

[0185]

[0186] Qstep i This represents the quantization step size of component i, where m and n are two variables.

[0187] i. In one example, m or n equals 1.

[0188] ii. In one example, m or n equals 0.

[0189] iii. In one example, m equals 2 and n equals 2.

[0190] iv. In one example, m equals 2 and n equals 6.

[0191] v. In one example, M or N equals 2 K or -2 K , where K is an integer value, such as a value in the range [-M, N], where M and N are not less than 0.

[0192] d. In one example, by D i esc The distortion of the quantized escaped codec samples is estimated using the average quantization distortion formula below:

[0193]

[0194] Where S i D represents the total number of samples for component i. i (j) represents the quantization distortion of the j-th sample point of component i.

[0195] e. In one example, by D i esc The distortion of the quantized escape codec sample points is represented by the use of

[0196] The following equation is estimated as the maximum distortion within a block:

[0197]

[0198] f. In one example, by D i esc The distortion of the quantized escaped codec samples is estimated as the minimum distortion within a block using the following equation:

[0199]

[0200] g. In one example, by R i esc The rate of the quantized escape codec sample is estimated as minRate, which represents the minimum bit length of the quantized escape codec sample.

[0201] i. In one example, if the quantization level of the escaped sample is binary-coded using a k-order Columbus exponent encoding / decoding, then R i esc It is estimated to be k+1.

[0202] ii. In one example, if the quantization level of the escaped sample is binary-coded using a fixed-length codec, then R i esc Estimated as (bitDepth) i -log2Qstep i Qstep i This indicates the quantization step size, which can be determined based on the associated quantization parameter QP. i of And it can be deduced that bitDepth i This represents the internal encoding / decoding bit depth of component i.

[0203] h. In one example, if the quantization level of the escaped sample is binary-coded using a k-order Columbus exponent encoding / decoding, then by R i esc The rate of the quantized escape codec sample is estimated as (P*minRate+Q*maxRate) / (P+Q), where maxRate represents the maximum bit length of the quantized escape codec sample, and P and Q are two variables that can represent weighting factors.

[0204] i. In one example, P and Q are equal to 1.

[0205] ii. In one example, P or Q equals 0.

[0206] i. In one example, by R i esc The rate of the quantized escape codec sample is estimated as f*maxRate, where f can represent a scaling factor in the range [0, 1].

[0207] j. In one example, by R i esc The rate of the quantized escape codec sample is estimated as Max[maxRate-offset, minRate], where offset can represent the offset parameter.

[0208] i. In one example, offset equals 4, and minRate equals 4. maxRate depends on Qstep. i and bitDepth i Determined maximum codec level l max .

[0209] ii. In one example, offset equals 2 K , where K is an integer value, such as a value in the range [0, N], where N is not less than zero.

[0210] k. In one example, depending on the binary conversion method, by R i esc The rate of the quantized escape codec sample is estimated as any value in the range [0, maxRate].

[0211] 4. Palette clustering step size pltQstep i You can look it up directly from the lookup table indexed by quantization parameter or quantization step size.

[0212] a. In one example, pltQstep can be obtained using the relevant basic quantization parameter QP for the dual-tree opening and / or dual-tree closing cases, according to Table 4-1, Table 4-2, Table 4-3, or Table 4-4. i .

[0213] b. In one example, the pltQstep0 of component C0 can be obtained according to Table 4-1, Table 4-2, or Table 4-3, and the pltQstep1 and pltQstep2 of components C1 and C2 can be obtained according to Table 4-4 or Table 4-5 using the relevant basic quantization parameters QP for the dual-tree on and / or dual-tree off cases.

[0214] 5. For lossy encoding / decoding or for QP values ​​greater than or not less than a threshold, when pltQstep i When the value is zero, the predictor palette can be used.

[0215] a. Alternatively, for lossy encoding and decoding, when pltQstep i When the value is zero, the predictor palette can be omitted.

[0216] 6. For lossless encoding / decoding modes or for QP values ​​less than or equal to a threshold, pltQstep i It can be set to zero.

[0217] Table 4-1: Palette Clustering Step Size and Basic QP

[0218] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 0 0 1 1 1 1 1 1 2 2 2 2 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 4 4 5 5 6 6 7 8 9 9 11 12 13 15 17 17 19 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 21 24 27 30 27 31 35 39 44 49 39 44 49 55 62 70

[0219] Table 4-2: Palette Clustering Step Size and Basic QP

[0220] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 0 0 1 1 1 1 1 1 1 2 2 2 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 3 3 4 4 5 6 6 7 8 9 11 12 13 15 17 19 22 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 24 27 31 35 39 44 49 55 62 70 78 88 99 111 125 140

[0221] Table 4-3: Palette Clustering Step Size and Basic QP

[0222] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 1 1 1 1 1 1 1 1 1 1 2 2 2 2 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 4 4 5 5 6 7 8 8 10 11 12 13 15 17 19 21 24 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 27 30 34 38 43 48 54 60 68 76 85 96 108 121 135 152

[0223] Table 4-4: Palette Clustering Step Size and Basic QP

[0224] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 0 1 1 1 1 1 1 2 2 2 2 3 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 4 4 5 6 6 6 7 8 9 10 11 12 13 15 17 19 19 21 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 24 27 30 34 31 35 39 44 49 55 44 49 55 62 70 78

[0225] Table 4-5: Palette Clustering Step Size and Basic QP

[0226]

[0227]

[0228] 5 Examples

[0229] An example of designing a palette clustering step size is shown below. The palette clustering step size is derived from the estimated rate and distortion cost of escaping the encoding / decoding samples. More specifically, pltQstep i This distance is typically set as the maximum permissible distance between color samples and cluster centroids. The distance between three-channel color samples (S0, S1, S2) and cluster centroids (C0, C1, C2) can be described as follows.

[0230] D0 = |S0 - C0|

[0231] D1 = |S1 - C1|

[0232] D2=|S2-C2|

[0233] Samples will be assigned to the nearest centroid. However, if the minimum distance of any color component is still greater than pltQstep... i If so, a new entry will be created.

[0234] 5.1 Example #1

[0235] In this embodiment, b-bit depth encoding / decoding is employed, and fixed-length binary encoding / decoding is used for the escaped sample. The quantization parameter for the escaped sample is QP. i pltQstep i The following equation is derived:

[0236]

[0237] in It can be obtained through the following formula,

[0238]

[0239] For both tree-on and tree-off cases, the three channels utilize the same pltQstep. i .

[0240] 5.2 Example #2

[0241] In this embodiment, b-bit depth encoding / decoding is used, and fixed-length binary encoding / decoding is employed for the escaped sample. The quantization parameter for the escaped sample is QP.i And the associated quantization step size is Qstep. i pltQstep i The following equation is derived:

[0242]

[0243] in It can be obtained through the following formula,

[0244]

[0245] For both tree-on and tree-off cases, the three channels utilize the same pltQstep. i .

[0246] 5.3 Example #3

[0247] In this embodiment, b-bit depth encoding / decoding is used, and fixed-length binary encoding / decoding is employed for the escaped sample. The quantization parameter for the escaped sample is QP. i And the associated quantization step size is Qstep. i pltQstep i The following equation is derived:

[0248]

[0249] in It can be obtained through the following formula,

[0250]

[0251] For both tree-on and tree-off cases, the three channels utilize the same pltQstep. i .

[0252] 5.4 Example #4

[0253] In this embodiment, b-bit depth encoding / decoding is used, and fixed-length binary encoding / decoding is employed for the escaped sample. The quantization parameter for the escaped sample is QP. i And the associated quantization step size is Qstep. i The basic quantization parameters are represented as qp, pltQstep. i It is derived from the index qp using a lookup table, as shown in Table 5-1. For both double-tree enabled and disabled cases, the three channels utilize the same pltQstep. i .

[0254] Table 5-1: Palette Clustering Step Size and Basic QP

[0255] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 0 0 1 1 1 1 1 1 2 2 2 2 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 4 4 5 5 6 6 7 8 9 9 11 12 13 15 17 17 19 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 21 24 27 30 27 31 35 39 44 49 39 44 49 55 62 70

[0256] 5.5 Example #5

[0257] In this embodiment, b-bit depth encoding / decoding is employed, and 3rd-order Columbus exponent binary encoding / decoding is used for escape sample encoding / decoding. pltQstep i The following equation is derived:

[0258]

[0259] in It can be obtained through the following formula,

[0260]

[0261] For both tree-on and tree-off cases, the three channels utilize the same pltQstep. i .

[0262] 5.6 Example #6

[0263] In this embodiment, b-bit depth encoding / decoding is employed, and 3rd-order Columbus exponent binary encoding / decoding is used for escape sample encoding / decoding. pltQstep i The following equation is derived:

[0264]

[0265] in It can be obtained through the following formula,

[0266]

[0267] For both tree-on and tree-off cases, the three channels utilize the same pltQstep. i .

[0268] 5.7 Example #7

[0269] In this embodiment, b-bit depth encoding / decoding is employed, and 3rd-order Columbus exponent binaryization is used for escape sample encoding / decoding. The basic quantization parameters are represented as qp, pltQstep. i It is derived from the index qp using a lookup table, as shown in Table 5-2. For both double-tree enabled and disabled cases, the three channels utilize the same pltQstep. i .

[0270] Table 5-2: Palette Clustering Step Size and Basic QP

[0271] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 0 0 1 1 1 1 1 1 1 2 2 2 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 3 3 4 4 5 6 6 7 8 9 11 12 13 15 17 19 22 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 24 27 31 35 39 44 49 55 62 70 78 88 99 111 125 140

[0272] 5.8 Example #8

[0273] In this embodiment, b-bit depth encoding / decoding is employed, and 3rd-order Columbus exponent binaryization is used for escape sample encoding / decoding. The basic quantization parameters are represented as qp, pltQstep. i It is derived from the index qp using a lookup table, as shown in Table 5-3. For both double-tree enabled and double-tree disabled cases, the three channels utilize the same pltQstep. i .

[0274] Table 5-3: Palette Clustering Step Size and Basic QP

[0275] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 1 1 1 1 1 1 1 1 1 1 2 2 2 2 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 4 4 5 5 6 7 8 8 10 11 12 13 15 17 19 21 24 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 27 30 34 38 43 48 54 60 68 76 85 96 108 121 135 152

[0276] 5.9 Example 9

[0277] In this embodiment, b-bit depth encoding / decoding is used, and fixed-length binary encoding / decoding is employed for the escaped sample. The quantization parameter for the escaped sample is QP. i And the associated quantization step size is Qstep. i The basic quantization parameters are represented as qp, pltQstep. i It is derived from the index qp using a lookup table. The luminance channel uses the same pltQstep as in Example #1. i The chromaticity components are obtained using the following formula: pltQstep i ,

[0278]

[0279] in It can be obtained through the following formula,

[0280]

[0281] 5.10 Example #10

[0282] In this embodiment, 10-bit depth encoding / decoding is used, and fixed-length binary encoding / decoding is applied to the escaped sample points. The quantization parameter for the escaped sample points is QP. i And the associated quantization step size is Qstep. i The basic quantization parameters are represented as qp, pltQstep. i It is derived from the index qp using a lookup table. The luminance channel uses the same pltQstep as in Table 5-1. i The chromaticity components are used as the pltQstep for dual-tree on-state conditions, based on Table 5-4. i .

[0283] Table 5-4: Palette Clustering Step Size and Basic QP

[0284] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 0 1 1 1 1 1 1 2 2 2 2 3 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 4 4 5 6 6 6 7 8 9 10 11 12 13 15 17 19 19 21 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 24 27 30 34 31 35 39 44 49 55 44 49 55 62 70 78

[0285] 5.11 Example #11

[0286] In this embodiment, 10-bit depth encoding / decoding is used, and fixed-length binary encoding / decoding is applied to the escaped sample points. The quantization parameter for the escaped sample points is QP. i And the associated quantization step size is Qstep. i The basic quantization parameters are represented as qp, pltQstep. i It is derived from the index qp using a lookup table. The luminance channel uses the same pltQstep as in Table 5-2. i The chromaticity components are used as pltQstep in Table 5-5 for dual-tree on-state conditions. i .

[0287] Table 5-5: Palette Clustering Step Size and Basic QP

[0288] qp 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 <![CDATA[pltQstep i ]]> 0 0 0 0 1 1 1 1 1 1 1 1 2 2 2 3 3 3 qp 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 <![CDATA[pltQstep i ]]> 3 4 4 5 5 6 7 8 8 10 11 12 13 15 17 19 21 24 qp 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 <![CDATA[pltQstep i ]]> 27 31 35 39 44 49 55 62 70 78 88 99 111 124 140 157

[0289] Figure 5 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0290] System 1900 may include codec component 1904, which may implement the various codec or encoding methods described in this document. Codec component 1904 may reduce the average bit rate of video from input 1902 to output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1906, the output of codec component 1904 may be stored or transmitted via connected communication. Component 1908 may use the stored or transmitted bitstream (or codec) representation of the video received at input 1902 to generate pixel values ​​or displayable video to be sent to display interface 1910. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and corresponding decoding tools or operations, the opposite of codec results, will be performed by the decoder.

[0291] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0292] Figure 6 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors(multiple) 3602 can be configured to implement one or more methods described in this document. The memories(multiple) 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.

[0293] Figure 7 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0294] like Figure 7As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0295] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0296] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.

[0297] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.

[0298] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to interface with an external display device.

[0299] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.

[0300] Figure 8 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 7 The video encoder 114 in the system 100 shown.

[0301] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 8 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0302] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0303] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0304] Furthermore, some components (such as motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but for interpretative purposes... Figure 8 The examples are shown separately.

[0305] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0306] The mode selection unit 203 can (e.g., based on error results) select one of the encoding / decoding modes (intra-frame or inter-frame) and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the codec block as a reference image. In some examples, the mode selection unit 203 can select an intra-frame and inter-frame combined prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution (e.g., sub-pixel or integer pixel precision) for the block's motion vector.

[0307] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0308] The motion estimation unit 204 and the motion compensation unit 205 can (for example, depending on whether the current video block is in an I-band, P-band, or B-band) perform different operations on the current video block.

[0309] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0310] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0311] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.

[0312] In some examples, motion estimation unit 204 may not output the full set of motion information for the current video. Instead, motion estimation unit 204 may refer to the motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0313] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0314] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0315] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0316] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0317] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0318] In other examples, there may be no residual data for the current video block (e.g., in skip mode), and the residual generation unit 207 may not perform the subtraction operation.

[0319] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0320] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0321] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video blocks, respectively, to reconstruct the residual video blocks from the transform coefficient video blocks. The reconstruction unit 212 can add the reconstructed residual video blocks to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block, which is then stored in the buffer 213.

[0322] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0323] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0324] Figure 9 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 7 The video decoder 114 in the system 100 shown.

[0325] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 8 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0326] exist Figure 9 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 8 The decoding process is the inverse of the encoding process described.

[0327] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-coded video data, and from the entropy-coded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge modes.

[0328] The motion compensation unit 302 can generate motion compensation blocks, thereby potentially performing interpolation based on an interpolation filter. Identifiers for the interpolation filter used for sub-pixel precision can be included in the syntax elements.

[0329] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolated values ​​of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0330] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how the images of the encoded video sequence are divided, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0331] Intra-prediction unit 303 can use an intra-prediction mode (e.g., received in the bitstream) to form prediction blocks from spatially neighboring blocks. Dequantization unit 303 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0332] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.

[0333] The following provides a list of preferred solutions for some of the embodiments.

[0334] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Item 1 or 6).

[0335] 1. A method for video processing (e.g., Figure 10 The method 1000 described in the text includes: determining (1002) one or more clustering steps for a palette pattern representation of one or more component blocks of a video region for a conversion between a video region and a video codec representation; and performing (1004) a conversion based on the one or more clustering steps; wherein the one or more clustering steps are derived from codec characteristics.

[0336] 2. The method according to Solution 1, wherein the encoding and decoding characteristics of the video region include component identifiers of component blocks.

[0337] 3. The method according to Solution 1, wherein the encoding / decoding features include a segmentation tree type for segmenting video regions.

[0338] 4. The method according to Solution 1, wherein the encoding / decoding characteristics include the stripe type of the stripe containing the video region.

[0339] 5. The method according to any one of solutions 1-4, wherein one or more clustering steps correspond to a clustering step for all component blocks of a video region.

[0340] 6. The method according to Solution 1, wherein one or more clustering steps have zero values ​​due to encoding / decoding characteristics, (1) the encoding / decoding representation is a lossless representation or (2) the quantization parameter used for conversion is less than or not greater than a threshold.

[0341] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 2).

[0342] 7. The method according to any one of the preceding solutions, wherein the encoding / decoding features include a binary method for representing escaped encoding / decoding samples of a video region in the encoding / decoding representation.

[0343] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 3).

[0344] 8. The method according to any one of the preceding solutions, wherein the palette clustering step size of the component block depends on the rate-distortion cost of the escaped codec samples of the component block.

[0345] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 4).

[0346] 9. A video processing method comprising: converting between a video region of a video and a codec representation of the video; determining one or more clustering steps for a palette mode representation of one or more component blocks of the video region; and performing the conversion based on the one or more clustering steps; wherein the one or more clustering steps are determined according to an index of one or more lists of quantization steps or quantization parameters of the video region.

[0347] 10. The method according to Solution 9, wherein one or more lists include a first list for the two-tree open case and a second list for the two-tree closed case.

[0348] The following solutions illustrate example embodiments of the techniques discussed in the previous chapter (e.g., item 5).

[0349] 11. A method for video processing, comprising: converting a video region of a video to a codec representation of the video; determining, based on codec conditions, whether to use a predictor palette for predicting a palette for the conversion; and performing the conversion based on the determination.

[0350] 12. The method according to solution 11, wherein the encoding / decoding condition corresponds to using a quantization parameter that is greater than or not less than a threshold.

[0351] 13. The method according to solution 12, wherein the encoding / decoding conditions further include a palette step size of zero.

[0352] 14. The method according to solution 11, wherein the determination includes determining not to use the predictor palette because the palette step size is equal to zero.

[0353] 15. The method according to any one of solutions 1-14, wherein performing the conversion includes encoding the video to generate a codec representation.

[0354] 16. The method according to any one of solutions 1-14, wherein performing the conversion includes parsing and decoding the codec representation to generate a video.

[0355] 17. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 16.

[0356] 18. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 16.

[0357] 19. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 16.

[0358] 20. A method, apparatus or system described in this document.

[0359] The following list of solutions can preferably be implemented by some embodiments.

[0360] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Items 1 and 6).

[0361] 1. A method for encoding and decoding video data (e.g., Figure 11The method described in the text (1100) includes: for the conversion between the current video block and the bitstream of the video, determining (1102) one or more clustering steps for the palette pattern representation of one or more component blocks of the current video block; and performing (1104) the conversion using the one or more clustering steps; wherein the one or more clustering steps are derived from the encoding / decoding characteristics according to rules.

[0362] 2. The method according to Solution 1, wherein the encoding / decoding characteristics of the current video block include component identifiers of the component blocks.

[0363] 3. The method according to any one of solutions 1-2, wherein the encoding / decoding features include a segmentation tree type for segmenting the current video block.

[0364] 4. The method according to any one of solutions 1-3, wherein the encoding / decoding characteristics include the stripe type of the stripe containing the current video block.

[0365] 5. The method according to any one of solutions 1-4, wherein the encoding / decoding feature includes the number of components associated with an escape value.

[0366] 6. The method according to any one of solutions 1-5, wherein the rule specifies that different components of the current video block share the same value of one or more clustering steps because the current video block is segmented using a dual-tree.

[0367] 7. The method according to any one of solutions 1-5, wherein the rule specifies that different components of the current video block share different values ​​of one or more clustering steps because the current video block is not segmented using a dual tree.

[0368] 8. The method according to any one of solutions 1-7, wherein one or more clustering steps correspond to a clustering step for all component blocks of the current video block.

[0369] 9. The method according to any one of solutions 1-7, wherein the rule specifies that two components of the video share a first clustering step size, and a third component uses a second step size different from the first step size.

[0370] 10. The method according to any one of solutions 1-7, wherein the rule specifies that each video component of the video uses a different clustering step size.

[0371] 11. The method according to any one of solutions 6-10, wherein the rule is based on whether dual-tree segmentation or single-tree segmentation is applied to the current video block.

[0372] 12. The method according to any one of solutions 1-11, wherein one or more clustering steps have a zero value due to encoding / decoding characteristics, (1) the bitstream is a lossless representation or (2) the quantization parameter used for conversion is less than or no greater than a threshold.

[0373] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 2).

[0374] 13. The method according to any one of solutions 1-12, wherein the encoding / decoding features include a binary method for representing the escaped encoding / decoding samples of the current video block in the bitstream.

[0375] 14. The method according to Solution 13, wherein the rule specifies that the first clustering step of the component blocks of the current video block with fixed-length codec escape points is different from the second clustering step of another component block of the current video block that is encoded and decoded using truncated binary codec escape points.

[0376] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 3).

[0377] 15. The method according to any one of solutions 1-14, wherein the palette clustering step size of the component blocks depends on the rate-distortion cost of the escaped codec samples of the component blocks.

[0378] 16. The method according to solution 15, wherein the rule specifies that the expression pltQstep is determined according to the following equation. i One or more clustering step sizes, where i is an integer:

[0379]

[0380] in This indicates that the escaped encoding / decoding sample points are determined according to the second rule.

[0381] 17. According to the method described in Solution 16, wherein the second rule specifies: A or B or C equals 1; or B equals 0 and C equals 2; or A equals 1 / 4, B equals 0, and C equals 2; or A equals 1, B equals 0, and C equals 2; or A or B or C equals 2. K or -2 K , where K is an integer value in the range [-M, N], where M and N are not less than 0.

[0382] 18. The method according to solution 15, wherein the rate-distortion cost of the escaping codec samples of the color components is estimated according to the following equation:

[0383]

[0384] Where D i esc λ represents the estimated distortion of the escaped encoding / decoding samples. i R represents the Lagrange parameter. i esc It is an estimated number of codec bits that escape from the codec sample, and X, Y, and Z are determined according to the third rule.

[0385] 19. The method according to Solution 18, wherein the second rule specifies: X or Y or Z equals 1; or X or Y or Z equals 0; or X and Y equal 1 and Z equals 0; or X equals 0, Y equals 1 and Z equals 0; or X or Y or Z equals 2. K or -2 K , where K is an integer value.

[0386] 20. The method according to any one of solutions 15 to 19, wherein the distortion of the quantized escaped codec sample is determined using the following formula:

[0387]

[0388] Qstep i The quantization step size of component i is represented by m and n, which are two variables according to the fourth rule.

[0389] 21. The method described in Solution 20, wherein the third rule specifies: m or n equals 1; or m or n equals 0; or m equals 2 and n equals 2; or m equals 2 and n equals 6; or m or n equals 2 K or -2 K , where K is an integer value.

[0390] 22. The method according to any one of solutions 19-21, wherein K is in the range [-M, N], where M and N are not less than 0.

[0391] 23. The method according to any one of solutions 15-19, wherein the distortion of the quantized escape symbol is determined according to the following formula:

[0392]

[0393] Where S i D represents the total number of samples for component i. i (j) represents the quantization distortion of the j-th sample point of component i.

[0394] 24. The method according to any one of solutions 15-19, wherein the distortion of the quantized escape symbol is determined by estimating the maximum distortion within a block using the following equation:

[0395]

[0396] The `max` option is used to find the maximum value.

[0397] 25. The method according to any one of solutions 15-19, wherein the distortion of the quantized escape symbol is determined by estimating the minimum distortion within a block using the following equation:

[0398]

[0399] The min operation is for finding the minimum value.

[0400] 26. The method according to any one of solutions 1-24, wherein R is represented as i esc The rate of the quantized escape codec sample is estimated as minRate, where minRate represents the minimum bit length of the quantized escape codec sample.

[0401] 27. According to the method described in Solution 25, wherein if the quantization level of the escaped sample is binary-coded using a k-order Columbus exponent encoder / decoder, then R i esc It is estimated to be k+1.

[0402] 28. The method according to solution 25, wherein if the quantization level of the escaped sample is binary-coded using a fixed-length codec, then R i esc Estimated as (bitDepth) i -log2Qstep i ), Qstep i This indicates the quantization step size, which is determined based on the associated quantization parameter QP. i of And it can be deduced that bitDepth i This represents the internal encoding / decoding bit depth of component i.

[0403] 29. The method according to any one of solutions 1-24, wherein if the quantization level of the escaped sample is binaryized using a k-order Columbus exponent encoder / decoder, then by R i esc The rate of the quantized escape codec sample is estimated as (P*minRate+Q*maxRate) / (P+Q), where maxRate represents the maximum bit length of the quantized escape codec sample, and P and Q are two variables representing weighting factors.

[0404] 30. The method according to solution 28, wherein P = Q = 1 or P = Q = 0.

[0405] 31. The method according to any one of solutions 1-24, wherein R i esc The rate of the quantized escape codec sample is estimated as f*maxRate, where f represents a scaling factor in the range [0, 1].

[0406] 32. The method according to any one of solutions 1-24, wherein R i esc The rate of the quantized escape codec sample is estimated as Max[maxRate-offset, minRate], where offset represents the offset parameter.

[0407] 33. The method according to solution 32, wherein offset equals 4, and minRate equals 4, and maxRate depends on Qstep. i and bitDepth i Determined maximum codec level l max .

[0408] 34. According to the method described in solution 32, where offset equals 2 K , where K is an integer value.

[0409] 35. The method described in solution 34, wherein K is in the range [0, N], and N is not less than zero.

[0410] 36. The method according to any one of solutions 1-24, wherein, depending on the binary conversion method, R i esc The rate of the quantized escape codec sample is estimated as any value in the range [0, maxRate].

[0411] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 4).

[0412] 37. A method for encoding and decoding video data (e.g., Figure 12 The method 1200 described herein includes: for the conversion between a current video block and a bitstream of video, determining (1202) one or more clustering steps for a palette mode representation of one or more component blocks of the current video block; and performing (1204) a conversion based on the one or more clustering steps; wherein the one or more clustering steps are determined according to an index of one or more lists of quantization steps or quantization parameters of the current video block.

[0413] 38. The method according to solution 37, wherein one or more lists include a first list for the two-tree open case and a second list for the two-tree closed case.

[0414] 39. The method according to solution 37, wherein the first list for determining the first clustering step size of the first component is different from the second list for determining the second clustering step size of the second component or the third component.

[0415] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 5).

[0416] 40. A method for encoding and decoding video data (e.g., Figure 13 The method described in the text (1300) includes: for the conversion between the current video block of the video and the bitstream of the video, determining (1302) whether to use a predictor palette for predicting a palette for the conversion based on encoding and decoding conditions; and performing (1304) the conversion based on the determination.

[0417] 41. The method according to solution 40, wherein the encoding / decoding condition corresponds to using a quantization parameter that is greater than or not less than a threshold.

[0418] 42. The method according to solution 41, wherein the encoding / decoding conditions further include a palette step size of zero.

[0419] 43. The method according to solution 41, wherein determining includes determining not to use the predictor palette because the palette step size is equal to zero.

[0420] 44. The method according to any one of solutions 1-43, wherein the conversion includes encoding the current video block into a bitstream.

[0421] 45. The method according to any one of solutions 1-43, wherein the conversion includes decoding the current video block from the bitstream.

[0422] 46. ​​A method for storing a bitstream representing a video into a computer-readable recording medium, comprising: generating a bitstream from the video according to the method described in any one or more of Solutions 1 to 43; and writing the bitstream into the computer-readable recording medium.

[0423] 47. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method described in one or more of solutions 1 to 43.

[0424] 48. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: generating a bitstream from a current video block according to the method described in any one or more of solutions 1 to 43.

[0425] 49. A computer-readable medium for storing a bit stream generated according to any one or more of solutions 1 to 16.

[0426] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed herein and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of materials implementing machine-readable propagating signals, or a combination of one or more of them. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0427] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as multiple collaborating files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer, or on multiple computers located in one place or distributed across multiple locations and interconnected through a communication network.

[0428] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating output. These processes and logic flows can also be executed by special-purpose logic circuits, and the devices can be implemented as special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0429] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, to receive data from, transfer data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM discs. The processor and memory may be supplemented or incorporated therein by dedicated logic circuitry.

[0430] While this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features characteristic of specific embodiments of a particular art. Certain features described in this patent document within the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0431] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0432] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A method of coding video data, comprising: determining, for a conversion between a current video block of a video and a bitstream of the video, one or more clustering step sizes for a palette mode representation of one or more component blocks of the current video block; and performing the conversion using the one or more clustering step sizes; wherein the one or more clustering step sizes are derived from a coding characteristic according to a rule; wherein the coding characteristic comprises a split tree type used to split the current video block; and wherein the rule specifies that different components of the current video block share a same value of the one or more clustering step sizes due to the current video block being split using dual tree.

2. The method of claim 1, wherein, the coding characteristic of the current video block comprises a component identification of the component blocks.

3. The method of claim 1, wherein, the coding characteristic comprises a slice type of a slice containing the current video block.

4. The method of claim 1, wherein, the coding characteristic comprises a number of components associated with one escape value.

5. The method of claim 1, wherein, the rule specifies that different components of the current video block share different values of the one or more clustering step sizes due to the current video block not being split using dual tree.

6. The method of claim 1, wherein, the one or more clustering step sizes correspond to one clustering step size for all component blocks of the current video block.

7. The method of claim 1, wherein, the rule specifies that two components of the video share a first clustering step size and a third component uses a second step size different from the first step size.

8. The method of claim 1, wherein, the rule specifies that each video component of the video uses a different clustering step size.

9. The method of claim 1, wherein, the rule is based on whether dual tree splitting or single tree splitting is applied to the current video block.

10. The method of claim 1, wherein, the one or more clustering step sizes have zero values due to the coding characteristic being that (1) the bitstream is lossless representation or (2) a quantization parameter used for the conversion is less than or not greater than a threshold.

11. The method of claim 1, wherein, the coding characteristic comprises a binarization method used to represent escape coded samples of the current video block in the bitstream.

12. The method of claim 11, wherein, the rule specifies that a first clustering step size of a component block of the current video block having fixed length coded escape samples is different from a second clustering step size of another component block of the current video block coded using truncated binary coded escape samples.

13. The method of claim 1, wherein, a palette clustering step size of a component block depends on a rate-distortion cost of escape coded samples of the component block.

14. The method of claim 13, wherein, The rules provide that one or more palette clustering steps, denoted as pltQstep i are determined according to the following equation, where i is an integer: wherein represents the escape coded sample, and A, B and C are determined according to the second rule; wherein the second rule specifies that: A or B or C equals 1; or B equals 0 and C equals 2; or A equals 1 / 4, B equals 0 and C equals 2; or A equals 1, B equals 0 and C equals 2; or A or B or C equals 2 K or -2 K wherein K is an integer value in the range of [-M, N], where M and N are not less than 0.

15. The method of claim 13, wherein, a rate-distortion cost of escape coded samples of a color component is estimated according to the following equation: where D i esc denotes the estimated distortion of the escape coded sample, λ i denotes the Lagrangian parameter, R i esc is the estimated number of coded bits of the escape coded sample, and X, Y and Z are determined according to a third rule; wherein the third rule provides that X or Y or Z equals 1 ; or X or Y or Z equals 0; or X and Y equal 1 and Z equals 0; or X equals 0, Y equals 1 and Z equals 0; or X or Y or Z equals 2 K or -2 K wherein K is an integer value.

16. The method of claim 13, wherein, a distortion of a quantized escape coded sample is determined using the following equation: where Qstep i denotes the quantization step length for component i, m and n are two variables according to a fourth rule; wherein the fourth rule provides that: m or n is equal to 1; or m or n is equal to 0; or m is equal to 2 and n is equal to 2; or m is equal to 2 and n is equal to 6; or m or n is equal to 2 K or -2 K where K is an integer value.

17. The method of claim 15, wherein, K is in a range of [-M, N], where M and N are not less than 0.

18. The method of claim 13, wherein, a distortion of a quantized escape symbol is determined according to the following equation: where S i represents the total number of samples of component i, D i (j) represents the quantization distortion of the jth sample of component i.

19. The method of claim 13, wherein, a distortion of a quantized escape symbol is determined according to estimating a maximum distortion within one block using the following equation: where max is a maximum operation.

20. The method of claim 13, wherein, a distortion of a quantized escape symbol is determined according to estimating a minimum distortion within one block using the following equation: where min is a minimum operation.

21. The method of claim 1, wherein, R is represented as i esc The rate of quantized escape coded samples is estimated as minRate, minRate representing the minimum bit length of quantized escape coded samples.

22. The method of claim 20, wherein, If the quantization levels of the escaped samples are binarized with a k-th order Golomb exponent, then R i esc is estimated as k + 1.

23. The method of claim 20, wherein, If the quantization levels of the escape samples are binarized with a fixed length coding, then R i esc is estimated as (bitDepth i - log2Qstep i ), Qstep i denotes the quantization step size, which is derived from the associated quantization parameter QP i in accordance with the following equation: bitDepth i denotes the internal coding bit depth of component i.

24. The method of claim 1, wherein, If the quantization level of the escaped sample is binaryized using a k-order Columbus exponent encoder / decoder, then by R i esc The rate of the quantized escape codec sample is estimated as (P*minRate+Q*maxRate) / (P+Q), where maxRate represents the maximum bit length of the quantized escape codec sample, and P and Q are two variables representing weighting factors. where minRate represents a minimum bit length of the quantized escape coded samples.

25. The method of claim 23, wherein P = Q = 1 or P = Q = 0.

26. The method of claim 1, wherein, R i esc The rate of quantized escape-coded samples represented by the escape codeword is estimated to be f * maxRate, where f represents a scaling factor in the range [0, 1].

27. The method of claim 1, wherein, R i esc The rate of quantized escape-coded samples represented by the escape code is estimated as Max[maxRate - offset, minRate], where offset represents an offset parameter; where minRate represents a minimum bit length of the quantized escape coded samples, and maxRate represents a maximum bit length of the quantized escape coded samples.

28. The method of claim 27, wherein, offset equals 4 and minRate equals 4, and wherein maxRate depends on the maximum coding level l determined by Qstep i and bitDepth i determined by Qstep max .

29. The method of claim 27, wherein, offset equals 2 K where K is an integer value.

30. The method of claim 29, wherein, K is in the range of [0, N], where N is not less than zero.

31. The method of claim 1, wherein, Depending on the binarization method, the rate of quantized escape-coded samples represented by R i esc The rate of quantized escape-coded samples represented by R is estimated to be any value in the range of [0, maxRate].

32. The method of claim 1, further comprising: wherein The one or more cluster steps are further determined according to an index of one or more lists of quantization steps or quantization parameters of the current video block; wherein the one or more lists include a first list for a dual tree on case and a second list for a dual tree off case.

33. The method of claim 32, wherein, The first list for determining the first cluster step of the first component is different from the second list for determining the second cluster step of the second component or the third cluster step of the third component.

34. The method of claim 1, further comprising: for conversion between a current video block of a video and a bitstream of the video, determining whether to use a predictor palette for a prediction palette for the conversion based on a coding condition; and performing the conversion based on the determination.

35. The method of claim 34, wherein, The coding condition corresponds to using a quantization parameter that is greater than or not less than a threshold.

36. The method of claim 35, wherein, The coding condition further includes that a step size of the palette is zero.

37. The method of claim 35, wherein, The determination includes determining not to use the predictor palette due to the palette step size being equal to zero.

38. The method of claim 1, wherein, The conversion includes encoding the current video block into the bitstream.

39. The method of claim 1, wherein, The conversion includes decoding the current video block from the bitstream.

40. A method of storing a bitstream representing a video to a computer readable recording medium, comprising: generating the bitstream from the video according to the method described in any one or more of claims 1 to 39; and writing the bitstream to the computer readable recording medium.

41. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method recited in one or more of claims 1 to 39.

42. A non-transitory computer readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises: generating the bitstream from a current video block according to the method described in any one or more of claims 1 to 39.

43. A computer readable medium storing a bitstream generated according to any one or more of claims 1 to 39.

Citation Information

Patent Citations

  • Quantized pulse code modulation in video coding

    US20120224640A1

  • Restriction of escape pixel signaled values in palette mode video coding

    US20170085891A1