Palette Initialization for Wavefront Parallel Processing
The implementation of palette modes with adaptive predictor palette management in video coding improves compression efficiency and parallel processing, addressing inefficiencies in current video coding technologies.
Patent Information
- Application Number
- CN202080064297.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-09-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-09-11
AI Technical Summary
The existing video encoding and decoding technology has problems such as insufficient accuracy and poor codec performance when processing motion vectors. Especially in screen content encoding and decoding, the predictor palette is not updated with a flexible enough, resulting in waste of resources and limited parallel processing capabilities.
Optimize the update process of the predictor palette by resetting or reinitializing the predictor palette at the beginning or end of the video unit, adaptively adjusting its size and number of entries, dynamically update the predictor palette based on block characteristics and color components, and maintaining entry usage frequency using shared predictor palettes and counters.
It improves the efficiency and accuracy of video encoding and codec, reduces resource waste, enhances parallel processing capabilities, and improves the performance of video encoding and codec.
Smart Images

Figure CN114365483B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] According to applicable patent laws and / or rules applicable to the Paris Convention, this application timely claims the priority and benefits of International Patent Application PCT / CN2019 / 105554 filed on September 12, 2019. For all legal purposes, the entire disclosure of the above - mentioned application is incorporated by reference into a part of the disclosure of this application. Technical field
[0003] This patent document relates to video coding and decoding technologies, devices, and systems. Background art
[0004] Currently, efforts are being made to improve the performance of current video codec technologies to provide a better compression ratio, or to provide video coding and decoding schemes that allow for lower complexity or parallel implementation. Industry experts have recently proposed several new video coding and decoding tools, and tests are currently being conducted to determine their effectiveness. Summary of the invention
[0005] Devices, systems, and methods related to digital video coding and decoding, particularly related to the management of motion vectors, are described. The described methods can be applied to existing video coding standards (e.g., High - Efficiency Video Coding (HEVC) or Versatile Video Coding) and future video coding standards or video codecs.
[0006] In a representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and an encoded / decoded representation of the video using a palette mode, wherein a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. Based on the characteristics of the current block, updating the predictor palette is prohibited after the conversion of the current block according to a rule.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and an encoded / decoded representation of the video using a palette mode, wherein a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. Whether to perform a change to the predictor palette is determined according to the color components of the current block.
[0008] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, multiple predictor palettes are used to predict the palette of representative sample values.
[0009] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. Before the conversion of the first block in a video unit or after the conversion of the last video block in a previous video unit, the predictor palette is reset or reinitialized according to a rule.
[0010] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a video unit of a video and an encoded / decoded representation of the video using a palette mode. The video unit includes multiple blocks. During the conversion, a shared predictor palette is used by all the multiple blocks to predict the palette of representative sample values for each of the multiple blocks in a predictor palette mode.
[0011] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and a counter indicating the usage frequency of a corresponding entry is maintained for each entry of the predictor palette.
[0012] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block to predict the palette of representative sample values of the current block. The number of entries of the palette signaled in the encoded / decoded representation is in the range of [0, the maximum allowed size of the palette - the number of palette entries derived during the conversion].
[0013] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and the size of the predictor palette is adaptively adjusted according to rules.
[0014] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and the size of the palette of representative samples or the predictor palette is determined according to rules, where the rules allow the size to change between video units of the video.
[0015] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. The predictor palette is reinitialized when a condition is met, where the condition is met when the video unit is the first video unit in a video unit row and a syntax element indicating that wavefront parallel processing is enabled for the video unit is included in the encoded / decoded representation.
[0016] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block. Further, the predictor palette is selectively reset before the conversion between the video block and the bitstream representation of the video block.
[0017] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block. Further, when multiple encoded / decoded units of the video unit have a common shared area, the predictor palette is a shared predictor palette.
[0018] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the size of the predictor palette is adaptively changed according to one or more conditions.
[0019] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the predictor palette is updated based on the size or number of entries in the predictor palette.
[0020] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the entries of the predictor palette are reordered or modified.
[0021] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the use of the predictor palette is indicated by maintaining a counter that tracks the number of times the predictor palette is used.
[0022] In another example aspect, the method described above can be implemented by a video decoder device including a processor.
[0023] In another example aspect, the method described above can be implemented by a video encoder device including a processor.
[0024] Further, in a representative aspect, a device in a video system is disclosed, the device including a processor and a non-transitory memory having instructions thereon. When the instructions are executed by the processor, the processor is caused to implement any one or more of the disclosed methods.
[0025] Further, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product including program code for performing any one or more of the disclosed methods.
[0026] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims. Description of the Drawings
[0027] Figure 1 An example of a block encoded and decoded in palette mode is shown.
[0028] Figure 2 An example of notifying palette entries using predictor palette signaling is shown.
[0029] Figure 3 Examples of horizontal traversal scan and vertical traversal scan are shown.
[0030] Figure 4 An example of the encoding and decoding of palette indices is shown.
[0031] Figure 5 An example of a picture having an 18 by 12 luminance CTU and divided into 12 slices and 3 raster scan stripes is shown.
[0032] Figure 6 An example of a picture having an 18 by 12 luminance CTU and divided into 24 slices and 9 rectangular stripes is shown.
[0033] Figure 7 An example of a picture divided into 4 slices, 11 bricks, and 4 rectangular stripes is shown.
[0034] Figure 8 An example of a picture having 28 sub - pictures is shown.
[0035] Figure 9 It is a block diagram of an example hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.
[0036] Figure 10 It is a block diagram of an example video processing system that can implement the disclosed technology.
[0037] Figure 11 A flowchart of an example method for video encoding and decoding is shown.
[0038] Figure 12 It is a flowchart representation of a method for video processing according to the present technology.
[0039] Figure 13 It is a flowchart representation of another method for video processing according to the present technology.
[0040] Figure 14 It is a flowchart representation of another method for video processing according to the present technology.
[0041] Figure 15It is a flowchart representation of another video processing method according to the present technology.
[0042] Figure 16 It is a flowchart representation of another video processing method according to the present technology.
[0043] Figure 17 It is a flowchart representation of another video processing method according to the present technology.
[0044] Figure 18 It is a flowchart representation of another video processing method according to the present technology.
[0045] Figure 19 It is a flowchart representation of another video processing method according to the present technology.
[0046] Figure 20 It is a flowchart representation of another video processing method according to the present technology.
[0047] Figure 21 It is a flowchart representation of yet another video processing method according to the present technology. Detailed implementation manners
[0048] 1. Video encoding and decoding of HEVC / H.265
[0049] Video encoding and decoding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) as well as the H.265 / HEVC standard. Since H.262, video encoding and decoding standards have been based on a hybrid video encoding and decoding structure, which utilizes temporal prediction and transform encoding and decoding. To explore future video encoding and decoding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, a JVET was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, aiming to reduce the bit rate by 50% compared to HEVC.
[0050] 2. Palette mode
[0051] 2.1 Palette mode of HEVC Screen Content Coding Extension (HEVC-SCC)
[0052] 2.1.1. Concept of the palette mode
[0053] The basic idea behind the palette mode is that the pixels in a CU are represented by a small set of representative color values. This set is called the palette. Also, out-of-palette samples can be indicated by signaling the escape symbols of the subsequent (possibly quantized) component values. Such pixels are called escape pixels. The palette mode is shown in Figure 1 As shown in Figure 1 , for each pixel with three color components (luma and two chroma components), an index of the palette is established, and the block can be reconstructed based on the values found in the palette.
[0054] 2.1.2. Encoding and Decoding of Palette Entries
[0055] For the palette-encoded / decoded blocks, the following key aspects are introduced:
[0056] 1) Construct the current palette based on the predictor palette and the new entries (if any) signaled for the current palette.
[0057] 2) Divide the current sample / pixel into two categories: one category (the first category) includes the samples / pixels in the current palette, and the other category (the second category) includes the samples / pixels outside the current palette.
[0058] a. For the samples / pixels in the second category, quantization (at the encoder) is applied to the samples / pixels and the quantization values are signaled; and dequantization (at the decoder) is applied.
[0059] 2.1.2.1. Predictor Palette
[0060] For the encoding and decoding of palette entries, a predictor palette is maintained, which is updated after decoding the palette-encoded / decoded blocks.
[0061] 2.1.2.1.1. Initialization of the Predictor Palette
[0062] Initialize the predictor palette at the start of each slice and each tile.
[0063] Signal the palette and the maximum size of the predictor palette in the SPS. In HEVC-SCC, the palette_predictor_initializer_present_flag is introduced in the PPS. When this flag is 1, the entries for initializing the predictor palette are signaled in the bitstream.
[0064] Depending on the value of palette_predictor_initializer_present_flag, the size of the predictor palette is reset to 0 or initialized using the predictor palette initializer entry signaled in the PPS. In HEVC-SCC, a predictor palette initializer of size 0 is enabled to allow explicit disabling of the predictor palette initialization at the PPS level.
[0065] The corresponding syntax, semantics, and decoding processes are defined as follows. Newly added text is shown in bold and underlined italics. Any deleted text is marked with [[ ]].
[0066] 7.3.2.2.3 Sequence parameter set screen content coding extension syntax
[0067]
[0068] palette_mode_enabled_flag being equal to 1 specifies that the decoding process for the palette mode can be used for intra blocks. palette_mode_enabled_flag being equal to 0 specifies that the decoding process for the palette mode shall not be applied. When not present, the value of palette_mode_enabled_flag is inferred to be equal to 0.
[0069] palette_max_size specifies the maximum allowed size of the palette. When not present, the value of palette_max_size is inferred to be 0.
[0070] delta_palette_max_predictor_size specifies the difference between the maximum allowed size of the palette predictor and the maximum allowed size of the palette. When not present, the value of delta_palette_max_predictor_size is inferred to be 0. The derivation of the variable PaletteMaxPredictorSize is as follows:
[0071] PaletteMaxPredictorSize = palette_max_size + delta_palette_max_predictor_size (2-1)
[0072] The bitstream conformance requirement is that when palette_max_size is equal to 0, the value of delta_palette_max_predictor_size shall be equal to 0.
[0073] The sps_palette_predictor_initializer_present_flag being equal to 1 specifies that the sequence palette predictor is initialized using the sps_palette_predictor_initializers specified in this clause. The sps_palette_predictor_initializer_flag being equal to 0 specifies that the entries in the sequence palette predictor are initialized to 0. When not present, the value of sps_palette_predictor_initializer_flag is inferred to be equal to 0.
[0074] The requirement for bitstream conformance is that when palette_max_size is equal to 0, the value of sps_palette_predictor_initializer_present_flag shall be equal to 0.
[0075] sps_num_palette_predictor_initializer_minus1 plus 1 specifies the number of entries in the sequence palette predictor initializer.
[0076] The requirement for bitstream conformance is that the value of sps_num_palette_predictor_initializer_minus1 plus 1 shall be less than or equal to PaletteMaxPredictorSize.
[0077] sps_palette_predictor_initializers[comp][i] specifies the value of the comp-th component of the i-th palette entry used to initialize the array PredictorPaletteEntries in the SPS. For values of i in the range 0 to sps_num_palette_predictor_initializer_minus1 (inclusive), the value of sps_palette_predictor_initializers[0][i] shall be in the range 0 to (1 << BitDepth Y ) – 1 (inclusive), and the values of sps_palette_predictor_initializers[1][i] and sps_palette_predictor_initializers[2][i] shall be in the range 0 to (1 << BitDepth C ) – 1 (inclusive).
[0078] 7.3.2.3.3 Picture Parameter Set Screen Content Coding Extension Syntax
[0079]
[0080] The pps_palette_predictor_initializer_present_flag being equal to 1 specifies that the palette predictor initializer derived from the PPS - specified palette predictor initializer is used for the pictures of the reference PPS. The pps_palette_predictor_initializer_flag being equal to 0 specifies that the palette predictor initializer for the pictures of the reference PPS is inferred to be equal to the palette predictor initializer specified by the active SPS. When not present, the value of pps_palette_predictor_initializer_present_flag is inferred to be equal to 0. The bit - stream conformance requirement is that when palette_max_size is equal to 0 or palette_mode_enabled_flag is equal to 0, the value of pps_palette_predictor_initializer_present_flag shall be equal to 0.
[0081] pps_num_palette_predictor_initializer specifies the number of entries in the picture palette predictor initializer.
[0082] The bit - stream conformance requirement is that the value of pps_num_palette_predictor_initializer shall be less than or equal to PaletteMaxPredictorSize.
[0083] Initialize the palette predictor variables as follows:
[0084] – If the coding tree unit is the first coding tree unit in the slice, the following applies:
[0085] – Call the initialization process of the palette predictor variables as specified in Clause 9.3.2.3.
[0086] – Otherwise, if entropy_coding_sync_enabled_flag is equal to 1, and CtbAddrInRs % PicWidthInCtbsY is equal to 0 or TileId[CtbAddrInTs] is not equal to TileId[CtbAddrRsToTs[CtbAddrInRs - 1]], the following applies:
[0087] – Derive the spatial neighboring block T using the position (x0, y0) of the top - left luma sample of the current coding tree block as followsFigure 2 ) the position (xNbT, yNbT) of the top - left luminance sample is as follows:
[0088] (xNbT,yNbT)=(x0 + CtbSizeY,y0 - CtbSizeY) (9 - 3)
[0089] - Invoke the availability derivation process of the blocks in the z - scan order specified in Clause 6.4.1, where the position (xCurr, yCurr) set to be equal to (x0, y0) and the neighboring position (xNbY, yNbY) set to be equal to (xNbT, yNbT) are used as inputs, and the output is assigned to availableFlagT.
[0090] - The synchronization process calls of context variables, Rice parameter initialization status, and palette predictor variables are as follows:
[0091] - If availableFlagT is equal to 1, then invoke the synchronization process of context variables, Rice parameter initialization status, and palette predictor variables specified in Clause 9.3.2.5, with TableStateIdxWpp, TableMpsValWpp, TableStatCoeffWpp, PredictorPaletteSizeWpp, and TablePredictorPaletteEntriesWpp as inputs.
[0092] - Otherwise, the following applies:
[0093] - Invoke the initialization process of the palette predictor variables specified in Clause 9.3.2.3.
[0094] - Otherwise, if CtbAddrInRs is equal to slice_segment_address and dependent_slice_segment_flag is equal to 1, then invoke the synchronization process of context variables and Rice parameter initialization status specified in Clause 9.3.2.5, with TableStateIdxDs, TableMpsValDs, TableStatCoeffDs, PredictorPaletteSizeDs, and TablePredictorPaletteEntriesDs as inputs.
[0095] - Otherwise, the following applies:
[0096] - Invoke the initialization process of the palette predictor variables specified in Clause 9.3.2.3.
[0097] 9.3.2.3 Initialization Process of Palette Predictor Entries
[0098] The output of this process is the initialized palette predictor variables PredictorPaletteSize and PredictorPaletteEntries.
[0099] The variable numComps is derived as follows:
[0100]
[0101] 2.1.2.1.2. Use of Predictor Palette
[0102] For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. This is shown in Figure 2 . The reuse flag is sent using run - length coding of zeros. After that, the number of new palette entries is signaled using an exponential Golomb (EG) code of order 0 (e.g., EG - 0). Finally, the component values of the new palette entries are signaled.
[0103] 2.1.2.2. Update of Predictor Palette
[0104] The update of the predictor palette is performed using the following steps:
[0105] (1) Before decoding the current block, there is a predictor palette, denoted as PltPred0
[0106] (2) The current palette table is constructed by first inserting the items in PltPred0 and then inserting the new entries of the current palette.
[0107] (3) Construct PltPred1:
[0108] a. First add the items in the current palette table (which may include items in PltPred0)
[0109] b. If not full, add the unreferenced items in PltPred0 in ascending order of entry index.
[0110] 2.1.3. Coding and Decoding of Palette Index
[0111] As Figure 3 shown, horizontal and vertical traversal scans are used to code and decode the palette index. The scan order is explicitly signaled in the bitstream using palette_transpose_flag. For the remainder of the sub - clause, it is assumed that the scan is horizontal.
[0112] Decode and encode palette indices using two palette sample modes: "COPY_LEFT" and "COPY_ABOVE". In the "COPY_LEFT" mode, the palette index is assigned to the decoded index. In the "COPY_ABOVE" mode, the palette index of the sample in the above row is copied. For both the "COPY_LEFT" and "COPY_ABOVE" modes, signal the run value, which specifies the number of subsequent samples encoded and decoded using the same mode.
[0113] In the palette mode, the index value of an escape sample is the number of palette entries. Also, when the escape symbol is part of a run in the "COPY_LEFT" or "COPY_ABOVE" mode, signal the escape component value for each escape symbol. The encoding and decoding of the palette index are shown in Figure 4 below.
[0114] The syntax order is completed as follows. First, signal the number of index values for the CU. Then signal the actual index values for the entire CU using truncated binary encoding. The number of indices and the index values are both encoded and decoded in bypass mode. This groups the bypass binary bits related to the indices together. Then, signal the palette sample mode (if necessary) and the run in an interleaved manner. Finally, the component escape values for the escape samples corresponding to the entire CU are grouped together and encoded and decoded in bypass mode. The binarization of the escape samples is EG encoding with three orders, such as EG-3.
[0115] Signal the additional syntax element last_run_type_flag after signaling the index values. This syntax element, combined with the number of indices, eliminates the need to signal the run value corresponding to the last run in the block.
[0116] In HEVC-SCC, the palette mode also applies to 4:2:2, 4:2:0, and monochrome chroma formats. For all chroma formats, the signaling of the palette entries and palette indices is almost the same. If it is a non-monochrome format, each palette entry consists of 3 components. For the monochrome format, each palette entry consists of a single component. For the subsampled chroma direction, the chroma samples are associated with the luminance sample indices divisible by 2. After reconstructing the palette index of the CU, if the sample has only a single component associated with it, only the first component of the palette entry is used. The only difference in signaling lies in the escape component value. For each escape sample, the number of escape component values signaled may be different, depending on the number of components associated with the sample.
[0117] In addition, there is an index adjustment process in palette index encoding and decoding. When signaling a palette index, the left adjacent index or the upper adjacent index should be different from the current index. Therefore, by removing one possibility, the range of the current palette index can be reduced by 1. After that, the index is signaled using truncated binary (TB) binarization.
[0118] The text related to this part is as follows, where CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the predicted index.
[0119] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is the index of the array represented by CurrentPaletteEntries. The array indices xC, yC specify the position (xC, yC) of the sample relative to the top - left luma sample of the picture. The value of PaletteIndexMap[xC][yC] should be in the range of 0 to MaxPaletteIndex (inclusive).
[0120] The derivation of the variable adjustedRefPaletteIndex is as follows:
[0121]
[0122]
[0123] When CopyAboveIndicesFlag[xC][yC] is equal to 0, the derivation of the variable CurrPaletteIndex is as follows:
[0124] if(CurrPaletteIndex >= adjustedRefPaletteIndex)
[0125] CurrPaletteIndex++
[0126] 2.1.3.1. Decoding process of palette encoding - decoding block
[0127] 1) Read the prediction information to mark which entries in the predictor palette will be reused;
[0128] (palette_predictor_run)
[0129] 2) Read the new palette entries of the current block
[0130] a.num_signaled_palette_entries
[0131] b.new_palette_entries
[0132] 3) Construct CurrentPaletteEntries based on a) and b)
[0133] 4) Read the escape symbol presence flag: palette_escape_val_present_flag to derive MaxPaletteIndex
[0134] 5) Encode and decode how many samples are not encoded in copy mode / run - length mode
[0135] a.num_palette_indices_minus1
[0136] b. For each sample not encoded in copy mode / run - length mode, encode and decode palette_idx_idc in the current plt table
[0137] 2.2. Palette Mode in VVC
[0138] 2.2.1. Palette in the Dual - Tree
[0139] In VVC, the dual - tree coding structure is used to code intra - slices, so the luminance component and the two chrominance components may have different palettes and palette indices. In addition, the two chrominance components share the same palette and palette index.
[0140] 2.2.2. Palette as a Separate Mode
[0141] In some embodiments, the prediction mode of the coding unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. The binarization of the prediction mode is changed accordingly.
[0142] When IBC is disabled, on I - slices, the first binary bit is used to indicate whether the current prediction mode is MODE_PLT. On P / B - slices, the first binary bit is used to indicate whether the current prediction mode is MODE_INTRA. If not, an additional binary bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTER.
[0143] When IBC is enabled, on an I slice, the first binary bit is used to indicate whether the current prediction mode is MODE_IBC. If not, the second binary bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. On a P / B slice, the first binary bit is used to indicate whether the current prediction mode is MODE_INTRA. If it is an intra mode, the second binary bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. If not, the second binary bit is used to indicate whether the current prediction mode is MODE_IBC or MODE_INTER.
[0144] An example syntax text is shown as follows.
[0145] Coding and decoding unit syntax
[0146]
[0147]
[0148] 2.3. Partitioning of pictures, sub - pictures, strips, slices, tiles and CTUs
[0149] Sub - picture: A rectangular region of one or more strips within a picture.
[0150] Strip: An integral number of tiles of a picture that is uniquely contained within a single NAL unit. A strip consists of a sequence of multiple complete slices or a continuous sequence of complete tiles of a slice.
[0151] Slice: A rectangular region of CTUs within a specific slice column and a specific slice row of a picture.
[0152] Tile: A rectangular region of CTU rows within a specific slice of a picture. A slice can be divided into multiple tiles, each tile consisting of one or more CTU rows within the slice. A slice that is not divided into multiple tiles is also called a tile. However, a tile that is a true subset of a slice is not called a slice.
[0153] Tile scan: A specific order sorting of CTU - partitioned pictures, where CTUs are sorted consecutively in the CTU raster scan of a tile, tiles within a slice are sorted consecutively in the raster scan of the tiles of the slice, and slices within a picture are sorted consecutively in the raster scan of the slices of the picture.
[0154] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular region of the picture.
[0155] A slice is divided into one or more tiles, each tile consisting of multiple CTU rows within the slice.
[0156] A slice that is not divided into multiple tiles is also referred to as a tile. However, a tile that is a proper subset of a slice is not called a slice.
[0157] A stripe contains multiple slices of a picture or multiple tiles of a slice.
[0158] A sub-picture contains one or more stripes that jointly cover a rectangular region of the picture.
[0159] Two stripe modes are supported, namely the raster scan stripe mode and the rectangular stripe mode. In the raster scan stripe mode, a stripe contains a sequence of slices in the slice raster scan of the picture. In the rectangular stripe mode, a stripe contains multiple tiles of the picture that jointly form a rectangular region of the picture. The tiles within a rectangular stripe are in the tile raster scan order of the stripe.
[0160] Figure 5 An example of the raster scan stripe segmentation of a picture is shown, where the picture is divided into 12 slices and 3 raster scan stripes.
[0161] Figure 6 An example of the rectangular stripe segmentation of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular stripes.
[0162] Figure 7 An example of a picture segmented into slices, tiles, and rectangular stripes is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the top-left slice contains 1 tile, the top-right slice contains 5 tiles, the bottom-left slice contains 2 tiles, and the bottom-right slice contains 3 tiles), and 4 rectangular stripes.
[0163] Figure 8 An example of the sub-picture segmentation of a picture is shown, where the picture is segmented into 28 sub-pictures of different dimensions.
[0164] When a picture is encoded and decoded using three separate color planes (separate_colour_plane_flag equals 1), a stripe contains only the CTUs of one color component, which is identified by the corresponding value of colour_plane_id, and each color component array of the picture consists of stripes with the same colour_plane_id value. The encoded and decoded stripes with different colour_plane_id values within the picture can be interleaved under the following constraint: for each value of colour_plane_id, the encoded and decoded stripe NAL units with that colour_plane_id value shall be in the order of increasing CTU addresses in the tile scan order of the first CTU of each encoded and decoded stripe NAL unit.
[0165] When separate_colour_plane_flag is equal to 0, each CTU of the picture is exactly contained in one strip. When separate_colour_plane_flag is equal to 1, each CTU of the color component is exactly contained in one strip (for example, the information of each CTU of the picture exists exactly in three strips, and these three strips have different colour_plane_id values).
[0166] 2.4. Wavefront with 1-CTU delay
[0167] In VVC, a one-CTU-delay wavefront parallel processing (WPP) is adopted instead of the two-CTU-delay in the HEVC design. The WPP processing enables multiple parallel processes with limited encoding and decoding losses, but the two-CTU-delay may hinder the parallel processing ability. Since the target resolution is becoming larger and the number of CPUs is increasing, it is asserted that improving the parallel processing ability by using the proposed one-CTU-delay is beneficial to reducing the encoding and decoding delay and will make full use of the CTU ability.
[0168] 3. Example problems in existing embodiments
[0169] DMVR and BIO do not involve the original signal during the refinement of the motion vector, which may lead to inaccurate motion information of the encoded and decoded blocks. In addition, DMVR and BIO sometimes adopt fractional motion vectors after motion refinement, while screen videos usually use integer motion vectors, which makes the current motion information more inaccurate and the encoding and decoding performance worse.
[0170] (1) The current palette is constructed based on the prediction of the previous encoded and decoded palette. The current palette is reinitialized only when entropy_coding_sync_enabled_flag is equal to 1, before decoding a new CTU row or a new slice. However, in practical applications, parallel encoders are preferred, where different CTU rows can be pre-encoded and decoded without referring to the information of other CTU rows.
[0171] (2) The way to process the predictor palette update process is fixed. That is, the entries inherited from the previous predictor palette and the new entries in the current palette are inserted in sequence. If the number of entries is still less than the size of the predictor palette, further entries not inherited from the previous predictor palette will be added. This design does not consider the importance of different entries in the current predictor palette and the previous predictor palette.
[0172] (3) The size of the predictor palette is fixed, and after decoding a block, it must be updated to fill all entries, which is sub-optimal because some of these entries may never be referenced.
[0173] (4) The size of the current palette is fixed, regardless of the color components, and fewer chroma samples may be used, for example, compared to luminance.
[0174] 4. Example Techniques and Embodiments
[0175] The detailed embodiments described below should be considered as examples explaining the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way.
[0176] In addition to DMVR and BIO mentioned below, the methods described below may also be applicable to other decoder motion information derivation techniques.
[0177] Re - graded Predictor Palette
[0178] 1. It is proposed to reset or re - initialize the predictor palette (e.g., entries and / or the size of the predictor palette) before decoding the first block in a new video unit.
[0179] a. Alternatively, after decoding the last block in a video unit, the predictor palette (e.g., entries and / or the size of the predictor palette) can be reset or re - initialized.
[0180] b. In one example, the video unit is a sub - region / CTU / CTB / multiple CTUs / multiple CUs / CTU row / slice / brick / sub - picture / view, etc. of a CTU (e.g., VPDU).
[0181] i. Alternatively, in addition, even if the wavefront is disabled (e.g., entropy_coding_sync_enabled_flag equals 0), the above - mentioned method is called.
[0182] c. In one example, the video unit is a chroma CTU row.
[0183] i. Alternatively, in addition, the predictor palette can be reset or re - initialized before decoding the first chroma CTB in a new chroma CTU row.
[0184] ii. Alternatively, in addition, when applying the dual - tree and the current split - tree is a chroma codec tree, the above - mentioned method is called.
[0185] d. In one example, the size of the predictor palette (e.g., PredictorPaletteSize in the specification) is reset to 0.
[0186] e. In one example, the size of the predictor palette (e.g., PredictorPaletteSize in the specification) is reset to the number of entries in the sequence palette predictor initializer (e.g., sps_num_palette_predictor_initializer_minus1 plus 1) or the maximum number of entries allowed in the predictor palette (e.g., PaletteMaxPredictorSize).
[0187] f. Initialization of the predictor palette (e.g., PredictorPaletteEntries) before encoding / decoding a new sequence / picture can be used to initialize the predictor palette before encoding / decoding a new video unit.
[0188] g. In one example, when the entropy_coding_sync_enabled_flag is equal to 1, the predictor palette after encoding / decoding the upper CTB / CTU can be used to initialize the predictor palette before encoding / decoding the current CTB / CTU.
[0189] 2. It is proposed to prohibit updating the predictor palette after encoding / decoding a specific palette coding / decoding block.
[0190] a. In one example, whether to update the predictor palette can depend on the decoded information of the current block.
[0191] i. In one example, whether to update the predictor palette can depend on the block dimensions of the current block.
[0192] 1. In one example, if the width of the current block is not greater than a first threshold (denoted by T1) and the height of the current block is not greater than a second threshold (denoted by T2), the update process is disabled.
[0193] 2. In one example, if the product of the width of the current block and the height of the block is not greater than a first threshold (denoted by T1), the update process is disabled.
[0194] 3. In one example, if the width of the current block is not less than a first threshold (denoted by T1) and the height of the current block is not less than a second threshold (denoted by T2), the update process is disabled.
[0195] 4. In one example, if the product of the width of the current block and the height of the block is not less than a first threshold (denoted by T1), the update process is disabled.
[0196] 5. In the above examples, T1 / T2 can be predefined or signaled.
[0197] a) In one example, T1 / T2 can be set to 4, 16, or 1024.
[0198] b) In one example, T1 / T2 can depend on the color component.
[0199] 3. A shared predictor palette can be defined, where the same predictor palette can be used for all CUs / PUs under the shared region.
[0200] a. In one example, a shared region can be defined for an MxN region (e.g., 16×4 or 4×16 region) partitioned using TT.
[0201] b. In one example, a shared region can be defined for an MxN region (e.g., 8×4 or 4×8 region) partitioned using BT.
[0202] c. In one example, a shared region can be defined for an MxN region (e.g., 8×8 region) partitioned using QT.
[0203] d. Alternatively, in addition, a shared predictor palette can be constructed before encoding / decoding all blocks within the shared region.
[0204] e. In one example, the indication of a prediction entry in the predictor palette (e.g., palette_predictor_run) can be signaled together with the first palette-coded block within the shared region.
[0205] i. Alternatively, in addition, for the remaining coded blocks within the shared region, signaling of the indication of a prediction entry in the predictor palette (e.g., palette_predictor_run) can be skipped.
[0206] f. Alternatively, in addition, after decoding / encoding the blocks within the shared region, the update of the predictor palette can always be skipped.
[0207] 4. A counter can be maintained for each entry of the predictor palette to indicate the frequency of its use.
[0208] a. In one example, for each new entry added to the predictor palette, the counter can be set to a constant K.
[0209] i. In one example, K can be set to 0.
[0210] b. In one example, when an entry is marked for reuse in the coded / decoded palette block, the corresponding counter can be incremented by a constant N.
[0211] i. In one example, N can be set to 1.
[0212] 5. It is proposed to adaptively change the size of the predictor palette instead of using a predictor palette of a fixed size.
[0213] a. In one example, it can be changed between a video unit (block / CU / CTU / slice / tile / sub-picture) and another video unit.
[0214] b. In one example, the size of the predictor palette can be updated according to the size of the current palette.
[0215] i. In one example, the size of the predictor palette can be set to the size of the current palette after decoding / encoding the current block.
[0216] ii. In one example, the size of the predictor palette can be set to the size of the current palette minus or plus an integer value represented as K.
[0217] 1. In one example, K can be signaled / instantaneously derived.
[0218] c. In one example, the size of the predictor palette can depend on the block size. Let S be the predefined size of the predictor palette for the palette-coded block.
[0219] i. In one example, a palette-coded block with a size less than or equal to T can use a predictor palette with a size less than S.
[0220] 1. In one example, the first K entries (K <= S) in the palette predictor can be used.
[0221] 2. In one example, a subsampled version of the palette predictor can be used.
[0222] ii. In one example, a palette-coded block with a size greater than or equal to T can use a predictor palette with a size equal to S.
[0223] iii. In the above examples, K and / or T are integers and can be based on
[0224] 1. Video content (e.g., screen content or natural content)
[0225] 2. Messages signaled in DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header / largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit
[0226] 3. The position of CU / PU / TU / block / video coding unit
[0227] 4. Indication of color format (e.g., 4:2:0, 4:4:4, RGB or YUV)
[0228] 5. Coding / decoding tree structure (e.g., dual tree or single tree)
[0229] 6. Slice / tile type and / or picture type
[0230] 7. Color component
[0231] 8. Temporal layer ID
[0232] 9. Standard profile / level / tier
[0233] d. In one example, after encoding / decoding the color palette section, the predictor palette can be customized according to the entry counter.
[0234] i. In one example, entries with a counter less than the threshold T can be discarded.
[0235] ii. In one example, entries with the minimum counter value can be discarded until the size of the predictor palette is less than the threshold T.
[0236] e. Alternatively, in addition, after decoding / encoding the palette coding block, the predictor palette can be updated based only on the current palette.
[0237] i. Alternatively, in addition, after decoding / encoding the palette coding block, the predictor palette can be updated to the current palette.
[0238] 6. Entries of the current palette and / or the predictor palette before encoding / decoding the current block can be reordered / modified before being used to update the predictor palette.
[0239] a. In one example, reordering can be applied according to the decoded information / reconstruction of the current sample.
[0240] b. In one example, reordering can be applied according to the counter value of the entry.
[0241] c. Alternatively, in addition, the number of occurrences of samples / pixels (in the current palette and / or outside the current palette) can be counted.
[0242] i. Alternatively, in addition, samples / pixels with a larger counter (e.g., occurring more frequently) can be placed before another sample / pixel with a smaller counter.
[0243] 7. Information of escape samples can be used to update the predictor palette.
[0244] a. Alternatively, in addition, an update of the predictor palette using the escape information can be conditionally invoked.
[0245] i. In one example, when the predictor palette is not full after inserting the current palette, the escape sample / pixel information can be added to the predictor palette.
[0246] 8. Updating / initializing / resetting the predictor palette can depend on the color component.
[0247] a. In one example, the rule for determining whether to update the predictor palette can depend on the color component, such as luminance or chrominance.
[0248] 9. A set of multiple predictor palettes can be maintained and / or updated.
[0249] a. In one example, one predictor palette can have information for one or all color components.
[0250] b. In one example, one predictor palette can have information for two color components (e.g., Cb and Cr).
[0251] c. In one example, at least one global palette and at least one local palette can be maintained.
[0252] i. In one example, the predictor palette can be updated based on the global palette and the local palette.
[0253] d. In one example, a palette associated with the last K palette-coded blocks (in encoding / decoding order) can be maintained.
[0254] e. In one example, palettes for the luminance component and the chrominance component can be predicted based on different predictor palettes (e.g., different indices of a set of multiple predictor palettes).
[0255] f. Alternatively, in addition, bullet point 1 can be applied to the set of predictor palettes.
[0256] g. Alternatively, in addition, the index / multiple indices of the predictor palette in the set of predictor palettes can be signaled for a sub-region of the CU / PU / CTU / CTB / CTU or CTB.
[0257] Regarding Palette / Predictor Palette Sizes
[0258] 10. The size of the palette can change between video units and another video unit.
[0259] a. In one example, it can change between a video unit (block / CU / CTU / slice / tile / sub-picture) and another video unit.
[0260] b. In one example, it can depend on the decoding information of the current block and / or neighboring (adjacent or non-adjacent) blocks.
[0261] 11. The size of the palette and / or predictor palette can depend on the block dimension and / or quantization parameter.
[0262] 12. The size (or the number of entries therein) of the palette and / or predictor palette can be different for different color components.
[0263] a. In one example, an indication of the size of the palette and / or predictor palette for the luminance component and chrominance component can be signaled explicitly or implicitly.
[0264] b. In one example, an indication of the size of the palette and / or predictor palette for each color component can be signaled explicitly or implicitly.
[0265] c. In one example, whether to signal an indication of multiple sizes can depend on the use of the dual tree and / or the slice / picture type.
[0266] Signaling of the Palette
[0267] 13. The compliant bitstream shall satisfy that the number of entries signaled directly for the current block (e.g., num_signaled_palette_entries) shall be within the range [0, palette_max_size - NumPredictedPaletteEntries], which is a closed range including 0 and palette_max_size - NumPredictedPaletteEntries.
[0268] a. How to binarize num_signaled_palette_entries can depend on the allowed range.
[0269] i. Truncated binarization coding can be used instead of EG-0 th 。
[0270] b. How to binarize num_signaled_palette_entries can depend on the decoding information (e.g., block dimension).
[0271] Regarding the Wavefront with 1 - CTU
[0272] 14. It is proposed to re-initialize the predictor palette (e.g., entries and / or size) when parsing of the CTU syntax ends (e.g., in Clause 7.3.8.2 of VVC), entropy_coding_sync_enabled_flag is equal to 1, and the current CTB is the first in a new CTU row or the current CTB is not in the same tile as its previous CTB.
[0273] a. Alternatively, in addition, maintain PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp to record the updated size and entries of the predictor palette after encoding / decoding of the above CTU is completed.
[0274] i. Alternatively, in addition, PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp can be used for encoding / decoding the current block in the current CTU.
[0275] b. In one example, when parsing of the CTU syntax in Clause 7.3.8.2 ends, entropy_coding_sync_enabled_flag is equal to 1, and CtbAddrInRs % PicWidthInCtbsY is equal to 0 or BrickId[CtbAddrInBs] is not equal to BrickId[CtbAddrRsToBs[CtbAddrInRs - 1]], call the storage procedure of the context variables specified in Clause 9.3.2.3, with TableStateIdx0Wpp, TableStateIdx1Wpp, TableMpsValWpp, PredictorPaletteSizeWpp, and PredictorPaletteEntriesWpp (when palette_mode_enabled_flag is equal to 1) as outputs.
[0276] Overview
[0277] 15. Whether and / or how to apply the above method may be based on the following:
[0278] a. Video content (e.g., screen content or natural content)
[0279] b. Messages signaled in DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header / largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit
[0280] c. Location of CU / PU / TU / block / video coding unit
[0281] d. Decoding information of the current block and / or its neighboring blocks
[0282] i. Block dimension / block shape of the current block and / or its neighboring blocks
[0283] e. Indication of color format (e.g., 4:2:0, 4:4:4, RGB or YUV)
[0284] f. Coding tree structure (e.g., dual tree or single tree)
[0285] g. Strip / slice group type and / or picture type
[0286] h. Color component (e.g., can be applied to only the luminance component and / or the chrominance component)
[0287] i. Temporal layer ID
[0288] j. Profile / level / hierarchy of the standard
[0289] 5. Additional embodiments
[0290] In the following embodiments, the newly added text is shown in bold and underlined italic text. Any deleted text is marked by [[ ]].
[0291] 5.1. Embodiment #1
[0292] 9.3.1 Overview
[0293] This procedure is called when parsing a syntax element using the descriptor ae(v) in clauses 7.3.8.1 to 7.3.8.12.
[0294] The input to this procedure is a request for the value of the syntax element and the values of previously parsed syntax elements.
[0295] The output of this procedure is the value of the syntax element.
[0296] When starting to parse one or more of the following, the initialization procedure specified in clause 9.3.2 is called:
[0297] 1. The strip segmentation data syntax specified in clause 7.3.8.1,
[0298] 2. The CTU syntax specified in clause 7.3.8.2 and the CTU is Brick [[slice]] the first CTU in,
[0299] 3. The CTU syntax specified in Clause 7.3.8.2, where [[entropy_coding_sync_enabled_flag is equal to 1 and]] the associated luma CTB is Brick [[slice]] the first luma CTB in the CTU row of
[0300] The parsing process of the syntax elements is as follows:
[0301] When cabac_bypass_alignment_enabled_flag is equal to 1, for a request for the value of a syntax element targeting the syntax element coeff_abs_level_remaining[] or coeff_sign_flag[] and escapeDataPresent is equal to 1, call the alignment process before calibration bypass decoding specified in Clause 9.3.4.3.6.
[0302] For each requested value of the syntax element, perform the derived binarization as specified in Clause 9.3.3.
[0303] The binarization of the syntax element and the determined binary bit sequence of parsing determine the decoding process, as described in Clause 9.3.4.
[0304] When processing a request for the value of a syntax element for pcm_flag and the decoded value of pcm_flag is equal to 1, initialize the decoding engine after decoding any pcm_alignment_zero_bit and all pcm_sample_luma and pcm_sample_chroma data specified in Clause 9.3.2.6. The storage process of the context variables is applied as follows:
[0305] – When the parsing of the CTU syntax in Clause 7.3.8.2 is completed, entropy_coding_sync_enabled_flag is equal to 1, and CtbAddrInRs % PicWidthInCtbsY is equal to 1, or both CtbAddrInRs are greater than 1 and TileId[CtbAddrInTs] is not equal to TileId[CtbAddrRsToTs[CtbAddrInRs - 2]], call the storage procedure for context variables, Rice parameter initialization status, and palette predictor variables specified in Clause 9.3.2.4, where TableStateIdxWpp, TableMpsValWpp, TableStatCoeffWpp (when persistent_rice_adaptation_enabled_flag is equal to 1), PredictorPaletteSizeWpp, and PredictorPaletteEntriesWpp (when palette_mode_enabled_flag is equal to 1) are used as outputs.
[0306] – When the parsing of the overall slice segment data syntax in Clause 7.3.8.1 is completed, dependent_slice_segments_enabled_flag is equal to 1 and end_of_slice_segment_flag is equal to 1, call the storage procedure for context variables, Rice parameter initialization status, and palette predictor variables specified in Clause 9.3.2.4, where TableStateIdxDs, TableMpsValDs, TableStatCoeffDs (when persistent_rice_adaptation_enabled_flag is equal to 1), PredictorPaletteSizeDs, and PredictorPaletteEntriesDs (when palette_mode_enabled_flag is equal to 1) are used as outputs.
[0307] 5.2. Example #2
[0308] 9.3 CABAC Parsing Process of Slice Data
[0309] 9.3.1 Overview
[0310] The input to this process is a request for the value of a syntax element and the value of previously parsed syntax elements.
[0311] The output of this process is the value of the syntax element.
[0312] When starting to parse the CTU syntax specified in Clause 7.3.8.2 and when one or more of the following conditions are true, call the initialization process specified in Clause 9.3.2,
[0313] – The CTU is the first CTU in the tile.
[0314] – The value of entropy_coding_sync_enabled_flag is equal to 1, and the CTU is the first CTU in the CTU row of the tile.
[0315] The parsing process of syntax elements is as follows:
[0316] For each requested value of a syntax element, obtain the binarization as specified in Subclause 9.3.3.
[0317] The binarization of the syntax element and the sequence of parsed bits determine the decoding process, as described in Subclause 9.3.4.
[0318] The storage process of context variables is applied as follows:
[0319] – When finishing the parsing of the CTU syntax in Clause 7.3.8.2, if entropy_coding_sync_enabled_flag is equal to 1 and CtbAddrInRs % PicWidthInCtbsY is equal to 0 or BrickId[CtbAddrInBs] is not equal to BrickId[CtbAddrRsToBs[CtbAddrInRs - 1]], call the storage process of context variables specified in Clause 9.3.2.3, where TableStateIdx0Wpp, TableStateIdx1Wpp, and TableMpsValWpp And PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp (when palette_mode_enabled_ flag equals 1) are used as the output.
[0320] 9.3.2 Initialization Process
[0321] 9.3.2.1 General
[0322] – The output of this process is the initialized CABAC internal variables.
[0323] – The context variables of the arithmetic decoding engine are initialized as follows:
[0324] – If the CTU is the first CTU in the tile, call the initialization process of context variables as specified in Clause 9.3.2.2, and the variables PredictorPaletteSize[0 / 1 / 2] are initialized to 0.
[0325] – Otherwise, if entropy_coding_sync_enabled_flag equals 1, and CtbAddrInRs % PicWidthInCtbsY equals 0 or BrickId[CtbAddrInBs] is not equal to BrickId[CtbAddrRsToBs[CtbAddrInRs-1]], the following applies:
[0326] – Derive the position (xNbT, yNbT) of the top-left luma sample of the spatial neighboring block T( Figure 9 - 2 ) using the position (x0, y0) of the top-left luma sample of the current CTB, as follows:
[0327] (xNbT,yNbT) = (x0,y0-CtbSizeY) (9-3)
[0328] – Invoke the derivation process of neighboring block availability specified in Clause 6.4.4, where the position (xCurr, yCurr) is set to be equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to be equal to (xNbT, yNbT), checkPredModeY is set to be equal to FALSE, cIdx is set to be equal to 0 as input, and the output is assigned to availableFlagT.
[0329] – The synchronization process of context variables is invoked as follows:
[0330] – If availableFlagT equals 1, invoke the synchronization process of context variables specified in Clause 9.3.2.4, where TableStateIdx0Wpp, TableStateIdx1Wpp, TableMpsValWpp are used as input, and the variable PredictorPaletteSize is initialized to 0.
[0331] – Otherwise, invoke the initialization process of context variables specified in Clause 9.3.2.2, and the variable PredictorPaletteSize is initialized to 0.
[0332] – Otherwise, invoke the initialization process of context variables specified in Clause 9.3.2.2, and the variable PredictorPaletteSize is initialized to 0.
[0333] – Initialize the decoding engine registers ivlCurrRange and ivlOffset with 16-bit register precision by invoking the initialization process of the arithmetic decoding engine specified in Sub-clause 9.3.2.5.
[0334] 9.3.2.3 Storage Procedure for Context Variables
[0335] The inputs to this procedure include:
[0336] – CABAC context variables indexed by ctxTable and ctxIdx.
[0337] The outputs to this procedure include:
[0338] – Variables tableStateSync0, tableStateSync1, and tableMPSSync that contain the values of variables pStateIdx0, pStateIdx1, and valMps used in the initialization of context variables, where these context variables are assigned to all syntax elements in clauses 7.3.8.1 to 7.3.8.11, except end_of_brick_one_bit and end_of_subset_one_bit.
[0339] – When palette_mode_enabled_flag equals 1, PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp
[0340] For each context variable, the corresponding entries pStateIdx0, pStateIdx1, and valMps in tables tableStateSync0, tableStateSync1, and tableMPSSync are initialized to the corresponding pStateIdx0, pStateIdx1, and valMps.
[0341] PredictorPaletteSizeWpp is set to 0, and PredictorPaletteEntriesWpp is set to be empty.
[0342] Alternatively, the following may apply:
[0343] PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp are respectively set to the corresponding PredictorPaletteSize and the entries of the predictor palette.
[0344] 5.3. Example #3
[0345]
[0346] Alternatively, the PredictorPaletteSize in the above table can be set to another integer value, such as a fixed value or the predictor palette size.
[0347] 6. Example Implementations of the Disclosed Technology
[0348] Figure 9is a block diagram of a video processing device 900. The device 900 can be used to implement one or more of the methods described herein. The device 900 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 900 can include one or more processors 902, one or more memories 904, and video processing hardware 906. The processor 902 can be configured to implement one or more of the methods described in this document. The memory 904 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 906 can be used to implement some of the techniques described in this document in hardware circuitry and can be partially or fully part of the processor 902 (e.g., a graphics processing unit core GPU or other signal processing circuitry).
[0349] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of a current video block can correspond to bits that are co-located within the bitstream or scattered at different locations as defined by the syntax. For example, a macroblock can be encoded based on the transformed and decoded error residual values and can also be decoded using bits in the header and other fields in the bitstream.
[0350] It will be understood that the disclosed methods and techniques will benefit video encoder and / or decoder embodiments incorporated into a video processing device, such as a smartphone, a laptop computer, a desktop computer, and similar devices, by allowing the use of the techniques disclosed in this document.
[0351] Figure 10 is a block diagram of an example video processing system 1000 that can implement the various techniques disclosed herein. Various embodiments can include some or all of the components of the system 1000. The system 1000 can include an input 1002 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 - or 10 - bit multi - component pixel values) or in a compressed or encoded format. The input 1002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi - Fi or cellular interfaces.
[0352] System 1000 may include an encoding / decoding component 1004 that may implement various encoding / decoding or coding methods described in this document. The encoding / decoding component 1004 may reduce the average bit rate of a video from the input 1002 to the output of the encoding / decoding component 1004 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding / decoding component 1004 may be stored or transmitted via a connected communication, represented by component 1006. The stored or transmitted bitstream (or encoded / decoded) representation of the video received at input 1002 may be used by component 1008 to generate pixel values or a displayable video that is sent to the display interface 1010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "encoding / decoding" operations or tools, it will be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding / decoding results will be performed by the decoder.
[0353] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), or a Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, an IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as a mobile phone, a laptop computer, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0354] Figure 11 is a flowchart of an example method 1100 of video processing. At 1110, method 1100 includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further wherein, the predictor palette is selectively reset before the conversion between the video block and the bitstream representation of the video block.
[0355] Some embodiments may be described using the following clause-based format.
[0356] 1. A method of video processing, comprising:
[0357] performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further wherein, the predictor palette is selectively reset before the conversion between the video block and the bitstream representation of the video block.
[0358] 2. The method according to Clause 1, wherein the video unit comprises one of the following: one or more codec tree units, one or more codec tree blocks, a sub-region of a codec tree unit or a codec tree block, or a view of a codec tree block row / slice / brick / sub-picture / codec tree unit.
[0359] 3. The method according to any one of Clauses 1-2, wherein the delayed wavefront parallel processing is disabled during the conversion.
[0360] 4. The method according to Clause 3, wherein the entropy_coding_sync_enabled_flag is set to be equal to 0.
[0361] 5. The method according to Clause 1, wherein the video unit is a row of chroma codec tree units.
[0362] 6. The method according to Clause 5, wherein the predictor palette is reset before decoding the first chroma codec tree block (CTB) in the new chroma CTU row.
[0363] 7. The method according to Clause 5, wherein when applying a dual codec tree and the current split of the dual codec tree is a chroma codec tree unit, the predictor palette is reset.
[0364] 8. The method according to Clause 1, wherein the size of the predictor palette is reset to zero.
[0365] 9. The method according to Clause 1, wherein the size of the predictor palette is reset to the number of entries in the sequence palette predictor initializer or the maximum number of allowed entries.
[0366] 10. The method according to Clause 9, wherein the sequence palette predictor initializer is used to initialize the palette predictor before applying it to the video unit.
[0367] 11. The method according to Clause 1, wherein when the entropy_coding_sync_enabled_flag is set to 1, the palette predictor applied to the previous video block is re-initialized before applying it to the video unit.
[0368] 12. The method according to Clause 1, wherein updating the predictor palette is prohibited based on the codec information associated with the video unit.
[0369] 13. The method according to Clause 12, wherein the codec information includes the dimensions of the video unit.
[0370] 14. The method according to clause 13, wherein updating of the predictor palette is prohibited based on the dimensions of a video unit reaching one or more threshold conditions.
[0371] 15. The method according to clause 14, wherein the one or more threshold conditions are predefined.
[0372] 16. The method according to clause 14, wherein the one or more threshold conditions are signaled explicitly or implicitly in the coded representation of the video unit.
[0373] 17. A method for video processing, comprising:
[0374] Performing a conversion between a video block in a video unit and a coded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict current palette information of the video block, and further wherein when multiple coded units of the video unit have a common shared area, the predictor palette is a shared predictor palette.
[0375] 18. The method according to clause 17, wherein the shared area is associated with any one of: TT partition, BT partition, QT partition.
[0376] 19. The method according to clause 17, wherein the shared predictor palette is constructed before being applied to multiple coded units.
[0377] 20. The method according to clause 17, wherein an indication of the use of the shared predictor palette is signaled explicitly or implicitly in the coded representation associated with a first palette coded unit of the shared area.
[0378] 21. The method according to clause 17, further comprising:
[0379] Skipping updating of the shared predictor palette after the coded unit applied to multiple coded units.
[0380] 22. A method for video processing, comprising:
[0381] Performing a conversion between a video block in a video unit and a coded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict current palette information of the video block, and further wherein the size of the predictor palette is adaptively changed according to one or more conditions.
[0382] 23. The method according to clause 22, wherein one or more conditions are associated with at least the following: the size of the previous palette information, the dimension of the video unit, the content of the video unit, the color format of the video unit, the color components of the video unit, the codec tree structure of the video block, the relative position of the video block in the codec representation, the temporal layer ID of the video block, the slice / tile type and / or picture type of the video block, or the profile / level / layer of the video block.
[0383] 24. A method for video processing, comprising:
[0384] Performing a conversion between a video block in a video unit and a codec representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, wherein the predictor palette is updated based on the size or number of entries in the predictor palette.
[0385] 25. The method according to clause 24, wherein the size of the predictor palette is updated from a previous video block to a current video block.
[0386] 26. The method according to clause 24, wherein the size of the predictor palette is signaled implicitly or explicitly in the codec representation.
[0387] 27. The method according to clause 24, wherein the size of the predictor palette depends on one or more of the following: the dimension of the video block, the quantization parameter of the video block, or one or more color components of the video block.
[0388] 28. A method for video processing, comprising:
[0389] Performing a conversion between a video block in a video unit and a codec representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, wherein the entries of the predictor palette are reordered or modified.
[0390] 29. The method according to clause 28, wherein the entries of the predictor palette are reordered or modified when the entropy_coding_sync_enabled_flag is equal to 1.
[0391] 30. The method according to clause 28, wherein the entries of the predictor palette are reordered or modified when the end of the codec tree unit syntax is encountered.
[0392] 31. The method according to clause 28, wherein the entries of the predictor palette are reordered or modified when the current CTB is the first CTB in a new CTU row or when the current CTB and the previous CTB are not in the same tile.
[0393] 32. A method for video processing, comprising:
[0394] Performing a conversion between a video block in a video unit and a coded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, wherein the use of the predictor palette is indicated by maintaining a counter that tracks the number of times the predictor palette is used.
[0395] 33. The method according to any of the preceding clauses, wherein enabling or disabling the predictor palette is associated with at least one of the following: the size of the previous palette information, the dimensions of the video block, the content of the video block, the color format of the video block, the color components of the video block, the codec tree structure of the video block, the relative position of the video block in the coded representation, the temporal layer ID of the video block, the slice / group-of-slices type and / or picture type of the video block, or the profile / level / tier of the video block.
[0396] 34. The method according to any of the preceding clauses, wherein more than one predictor palette is used during the conversion.
[0397] 35. A video decoding device, comprising a processor configured to implement the method recited in one or more of clauses 1 to 34.
[0398] 36. A video encoding device, comprising a processor configured to implement the method recited in one or more of clauses 1 to 34.
[0399] 37. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method recited in any one of clauses 1 to 34.
[0400] 38. A method, device, or system described in this document.
[0401] Figure 12 It is a flowchart representation of a method 1200 for video processing according to the present technology. At operation 1210, method 1200 includes performing a conversion between a current block of video and a coded representation of the video using a palette mode, in which a palette of representative sample values is used to code the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and based on the characteristics of the current block, updating of the predictor palette is prohibited according to a rule after the conversion of the current block.
[0402] In some embodiments, the characteristics of the current block include codec information associated with the current block. In some embodiments, the characteristics of the current block include the dimensions of the current block. In some embodiments, the rule stipulates that updating the predictor palette is prohibited when the width of the current block is less than or equal to a first threshold and the height of the current block is less than or equal to a second threshold. In some embodiments, the rule stipulates that updating the predictor palette is prohibited when the height of the current block is less than or equal to the first threshold. In some embodiments, the rule stipulates that updating the predictor palette is prohibited when the width of the current block is greater than or equal to the first threshold and the height of the current block is greater than or equal to the second threshold. In some embodiments, the rule stipulates that updating the predictor palette is prohibited when the height of the current block is greater than or equal to the first threshold.
[0403] In some embodiments, the first threshold or the second threshold is predefined or signaled in the codec representation. In some embodiments, the first threshold is 4, 16, or 1024. In some embodiments, the second threshold is 4, 16, or 1024. In some embodiments, the first threshold or the second threshold is based on the color components of the current block.
[0404] Figure 13 is a flowchart representation of a video processing method 1300 according to the present technology. At operation 1310, method 1300 includes performing a conversion between a current block of a video and a codec representation of the video using a palette mode, in which a palette of representative sample values is used to codec the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and it is determined whether to change the predictor palette according to the color components of the current block.
[0405] In some embodiments, changing the predictor palette includes updating, initializing, or resetting the predictor palette. In some embodiments, the color components include luminance or chrominance components. In some embodiments, the predictor palette includes information corresponding to the color components of the current block. In some embodiments, the predictor palette includes information corresponding to all color components of the current block. In some embodiments, the predictor palette includes information corresponding to two chrominance components of the current block.
[0406] Figure 14 is a flowchart representation of a video processing method 1400 according to the present technology. At operation 1410, method 1400 includes performing a conversion between a current block in a video unit of a video and a codec representation of the video using a palette mode, in which a palette of representative sample values is used to codec the current block. During the conversion, multiple predictor palettes are used to predict the palette of representative sample values.
[0407] In some embodiments, the predictor palette for the current block is updated based at least on a global palette and a local palette. In some embodiments, multiple predictor palettes are associated with K blocks in a video unit that has been coded or decoded using a palette mode. In some embodiments, palettes for different color components are determined based on different predictor palettes of the multiple predictor palettes. In some embodiments, the multiple predictor palettes are reset or reinitialized before a first block in a video unit is transformed or after a last block in a previously transformed video unit is transformed. In some embodiments, an index of the predictor palette of the multiple predictor palettes is signaled in a coding unit, a prediction unit, a coding tree unit, a coding tree block, a sub-region of a coding tree unit, or a sub-region of a coding tree block in a coded representation.
[0408] Figure 15 is a flowchart representation of a method 1500 of video processing according to the present technology. At operation 1510, method 1500 includes performing a transformation between a current block in a video unit of a video and a coded representation of the video using a palette mode, in which a palette of representative sample values is used to code or decode the current block. During the transformation, a predictor palette is used to predict the palette of representative sample values. According to a rule, the predictor palette is reset or reinitialized before a transformation of a first block in a video unit or after a transformation of a last video block in a previous video unit.
[0409] In some embodiments, a video unit includes a sub-region of a coding tree unit, a virtual pipeline data unit, one or more coding tree units, a coding tree block, one or more coding units, a coding tree unit row, a slice, a tile, a sub-picture, or a view of a video. In some embodiments, the rule that the predictor palette is reset or reinitialized applies to a video unit regardless of whether wavefront parallel processing of multiple video units is enabled. In some embodiments, a video unit includes a coding tree unit row corresponding to a chrominance component. In some embodiments, the first block includes a first coding tree block corresponding to a chrominance component in a coding tree unit row. In some embodiments, the rule that the predictor palette is reset or reinitialized applies to a video unit in the case where a dual-tree split is applied and the current split tree is a coding tree corresponding to a chrominance component. In some embodiments, the size of the predictor palette is reset or reinitialized to 0. In some embodiments, the size of the predictor palette is reset or reinitialized to the number of entries in a sequence palette predictor initializer or the maximum number of entries allowed in a predictor palette signaled in a coded representation.
[0410] In some embodiments, the predictor palette is further reset or reinitialized before converting a new video unit. In some embodiments, in the case of enabling wavefront parallel processing of multiple video units, the predictor palette for converting the current coding tree block or the current coding tree unit is determined based on the converted coding tree blocks or coding tree units.
[0411] Figure 16 is a flowchart representation of a video processing method 1600 according to the present technology. At operation 1610, method 1600 includes performing a conversion between a video unit of a video and a coded representation of the video using a palette mode. The video unit includes a plurality of blocks. During the conversion, a shared predictor palette is used by all of the plurality of blocks to predict a palette of representative sample values for each of the plurality of blocks in the predictor palette mode.
[0412] In some embodiments, a ternary tree segmentation is applied to the video unit, and wherein the shared predictor palette is used for video units with dimensions of 16×4 or 4×16. In some embodiments, a binary tree segmentation is applied to the video unit, and the shared predictor palette is used for video units with dimensions of 8×4 or 4×8. In some embodiments, a quadtree segmentation is applied to the video unit, and the shared predictor palette is used for video units with dimensions of 8×8. In some embodiments, the shared predictor palette is constructed before the conversion of all of the plurality of blocks within the video unit.
[0413] In some embodiments, for a first coded block of a plurality of blocks in a region, an indication of a predicted entry in the shared predictor palette is signaled in the coded representation. In some embodiments, for the remainder of the plurality of blocks in the region, the indication of the predicted entry in the shared predictor palette is omitted in the coded representation. In some embodiments, the update of the shared predictor palette is skipped after the conversion of one of the plurality of blocks in the region.
[0414] Figure 17 is a flowchart representation of a video processing method 1700 according to the present technology. At operation 1710, method 1700 includes performing a conversion between a current block of a video and a coded representation of the video using a palette mode, in which a palette of representative sample values is used to code the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and a counter indicating the usage frequency of the corresponding entry is maintained for each entry of the predictor palette.
[0415] In some embodiments, for a new entry to be added to the predictor palette, the counter is set to K, where K is an integer. In some embodiments, K = 0. In some embodiments, whenever the corresponding entry is reused during the conversion of the current block, the counter is incremented by N, where N is a positive integer. In some embodiments, N = 1.
[0416] In some embodiments, before the predictor palette is used for transformation, the entries of the predictor palette are reordered according to a rule. In some embodiments, the rule stipulates that the entries of the predictor palette are reordered according to the codec information of the current sample. In some embodiments, the rule stipulates that the entries of the predictor palette are reordered according to the counter of each corresponding entry of the entries in the predictor palette.
[0417] In some embodiments, a second counter is used to indicate the occurrence frequency of samples. In some embodiments, in the predictor palette, a first sample with a higher occurrence frequency is located before a second sample with a lower occurrence frequency. In some embodiments, the predictor palette is updated according to a rule using the escape samples in the current block. In some embodiments, the rule stipulates that the predictor palette is updated using the escape samples when a condition is met. In some embodiments, the condition is met when the predictor palette is not full after inserting the current block of the current block.
[0418] Figure 18 is a flowchart representation of a method 1800 for video processing according to the present technology. At operation 1810, method 1800 includes performing a transformation between a current block of a video and a codec representation of the video using a palette mode, in which a palette of representative sample values is used to codec the current block to predict the palette of representative sample values of the current block. The number of palette entries signaled in the codec representation is in the range of [0, the maximum allowable size of the palette - the number of palette entries derived during transformation], which is a closed range including 0 and the maximum allowable size of the palette - the number of palette entries derived during transformation.
[0419] In some embodiments, the number of palette entries signaled in the codec representation is binarized based on this range. In some embodiments, a truncated binary codec process is used to binarize the number of entries signaled in the codec representation. In some embodiments, the number of entries signaled in the codec representation is binarized based on the characteristics of the current block. In some embodiments, the characteristics include the dimensions of the current block.
[0420] Figure 19 is a flowchart representation of a method 1900 for video processing according to the present technology. At operation 1910, method 1900 includes performing a transformation between a current block in a video unit of a video and a codec representation of the video using a palette mode, in which a palette of representative sample values is used to codec the current block. During the transformation, a predictor palette is used to predict the palette of representative sample values, and wherein the size of the predictor palette is adaptively adjusted according to a rule.
[0421] In some embodiments, the size of the predictor palette in the dual-tree segmentation is different from that in the single-tree segmentation. In some embodiments, a video unit includes a block, a coding / decoding unit, a coding / decoding tree unit, a slice, a tile, or a sub-picture. In some embodiments, the rule specifies that the predictor palette has a first size for a video unit and a different second size for the transformation of a subsequent video unit. In some embodiments, the rule specifies adjusting the size of the predictor palette according to the size of the current palette used for the transformation. In some embodiments, the size of the predictor palette is equal to the size of the current palette determined after the transformation of the current block. In some embodiments, the size of the predictor palette is equal to the size of the current palette determined after the transformation of the current block plus or minus an offset, where the offset is an integer. In some embodiments, the offset is signaled in the coded representation. In some embodiments, the offset is derived during the transformation.
[0422] In some embodiments, the rule specifies a predefined size S for the predictor palette of the current block, and the rule further specifies adjusting the size of the predictor palette according to the size of the current block. In some embodiments, when the size of the current block is less than or equal to T, the size of the predictor palette is adjusted to be less than the predefined size S, where T and S are integers. In some embodiments, the first K entries in the predictor palette are used for the transformation, where K is an integer and K ≤ S. In some embodiments, a subsampled predictor palette with a size less than the predefined size S is used for the transformation. In some embodiments, when the size of the current block is greater than or equal to T, the size of the predictor palette is adjusted to the predefined size S.
[0423] In some embodiments, K or T is determined based on characteristics of the video. In some embodiments, the characteristics of the video include the content of the video. In some embodiments, the characteristics of the video include information signaled in any of the following: a decoder parameter set, a strip parameter set, a video parameter set, a picture parameter set, an adaptive parameter set, a picture header, a strip header, a slice group header, a largest coding unit (LCU), a coding / decoding unit, an LCU row, an LCU group, a transform unit, a picture unit, or a video coding / decoding unit in the coded representation. In some embodiments, the characteristics of the video include the position of a coding / decoding unit, a picture unit, a transform unit, a block, or a video coding / decoding unit within the video. In some embodiments, the characteristics of the video include an indication of the color format of the video. In some embodiments, the characteristics of the video include the coding / decoding tree structure applicable to the video. In some embodiments, the characteristics of the video include the strip type, slice group type, or picture type of the video. In some embodiments, the characteristics of the video include the color components of the video. In some embodiments, the characteristics of the video include the temporal layer identifier of the video. In some embodiments, the characteristics of the video include the profile, level, or tier of the video standard.
[0424] In some embodiments, the rule stipulates that the size of the predictor palette is adjusted according to one or more counters of each entry in the predictor palette. In some embodiments, during conversion, entries with a counter less than a threshold T, where T is an integer, are discarded. In some embodiments, the entry with the smallest counter is discarded until the size of the predictor palette is less than a threshold T, where T is an integer.
[0425] In some embodiments, the rule stipulates that the predictor palette is updated only based on the current palette used for conversion. In some embodiments, the predictor palette is updated to the current palette of a subsequent block after conversion.
[0426] Figure 20 is a flowchart of a method 2000 for video processing according to the present technology. Method 2000 includes, at operation 2010, performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and the size of the palette of representative samples or the predictor palette is determined according to a rule that allows the size to vary between video units of the video.
[0427] In some embodiments, a video unit includes a block, a coding / decoding unit, a coding / decoding tree unit, a block, a tile, or a sub-picture. In some embodiments, the size of the palette of representative samples or the predictor palette is further determined based on the characteristics of the current block or neighboring blocks of the current block. In some embodiments, the characteristics include the dimensions of the current block or neighboring blocks. In some embodiments, the characteristics at least include the quantization parameter of the current block or neighboring blocks. In some embodiments, the characteristics include the color components of the current block or neighboring blocks.
[0428] In some embodiments, different sizes of the palette of representative samples or the predictor palette are used for different color components. In some embodiments, the size of the palette of representative samples or the predictor palette of the luminance component and the chrominance component is indicated in the encoded / decoded representation. In some embodiments, the size of the palette of representative samples or the predictor palette of each color component is indicated in the encoded / decoded representation. In some embodiments, the signaling of different sizes in the encoded / decoded representation is based on the use of a dual-tree segmentation, a slice type, or a picture type used for conversion.
[0429] Figure 21It is a flowchart representation of a video processing method 2100 according to the present technology. At operation 2110, method 2100 includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode / decoded the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. The predictor palette is reinitialized when a condition is met, where the condition is met when the video unit is the first video unit in a video unit row and an syntax element indicating that wavefront parallel processing is enabled for the video unit is included in the encoded / decoded representation.
[0430] In some embodiments, a video unit includes a coding tree unit or a coding tree block. In some embodiments, the condition is met if the current block and a previous block are not in the same tile. In some embodiments, after the conversion of the video unit, at least one syntax element is maintained to record the size of the predictor palette and / or the number of entries in the predictor palette. In some embodiments, at least one syntax element is used for the conversion of the current block.
[0431] In some embodiments, in case (1) the current block is in the first column of a picture or case (2) the current block and a previous block are not in the same tile, a storage procedure of context variables of the video is invoked. In some embodiments, the output of the storage procedure includes at least the size of the predictor palette or the number of predictor palette entries.
[0432] In some embodiments, the applicability of one or more of the above methods is based on characteristics of the video. In some embodiments, characteristics of the video include the content of the video. In some embodiments, characteristics of the video include information signaled in any of the following: decoder parameter sets, slice parameter sets, video parameter sets, picture parameter sets, adaptive parameter sets, picture headers, slice headers, picture group headers, largest coding units (LCUs), coding units, LCU rows, LCU groups, transform units, picture units, or video coding units in the coded representation. In some embodiments, characteristics of the video include the position of a coding unit, picture unit, transform unit, block, or video coding unit within the video. In some embodiments, characteristics of the video include characteristics of a current block or neighboring blocks of the current block. In some embodiments, characteristics of the current block or neighboring blocks of the current block include the dimensions of the current block or the dimensions of the neighboring blocks of the current block. In some embodiments, characteristics of the video include an indication of the color format of the video. In some embodiments, characteristics of the video include the coding tree structure applicable to the video. In some embodiments, characteristics of the video include the slice type, group type, or picture type of the video. In some embodiments, characteristics of the video include the color components of the video. In some embodiments, characteristics of the video include the temporal layer identifier of the video. In some embodiments, characteristics of the video include the profile, level, or tier of the video standard.
[0433] In some embodiments, the transformation includes encoding the video into a coded representation. In some embodiments, the transformation includes decoding the coded representation to generate pixel values of the video.
[0434] Some embodiments of the disclosed techniques include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the transformation from video blocks to the bitstream representation of the video will use the video processing tool or mode when the video processing tool or mode is enabled based on the decision or determination. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the transformation from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the decision or determination.
[0435] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode when converting video blocks into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode enabled based on the decision or determination.
[0436] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations of one or more of them. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or combinations of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, e.g., including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated for encoding information for transmission to a suitable receiver device.
[0437] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the relevant program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or can be deployed to execute on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0438] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be executed by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special-purpose logic circuitry.
[0439] For example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0440] Although this patent document contains many details, these details should not be construed as limiting the scope of any subject matter or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be excluded from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.
[0441] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0442] Only some embodiments and examples are described, and other embodiments, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: For the conversion between the current video block of a video and the bitstream of the video, determining to apply a prediction mode to the current video block, wherein, in the prediction mode, the reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Performing the conversion based on the reconstructed samples, wherein, when the value of the entropy coding / decoding synchronization flag is equal to 1, the coding / decoding tree block (CTB) including the current video block is the first CTB in the CTB row of a slice and the predetermined spatial neighboring block is not available, a variable specifying the size of the palette predictor is initialized to 0.
2. The method according to claim 1, wherein The position of the upper-left luminance sample of the predetermined spatial neighboring block is (x0, y0 - CtbSizeY), where (x0, y0) represents the position of the upper-left luminance sample of the CTB including the current video block, and where CtbSizeY represents the size of the CTB.
3. The method according to claim 2, wherein The variable is initialized to 0 when starting to parse the CTB syntax.
4. The method according to claim 1, wherein When the value of the entropy coding / decoding synchronization flag is equal to 1, the predetermined spatial neighboring block is available and the value of the sequence palette enable flag is equal to 1, a synchronization process of the palette predictor is called.
5. The method according to claim 1, wherein The entropy coding / decoding synchronization flag specifying whether to call a specific synchronization process and a specific storage process of context variables is included in the bitstream.
6. The method according to claim 1, wherein, The palette predictor has different maximum sizes for video blocks with tree type being single tree and tree type being double tree.
7. The method according to claim 6, wherein The maximum size of the palette predictor is a fixed integer value.
8. The method according to claim 1, wherein When applying a single tree to the current video block, the palette predictor includes three color components, wherein, when applying a double tree to the current video block and the current video block is a chrominance block, the palette predictor includes two chrominance color components, and wherein, when applying a double tree to the current video block and the current video block is a luminance block, the palette predictor includes one color component.
9. The method according to claim 1, wherein, The conversion includes encoding the current video block into the bitstream.
10. The method according to claim 1, wherein The conversion includes decoding the current video block from the bitstream.
11. A device for processing video data, the device comprising a processor and a non-transitory memory storing instructions thereon, wherein, The instruction, when run by the processor, causes the processor to: For the conversion between the current video block of a video and the bitstream of the video, determining to apply a prediction mode to the current video block, wherein, in the prediction mode, the reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Performing the conversion based on the reconstructed samples, wherein, when the value of the entropy coding / decoding synchronization flag is equal to 1, the coding / decoding tree block (CTB) including the current video block is the first CTB in the CTB row of a slice and the predetermined spatial neighboring block is not available, a variable specifying the size of the palette predictor is initialized to 0.
12. The apparatus according to claim 11, wherein, The position of the top-left luma sample of the predetermined spatial neighboring block is (x0, y0 - CtbSizeY), where (x0, y0) represents the position of the top-left luma sample of the CTB including the current video block, and where CtbSizeY represents the size of the CTB.
13. The apparatus according to claim 12, wherein, When the value of the entropy coding / decoding synchronization flag is equal to 1, the predetermined spatial neighboring block is available, and the value of the sequence palette enable flag is equal to 1, the synchronization process of the palette predictor is invoked.
14. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing a processor to: For the conversion between the current video block of a video and the bitstream of the video, it is determined to apply a prediction mode to the current video block, where, In the prediction mode, the reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Perform the transformation based on the reconstructed samples, where, when the value of the entropy coding / decoding synchronization flag is equal to 1, the coding / decoding tree block (CTB) including the current video block is the first CTB in the CTB row of the slice and the predetermined spatial neighboring block is not available, the variable specifying the size of the palette predictor is initialized to 0.
15. The non-transitory computer-readable storage medium according to claim 14, wherein, The position of the top-left luma sample of the predetermined spatial neighboring block is (x0, y0 - CtbSizeY), where (x0, y0) represents the position of the top-left luma sample of the CTB including the current video block, and where CtbSizeY represents the size of the CTB.
16. The non-transitory computer-readable storage medium according to claim 15, wherein, When the value of the entropy coding / decoding synchronization flag is equal to 1, the predetermined spatial neighboring block is available, and the value of the sequence palette enable flag is equal to 1, the synchronization process of the palette predictor is invoked.
17. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein, The method includes: For the transformation between the current video block of a video and the bitstream of the video, determine to apply a prediction mode to the current video block, where, in the prediction mode, the reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Generate the bitstream based on the reconstructed samples, where, when the value of the entropy coding / decoding synchronization flag is equal to 1, the coding / decoding tree block (CTB) including the current video block is the first CTB in the CTB row of the slice and the predetermined spatial neighboring block is not available, the variable specifying the size of the palette predictor is initialized to 0.
18. A method for storing a bitstream of a video, including: For the transformation between the current video block of a video and the bitstream of the video, determine to apply a prediction mode to the current video block, where, in the prediction mode, the reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Generate the bitstream based on the reconstructed samples; and Store the bitstream in a non-transitory computer-readable recording medium, Among them, when the value of the entropy encoding / decoding synchronization flag is equal to 1, including that the coding tree block (CTB) of the current video block is the first CTB in the CTB row of the slice and the predetermined spatial neighboring block is not available, the variable specifying the size of the palette predictor is initialized to 0.
Citation Information
Patent Citations
Palette predictor initialization and merge for video coding
CN108028932A
Image Processing Method and Apparatus
US20150010087A1