Palette Predictor Size Adaptation in Video Coding and Decoding
By using palette mode and predictor palette optimization methods in video encoding and codec, the problem of inefficient motion vector management and predictor palette update in the prior art is solved, and more efficient video encoding and codec performance is achieved.
Patent Information
- Application Number
- CN202080064453.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-09-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-09-11
AI Technical Summary
Existing video codec technology has inefficient problems in motion vector management and predictor palette updates, resulting in poor codec performance.
The palette mode is used to convert the video block and the codec representation, the palette of representative sample point values is used for codec, and the sample point value is predicted through the predictor palette. The update and size of the predictor palette is prohibited or adaptively adjusted according to the characteristics of the current block.
Improves the efficiency and performance of video encoding and codec, optimizes the update and size of the predictor palette, reduces unnecessary computing and storage, and improves compression ratio and parallel implementation capabilities.
Smart Images

Figure CN114402612B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] Under the applicable patent laws and / or the rules applicable to the Paris Convention, this application timely claims the priority and benefits of International Patent Application PCT / CN2019 / 105554, filed on September 12, 2019. For all legal purposes, the entire disclosure of the above - mentioned application is incorporated by reference into a part of the disclosure of this application. Technical field
[0003] This patent document relates to video coding and decoding technologies, devices, and systems. Background art
[0004] Currently, efforts are being made to improve the performance of current video codec technologies to provide a better compression ratio, or to provide video coding and decoding schemes that allow for lower complexity or parallel implementation. Industry experts have recently proposed several new video coding and decoding tools, and tests are currently being conducted to determine their effectiveness. Summary of the invention
[0005] Devices, systems, and methods related to digital video coding and decoding, particularly related to the management of motion vectors, are described. The described methods can be applied to existing video coding and decoding standards (e.g., High Efficiency Video Coding (HEVC) or Versatile Video Coding) and future video coding and decoding standards or video codecs.
[0006] In a representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and a coded - decoded representation of the video using a palette mode, where a palette of representative sample values is used to code - decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. Based on the characteristics of the current block, updating the predictor palette is prohibited after the conversion of the current block according to a rule.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block of a video and a coded - decoded representation of the video using a palette mode, where a palette of representative sample values is used to code - decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. Whether to perform a change to the predictor palette is determined according to the color components of the current block.
[0008] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, multiple predictor palettes are used to predict the palette of representative sample values.
[0009] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. Before the conversion of the first block in the video unit or after the conversion of the last video block in a previous video unit, the predictor palette is reset or reinitialized according to a rule.
[0010] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a video unit of a video and an encoded / decoded representation of the video using a palette mode. The video unit includes multiple blocks. During the conversion, a shared predictor palette is used by all the multiple blocks to predict the palette of representative sample values for each of the multiple blocks in a predictor palette mode.
[0011] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and a counter indicating the usage frequency of a corresponding entry is maintained for each entry of the predictor palette.
[0012] In another representative aspect, the disclosed techniques can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block to predict the palette of representative sample values of the current block. The number of entries of the palette signaled in the encoded / decoded representation is in the range of [0, the maximum allowed size of the palette - the number of palette entries derived during the conversion].
[0013] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and the size of the predictor palette is adaptively adjusted according to rules.
[0014] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and the size of the palette of representative samples or the predictor palette is determined according to rules, where the rules allow the size to change between video units of the video.
[0015] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, where a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. The predictor palette is reinitialized when a condition is met, where the condition is met when the video unit is the first video unit in a video unit row and a syntax element indicating that wavefront parallel processing is enabled for the video unit is included in the encoded / decoded representation.
[0016] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block. Further, the predictor palette is selectively reset before the conversion between the video block and the bitstream representation of the video block.
[0017] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block. Further, when multiple encoded / decoded units of the video unit have a common shared area, the predictor palette is a shared predictor palette.
[0018] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the size of the predictor palette is adaptively changed according to one or more conditions.
[0019] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the predictor palette is updated based on the size or number of entries in the predictor palette.
[0020] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the entries of the predictor palette are reordered or modified.
[0021] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, where, during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, the use of the predictor palette is indicated by maintaining a counter that tracks the number of times the predictor palette is used.
[0022] In another example aspect, the method described above can be implemented by a video decoder device including a processor.
[0023] In another example aspect, the method described above can be implemented by a video encoder device including a processor.
[0024] Further, in a representative aspect, a device in a video system is disclosed, the device including a processor and a non-transitory memory having instructions thereon. When the instructions are executed by the processor, the processor is caused to implement any one or more of the disclosed methods.
[0025] Further, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product including program code for performing any one or more of the disclosed methods.
[0026] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings, the specification, and the claims. Description of the Drawings
[0027] Figure 1 Shows an example of a block encoded and decoded in palette mode.
[0028] Figure 2 Shows an example of notifying palette entries using predictor palette signaling.
[0029] Figure 3 Shows examples of horizontal traversal scan and vertical traversal scan.
[0030] Figure 4 Shows an example of encoding and decoding of palette indices.
[0031] Figure 5 Shows an example of a picture with an 18 by 12 luminance CTU divided into 12 slices and 3 raster scan stripes.
[0032] Figure 6 Shows an example of a picture with an 18 by 12 luminance CTU divided into 24 slices and 9 rectangular stripes.
[0033] Figure 7 Shows an example of a picture divided into 4 slices, 11 bricks, and 4 rectangular stripes.
[0034] Figure 8 Shows an example of a picture with 28 sub - pictures.
[0035] Figure 9 Is a block diagram of an example hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.
[0036] Figure 10 Is a block diagram of an example video processing system that can implement the disclosed technology.
[0037] Figure 11 Shows a flowchart of an example method for video encoding and decoding.
[0038] Figure 12 Is a flowchart representation of a method for video processing according to the present technology.
[0039] Figure 13 Is a flowchart representation of another method for video processing according to the present technology.
[0040] Figure 14 Is a flowchart representation of another method for video processing according to the present technology.
[0041] Figure 15It is a flowchart representation of another video processing method according to the present technology.
[0042] Figure 16 It is a flowchart representation of another video processing method according to the present technology.
[0043] Figure 17 It is a flowchart representation of another video processing method according to the present technology.
[0044] Figure 18 It is a flowchart representation of another video processing method according to the present technology.
[0045] Figure 19 It is a flowchart representation of another video processing method according to the present technology.
[0046] Figure 20 It is a flowchart representation of another video processing method according to the present technology.
[0047] Figure 21 It is a flowchart representation of yet another video processing method according to the present technology. Detailed implementation manners
[0048] 1. Video encoding and decoding of HEVC / H.265
[0049] Video encoding and decoding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) as well as the H.265 / HEVC standard. Since H.262, video encoding and decoding standards have been based on a hybrid video encoding and decoding structure that utilizes temporal prediction and transform coding. To explore future video encoding and decoding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, JVET was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC.
[0050] 2. Palette mode
[0051] 2.1 Palette mode of HEVC Screen Content Coding Extension (HEVC-SCC)
[0052] 2.1.1. Concept of palette mode
[0053] The basic idea behind the palette mode is that the pixels in a CU are represented by a small set of representative color values. This set is called the palette. Also, out-of-palette samples can be indicated by signaling the escape symbol for subsequent (possibly quantized) component values. Such pixels are called escape pixels. The palette mode is shown in Figure 1 As shown in Figure 1 , for each pixel with three color components (luma and two chroma components), an index into the palette is established, and the block can be reconstructed based on the values found in the palette.
[0054] 2.1.2. Encoding and Decoding of Palette Entries
[0055] For the palette encoding / decoding block, the following key aspects are introduced:
[0056] 1) Construct the current palette based on the predictor palette and the new entries (if any) signaled for the current palette.
[0057] 2) Divide the current sample / pixel into two categories: one category (the first category) includes the samples / pixels in the current palette, and the other category (the second category) includes the samples / pixels outside the current palette.
[0058] a. For the samples / pixels in the second category, quantization (at the encoder) is applied to the samples / pixels and the quantization values are signaled; and dequantization (at the decoder) is applied.
[0059] 2.1.2.1. Predictor Palette
[0060] For the encoding and decoding of palette entries, a predictor palette is maintained, which is updated after decoding the palette encoding / decoding block.
[0061] 2.1.2.1.1. Initialization of the Predictor Palette
[0062] Initialize the predictor palette at the start of each slice and each tile.
[0063] The palette and the maximum size of the predictor palette are signaled in the SPS. In HEVC-SCC, the palette_predictor_initializer_present_flag is introduced in the PPS. When this flag is 1, the entries used to initialize the predictor palette are signaled in the bitstream.
[0064] Depending on the value of palette_predictor_initializer_present_flag, the size of the predictor palette is reset to 0 or initialized using the predictor palette initializer entry signaled in the PPS. In HEVC-SCC, a predictor palette initializer of size 0 is enabled to allow explicit disabling of the predictor palette initialization at the PPS level.
[0065] The corresponding syntax, semantics, and decoding process are defined as follows. Newly added text is shown in bold and underlined italics. Any deleted text is marked with [[ ]].
[0066] 7.3.2.2.3 Sequence parameter set screen content coding extension syntax
[0067]
[0068] Equal to 1 specifies that the decoding process for the palette mode can be used for intra blocks. palette_mode_enabled_flag equal to 0 specifies that the decoding process for the palette mode is not applied. When not present, the value of palette_mode_enabled_flag is inferred to be equal to 0.
[0069] Specifies the maximum allowed size of the palette. When not present, the value of palette_max_size is inferred to be 0.
[0070] Specifies the difference between the maximum allowed size of the palette predictor and the maximum allowed size of the palette. When not present, the value of delta_palette_max_predictor_size is inferred to be 0. The derivation of the variable PaletteMaxPredictorSize is as follows:
[0071] PaletteMaxPredictorSize = palette_max_size + delta_palette_max_predictor_size (2 - 1)
[0072] The requirement for bitstream conformance is that when palette_max_size is equal to 0, the value of delta_palette_max_predictor_size should be equal to 0.
[0073] Equal to 1 specifies that the sequence palette predictor is initialized using the sps_palette_predictor_initializers specified in this clause. An sps_palette_predictor_initializer_flag equal to 0 specifies that the entries in the sequence palette predictor are initialized to 0. When not present, the value of sps_palette_predictor_initializer_flag is inferred to be equal to 0.
[0074] The requirement for bitstream conformance is that when palette_max_size is equal to 0, the value of sps_palette_predictor_initializer_present_flag shall be equal to 0.
[0075] Plus 1 specifies the number of entries in the sequence palette predictor initializer.
[0076] The requirement for bitstream conformance is that the value of sps_num_palette_predictor_initializer_minus1 plus 1 shall be less than or equal to PaletteMaxPredictorSize.
[0077] Specifies the value of the comp-th component of the i-th palette entry used to initialize the array PredictorPaletteEntries in the SPS. For values of i in the range 0 to sps_num_palette_predictor_initializer_minus1 (inclusive), the value of sps_palette_predictor_initializers[0][i] shall be in the range 0 to (1 << BitDepth Y ) – 1 (inclusive), and the values of sps_palette_predictor_initializers[1][i] and sps_palette_predictor_initializers[2][i] shall be in the range 0 to (1 << BitDepth C ) – 1 (inclusive).
[0078] 7.3.2.3.3 Picture Parameter Set Screen Content Coding Extension Syntax
[0079]
[0080] When pps_palette_predictor_initializer_flag equals 1, the palette predictor initializer for the picture that references the PPS is derived from the palette predictor initializer of the PPS. When pps_palette_predictor_initializer_flag equals 0, the palette predictor initializer for the picture that references the PPS is inferred to be equal to the palette predictor initializer specified by the active SPS. When not present, the value of pps_palette_predictor_initializer_present_flag is inferred to be equal to 0.
[0081] For bitstream conformance, when palette_max_size equals 0 or palette_mode_enabled_flag equals 0, the value of pps_palette_predictor_initializer_present_flag shall equal 0.
[0082] Specifies the number of entries in the picture palette predictor initializer.
[0083] For bitstream conformance, the value of pps_num_palette_predictor_initializer shall be less than or equal to PaletteMaxPredictorSize.
[0084] Initialize the palette predictor variables as follows:
[0085] – If the coding tree unit is the first coding tree unit in the slice, the following applies:
[0086] – Call the initialization process of the palette predictor variables as specified in Clause 9.3.2.3.
[0087] – Otherwise, if equals 1, and CtbAddrInRs % PicWidthInCtbsY equals 0 or TileId[CtbAddrInTs] is not equal to TileId[CtbAddrRsToTs[CtbAddrInRs - 1]], the following applies:
[0088] – Derive the position (xNbT, yNbT) of the top - left luma sample of the spatial neighboring block T from the position (x0, y0) of the top - left luma sample of the current coding tree block as follows:
[0089] (xNbT,yNbT) = (x0 + CtbSizeY,y0 - CtbSizeY)(9 - 3)
[0090] – Call the availability derivation process of the block that calls the z-scan order specified in Clause 6.4.1, where the position (xCurr, yCurr) set to be equal to (x0, y0) and the neighboring positions (xNbY, yNbY) set to be equal to (xNbT, yNbT) are used as inputs, and the output is assigned to availableFlagT.
[0091] – The synchronization process calls of context variables, Rice parameter initialization status, and palette predictor variables are as follows:
[0092] – If availableFlagT is equal to 1, call the synchronization process of context variables, Rice parameter initialization status, and palette predictor variables specified in Clause 9.3.2.5, with TableStateIdxWpp, TableMpsValWpp, TableStatCoeffWpp, PredictorPaletteSizeWpp, and TablePredictorPaletteEntriesWpp as inputs.
[0093] – Otherwise, the following applies:
[0094] – Call the initialization process of the palette predictor variables specified in Clause 9.3.2.3.
[0095] – Otherwise, if CtbAddrInRs is equal to slice_segment_address and dependent_slice_segment_flag is equal to 1, call the synchronization process of context variables and Rice parameter initialization status specified in Clause 9.3.2.5, with TableStateIdxDs, TableMpsValDs, TableStatCoeffDs, PredictorPaletteSizeDs, and TablePredictorPaletteEntriesDs as inputs.
[0096] – Otherwise, the following applies:
[0097] – Call the initialization process of the palette predictor variables specified in Clause 9.3.2.3.
[0098] 9.3.2.3 Initialization Process of Palette Predictor Entries
[0099] The output of this process is the initialized palette predictor variables PredictorPaletteSize and PredictorPaletteEntries.
[0100] The derivation of the variable numComps is as follows:
[0101] numComps = (ChromaArrayType == 0)? 1 : 3(9 - 8)
[0102] – If pps_palette_predictor_initializer_present_flag is equal to 1, the following applies:
[0103] – PredictorPaletteSize is set to be equal to pps_num_palette_predictor_initializer.
[0104] – The derivation of the array PredictorPaletteEntries is as follows:
[0105]
[0106] – Otherwise (pps_palette_predictor_initializer_present_flag is equal to 0), if sps_palette_predictor_initializer_present_flag is equal to 1, the following applies:
[0107] – PredictorPaletteSize is set to be equal to sps_num_palette_predictor_initializer_minus1 plus 1.
[0108] – The derivation of the array PredictorPaletteEntries is as follows:
[0109]
[0110] – Otherwise (pps_palette_predictor_initializer_present_flag is equal to 0 and sps_palette_predictor_initializer_present_flag is equal to 0), PredictorPaletteSize is set to be equal to 0.
[0111] 2.1.2.1.2. Use of the Predictor Palette
[0112] For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. This is in Figure 2It is shown in. The run - length encoding and decoding of zero is used to send the reuse flag. After that, the number of new palette entries is signaled using an exponential Golomb (EG) code of order 0 (e.g., EG - 0). Finally, the component values of the new palette entries are signaled.
[0113] 2.1.2.2. Update of the predictor palette
[0114] The update of the predictor palette is performed using the following steps:
[0115] (1) Before decoding the current block, there is a predictor palette, denoted as PltPred0
[0116] (2) The current palette table is constructed by first inserting the items in PltPred0 and then inserting the new entries of the current palette.
[0117] (3) Construct PltPred1:
[0118] a. First, add the items in the current palette table (which may include the items in PltPred0)
[0119] b. If not full, add the unreferenced items in PltPred0 in ascending order of entry index.
[0120] 2.1.3. Encoding and decoding of palette indices
[0121] As Figure 3 shown, horizontal and vertical traversal scans are used to encode and decode the palette indices. The scan order is explicitly signaled in the bitstream using palette_transpose_flag. For the remainder of the sub - clause, assume the scan is horizontal.
[0122] Two palette sample modes are used to encode and decode the palette indices: "COPY_LEFT" and "COPY_ABOVE". In the "COPY_LEFT" mode, the palette index is assigned to the decoded index. In the "COPY_ABOVE" mode, the palette index of the sample in the above row is copied. For both the "COPY_LEFT" and "COPY_ABOVE" modes, a run value is signaled, which specifies the number of subsequent samples encoded using the same mode.
[0123] In the palette mode, the index value of an escape sample is the number of palette entries. Also, when the escape symbol is part of a run in the "COPY_LEFT" or "COPY_ABOVE" mode, an escape component value is signaled for each escape symbol. The encoding and decoding of the palette indices are shown in Figure 4 in.
[0124] This syntax order is completed as follows. First, the number of index values for the CU is signaled. Then the actual index values for the entire CU are signaled using truncated binary coding. Both the number of indexes and the index values are coded in bypass mode. This groups the bypass binary bits related to the indexes together. Then, the palette sample mode (if necessary) and run lengths are signaled in an interleaved manner. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and coded in bypass mode. The binarization of the escape samples is EG coding with three orders, e.g., EG-3.
[0125] After signaling the index values, the additional syntax element last_run_type_flag is signaled. This syntax element, in combination with the number of indexes, eliminates the need to signal the run value corresponding to the last run in the block.
[0126] In HEVC-SCC, the palette mode also applies to 4:2:2, 4:2:0, and monochrome chroma formats. For all chroma formats, the signaling of palette entries and palette indexes is almost the same. If it is a non-monochrome format, each palette entry consists of 3 components. For the monochrome format, each palette entry consists of a single component. For subsampled chroma directions, chroma samples are associated with luminance sample indexes divisible by 2. After reconstructing the palette index of the CU, if the sample has only a single component associated with it, only the first component of the palette entry is used. The only difference in signaling lies in the escape component values. For each escape sample, the number of escape component values signaled may be different, depending on the number of components associated with that sample.
[0127] In addition, there is an index adjustment process in palette index coding. When signaling the palette index, the left neighboring index or the upper neighboring index should be different from the current index. Therefore, by removing one possibility, the range of the current palette index can be reduced by 1. After that, the index is signaled using truncated binary (TB) binarization.
[0128] The text related to this part is shown below, where CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the predicted index.
[0129] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is the index of the array represented by CurrentPaletteEntries. The array indexes xC, yC specify the position (xC, yC) of the sample relative to the top-left luminance sample of the picture. The value of PaletteIndexMap[xC][yC] should be in the range from 0 to MaxPaletteIndex (inclusive).
[0130] The derivation of the variable adjustedRefPaletteIndex is as follows:
[0131]
[0132]
[0133] When CopyAboveIndicesFlag[xC][yC] is equal to 0, the derivation of the variable CurrPaletteIndex is as follows:
[0134] if(CurrPaletteIndex >= adjustedRefPaletteIndex)
[0135] CurrPaletteIndex++
[0136] 2.1.3.1. Decoding process of the palette coding / decoding block
[0137] 1) Read the prediction information to mark which entries in the predictor palette will be reused;
[0138] (palette_predictor_run)
[0139] 2) Read the new palette entries of the current block
[0140] a.num_signaled_palette_entries
[0141] b.new_palette_entries
[0142] 3) Construct CurrentPaletteEntries based on a) and b)
[0143] 4) Read the escape symbol presence flag: palette_escape_val_present_flag to derive MaxPaletteIndex
[0144] 5) Code / Decode how many samples that are not coded in copy mode / run mode
[0145] a.num_palette_indices_minus1
[0146] b. For each sample that is not coded in copy mode / run mode, code / Decode palette_idx_idc in the current plt table
[0147] 2.2. Palette mode in VVC
[0148] 2.2.1. Palette in the dual tree
[0149] In VVC, the dual - tree coding / decoding structure is used to code / decode intra - frame strips. Therefore, the luminance component and the two chrominance components may have different palettes and palette indices. Additionally, the two chrominance components share the same palette and palette index.
[0150] 2.2.2. Palette as a separate mode
[0151] In some embodiments, the prediction mode of the coding / decoding unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. The binarization of the prediction mode is changed accordingly.
[0152] When IBC is disabled, on I - slices, the first binary bit is used to indicate whether the current prediction mode is MODE_PLT. On P / B - slices, the first binary bit is used to indicate whether the current prediction mode is MODE_INTRA. If not, an additional binary bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTER.
[0153] When IBC is enabled, on I - slices, the first binary bit is used to indicate whether the current prediction mode is MODE_IBC. If not, the second binary bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. On P / B - slices, the first binary bit is used to indicate whether the current prediction mode is MODE_INTRA. If it is an intra - mode, the second binary bit is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. If not, the second binary bit is used to indicate whether the current prediction mode is MODE_IBC or MODE_INTER.
[0154] An example syntax text is shown as follows.
[0155] Coding / decoding unit syntax
[0156]
[0157]
[0158] 2.3. Division of pictures, sub - pictures, strips, slices, tiles, and CTUs
[0159] Sub - picture: A rectangular area of one or more strips within a picture.
[0160] Strip: An integer number of tiles of a picture that are uniquely contained within a single NAL unit. A strip consists of a sequence of multiple complete slices or complete tiles of a slice.
[0161] Slice: A rectangular region of CTUs within a specific slice column and a specific slice row in a picture.
[0162] Tile: A rectangular region of CTU rows within a specific slice in a picture. A slice can be divided into multiple tiles, each tile consisting of one or more CTU rows within the slice. A slice that is not divided into multiple tiles is also referred to as a tile. However, a tile that is a proper subset of a slice is not called a slice.
[0163] Tile scan: A specific order sorting of CTU partitions of a picture, where CTUs are sorted consecutively in the CTU raster scan of a tile, tiles within a slice are sorted consecutively in the raster scan of the tiles of the slice, and slices within a picture are sorted consecutively in the raster scan of the slices of the picture.
[0164] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that cover a rectangular region of the picture.
[0165] A slice is divided into one or more tiles, each tile consisting of multiple CTU rows within the slice.
[0166] A slice that is not divided into multiple tiles is also referred to as a tile. However, a tile that is a proper subset of a slice is not called a slice.
[0167] A strip contains multiple slices of a picture or multiple tiles of a slice.
[0168] A sub-picture contains one or more strips that jointly cover a rectangular region of the picture.
[0169] Two strip modes are supported, namely the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains a sequence of slices in the slice raster scan of the picture. In the rectangular strip mode, a strip contains multiple tiles of the picture that jointly form a rectangular region of the picture. The tiles within a rectangular strip are in the raster scan order of the tiles of the strip.
[0170] Figure 5 An example of the raster scan strip partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0171] Figure 6 An example of the rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0172] Figure 7An example of a picture divided into slices, tiles, and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular strips.
[0173] Figure 8 An example of sub-picture segmentation of a picture is shown, where the picture is segmented into 28 sub-pictures of different dimensions.
[0174] When encoding and decoding a picture using three separate color planes (separate_colour_plane_flag equals 1), a strip contains only CTUs of one color component, which is identified by the corresponding value of colour_plane_id, and each color component array of the picture consists of strips with the same colour_plane_id value. The encoded and decoded strips with different colour_plane_id values within a picture can be interleaved with each other under the following constraints. For each value of colour_plane_id, the encoded and decoded strip NAL units with that colour_plane_id value shall be in the order of increasing CTU addresses in the tile scan order of the first CTU of each encoded and decoded strip NAL unit.
[0175] When separate_colour_plane_flag equals 0, each CTU of the picture is exactly contained in one strip. When separate_colour_plane_flag equals 1, each CTU of a color component is exactly contained in one strip (for example, the information of each CTU of the picture exists exactly in three strips, and these three strips have different colour_plane_id values).
[0176] 2.4. Wavefront with 1-CTU Delay
[0177] In VVC, a one-CTU delay wavefront parallel processing (WPP) is adopted instead of the two-CTU delay in the HEVC design. WPP processing enables multiple parallel processes with limited encoding and decoding losses, but the two-CTU delay may hinder the parallel processing ability. Since the target resolution is becoming larger and the number of CPUs is increasing, it is asserted that improving the parallel processing ability by using the proposed one-CTU delay is beneficial for reducing the encoding and decoding delay and will make full use of the CTU ability.
[0178] 3. Example Problems in Existing Implementations
[0179] DMVR and BIO do not involve the original signal during the refinement of motion vectors, which may lead to inaccurate motion information of the coded / decoded blocks. In addition, DMVR and BIO sometimes adopt fractional motion vectors after motion refinement, while screen video usually uses integer motion vectors, which makes the current motion information more inaccurate and the coding / decoding performance worse.
[0180] (1) The current palette is constructed based on the prediction of the previous coded / decoded palette. The current palette is only reinitialized before decoding a new CTU row or a new slice when the entropy_coding_sync_enabled_flag is equal to 1. However, in practical applications, parallel encoders are preferred, where different CTU rows can be pre-coded / decoded without referring to the information of other CTU rows.
[0181] (2) The way to handle the predictor palette update process is fixed. That is, the entries inherited from the previous predictor palette and the new entries in the current palette are inserted in sequence. If the number of entries is still less than the size of the predictor palette, further entries not inherited from the previous predictor palette will be added. This design does not consider the importance of different entries in the current predictor palette and the previous predictor palette.
[0182] (3) The size of the predictor palette is fixed, and after decoding a block, it must be updated to fill all entries, which is suboptimal because some of these entries may never be referenced.
[0183] (4) The size of the current palette is fixed, regardless of the color component. For example, fewer chroma samples can be used compared to luminance.
[0184] 4. Example Techniques and Embodiments
[0185] The detailed embodiments described below should be regarded as examples for explaining the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way.
[0186] In addition to DMVR and BIO mentioned below, the methods described below may also be applicable to other decoder motion information derivation techniques.
[0187] Re - hierarchical predictor palette
[0188] 1. It is proposed to reset or reinitialize the predictor palette (e.g., entries and / or the size of the predictor palette) before decoding the first block in a new video unit.
[0189] a. Alternatively, after decoding the last block in a video unit, the predictor palette (e.g., the entries and / or the size of the predictor palette) can be reset or re-initialized.
[0190] b. In one example, the video unit is a sub-region of a CTU (e.g., a VPDU)
[0191] / CTU / CTB / multiple CTUs / multiple CUs / CTU rows / slices / bricks / sub-pictures / views, etc.
[0192] i. Alternatively, in addition, even if the wavefront is disabled (e.g.,
[0193] the entropy_coding_sync_enabled_flag is equal to 0), the above method is called.
[0194] c. In one example, the video unit is a chrominance CTU row.
[0195] i. Alternatively, in addition, before decoding the first chrominance CTB in a new chrominance CTU row, the predictor palette can be reset or re-initialized.
[0196] ii. Alternatively, in addition, when applying the dual-tree and the current split tree is a chrominance codec tree, the above method is called.
[0197] d. In one example, the size of the predictor palette (e.g., in the specification
[0198] ) is reset to 0.
[0199] e. In one example, the size of the predictor palette (e.g., in the specification
[0200] ) is reset to the number of entries in the sequence palette predictor initializer (e.g., plus 1) or the maximum number of entries allowed in the predictor palette (e.g.,
[0201] ).
[0202] f. The initialization of the predictor palette (e.g., PredictorPaletteEntries) before encoding / decoding a new sequence / picture can be used to initialize the predictor palette before encoding / decoding a new video unit.
[0203] g. In one example, when the entropy_coding_sync_enabled_flag is equal to 1, the predictor palette after encoding / decoding the upper CTB / CTU can be used to initialize the predictor palette before encoding / decoding the current CTB / CTU.
[0204] 2. It is proposed to prohibit updating the predictor palette after encoding / decoding a specific palette coding block.
[0205] a. In one example, whether to update the predictor palette can depend on the decoding information of the current block.
[0206] i. In one example, whether to update the predictor palette can depend on the block dimension of the current block.
[0207] 1. In one example, if the width of the current block is not greater than a first threshold (denoted by T1) and the height of the current block is not greater than a second threshold (denoted by T2), the update process is disabled.
[0208] 2. In one example, if the product of the width of the current block and the height of the block is not greater than a first threshold (denoted by T1), the update process is disabled.
[0209] 3. In one example, if the width of the current block is not less than a first threshold (denoted by T1) and the height of the current block is not less than a second threshold (denoted by T2), the update process is disabled.
[0210] 4. In one example, if the product of the width of the current block and the height of the block is not less than a first threshold (denoted by T1), the update process is disabled.
[0211] 5. In the above examples, T1 / T2 can be predefined or signaled.
[0212] a) In one example, T1 / T2 can be set to 4, 16, or 1024.
[0213] b) In one example, T1 / T2 can depend on the color component.
[0214] 3. A shared predictor palette can be defined, where the same predictor palette can be used for all CUs / PUs under the shared region.
[0215] a. In one example, a shared region can be defined for an MxN region (e.g., 16×4 or 4×16 region) using TT partitioning.
[0216] b. In one example, a shared region can be defined for an MxN region (e.g., 8×4 or 4×8 region) using BT partitioning.
[0217] c. In one example, a shared region may be defined for an MxN region (e.g., an 8×8 region) partitioned using QT.
[0218] d. Alternatively, in addition, a shared predictor palette may be constructed before encoding / decoding all blocks within the shared region.
[0219] e. In one example, an indication of a prediction entry in the predictor palette (e.g., palette_predictor_run) may be signaled together with the first palette-coded block within the shared region.
[0220] i. Alternatively, in addition, for the remaining coded blocks within the shared region, signaling of an indication of a prediction entry in the predictor palette (e.g.,
[0221] palette_predictor_run) may be skipped.
[0222] f. Alternatively, in addition, after decoding / encoding blocks within the shared region, updating of the predictor palette may always be skipped.
[0223] 4. A counter may be maintained for each entry of the predictor palette to indicate the frequency with which it is used.
[0224] a. In one example, for each new entry added to the predictor palette, the counter may be set to a constant K.
[0225] i. In one example, K may be set to 0.
[0226] b. In one example, when an entry is marked for reuse in the coded / decoded palette block, the corresponding counter may be incremented by a constant N.
[0227] i. In one example, N may be set to 1.
[0228] 5. It is proposed to adaptively change the size of the predictor palette instead of using a predictor palette of a fixed size.
[0229] a. In one example, it may change between video units (blocks / CUs / CTUs / slices / bricks / sub-pictures) and another video unit.
[0230] b. In one example, the size of the predictor palette may be updated according to the size of the current palette.
[0231] i. In one example, the size of the predictor palette may be set to the size of the current palette after decoding / encoding the current block.
[0232] ii. In one example, the size of the predictor palette can be set to the size of the current palette minus or plus an integer value represented as K.
[0233] 1. In one example, K can be signaled / derived on the fly.
[0234] c. In one example, the size of the predictor palette can depend on the block size. Let S be the predefined size of the predictor palette for the palette-coded block.
[0235] i. In one example, a palette-coded block with a size less than or equal to T can use a predictor palette with a size less than S.
[0236] 1. In one example, the first K entries (K <= S) in the palette predictor can be used.
[0237] 2. In one example, a subsampled version of the palette predictor can be used.
[0238] ii. In one example, a palette-coded block with a size greater than or equal to T can use a predictor palette with a size equal to S.
[0239] iii. In the above examples, K and / or T are integers and can be based on
[0240] 1. Video content (e.g., screen content or natural content)
[0241] 2. Messages signaled in the DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header /
[0242] largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit
[0243] 3. The position of the CU / PU / TU / block / video coding unit
[0244] 4. Indication of the color format (e.g., 4:2:0, 4:4:4, RGB, or YUV)
[0245] 5. Coding tree structure (e.g., binary tree or single tree)
[0246] 6. Strip / slice group type and / or picture type
[0247] 7. Color component
[0248] 8. Temporal layer ID
[0249] 9. Profile / level / tier of the standard
[0250] d. In one example, after encoding / decoding the color palette block, the predictor palette can be customized according to the counter of the entries.
[0251] i. In one example, entries with a counter less than the threshold T can be discarded.
[0252] ii. In one example, entries with the minimum counter value can be discarded until the size of the predictor palette is less than the threshold T.
[0253] e. Alternatively, in addition, after decoding / encoding the palette codec block, the predictor palette can be updated based only on the current palette.
[0254] i. Alternatively, in addition, after decoding / encoding the palette codec block, the predictor palette can be updated to the current palette.
[0255] 6. Entries in the current palette and / or the predictor palette before encoding / decoding the current block can be reordered / modified before being used to update the predictor palette.
[0256] a. In one example, reordering can be applied according to the decoding information / reconstruction of the current sample.
[0257] b. In one example, reordering can be applied according to the counter value of the entries.
[0258] c. Alternatively, in addition, the number of occurrences of samples / pixels (in the current palette and / or outside the current palette) can be counted.
[0259] i. Alternatively, in addition, samples / pixels with a larger counter (e.g., occurring more frequently) can be placed before another sample / pixel with a smaller counter.
[0260] 7. Information about escaped samples can be used to update the predictor palette.
[0261] a. Alternatively, in addition, an update of the predictor palette using the escape information can be conditionally invoked.
[0262] i. In one example, when the predictor palette is not full after inserting the current palette, the escaped sample / pixel information can be added to the predictor palette.
[0263] 8. Updating / initializing / reseting the predictor palette can depend on the color component.
[0264] a. In one example, the rule for determining whether to update the predictor palette can depend on the color component, such as luminance or chrominance.
[0265] 9. A set of multiple predictor palettes can be maintained and / or updated.
[0266] a. In one example, a predictor palette can have information for one or all color components.
[0267] b. In one example, a predictor palette can have information for two color components (e.g., Cb and Cr).
[0268] c. In one example, at least one global palette and at least one local palette can be maintained.
[0269] i. In one example, the predictor palette can be updated based on the global palette and the local palette.
[0270] d. In one example, a palette associated with the last K palette encoding / decoding blocks (in encoding / decoding order) can be maintained.
[0271] e. In one example, palettes for the luminance component and the chrominance component can be predicted based on different predictor palettes (e.g., different indices of a set of multiple predictor palettes).
[0272] f. Alternatively, in addition, bullet point 1 can be applied to the set of predictor palettes.
[0273] g. Alternatively, in addition, the index / multiple indices of the predictor palette in the set of predictor palettes can be signaled for a sub-region of a CU / PU / CTU / CTB / CTU or CTB.
[0274] Regarding palette / predictor palette size
[0275] 10. The size of the palette can vary between one video unit and another video unit.
[0276] a. In one example, it can vary between a video unit (block / CU / CTU / slice / tile / sub-picture)
[0277] and another video unit.
[0278] b. In one example, it can depend on the decoding information of the current block and / or neighboring (adjacent or non-adjacent)
[0279] blocks.
[0280] 11. The size of the palette and / or the predictor palette can depend on the block dimension and / or the quantization parameter.
[0281] 12. The size (or the number of entries therein) of the palette and / or the predictor palette can be different for different color components.
[0282] a. In one example, an indication of the size of the palette and / or predictor palette for the luminance component and chrominance components may be signaled explicitly or implicitly.
[0283] b. In one example, an indication of the size of the palette and / or predictor palette for each color component may be signaled explicitly or implicitly.
[0284] c. In one example, whether to signal an indication of multiple sizes may depend on the use of the dual tree and / or the slice / picture type.
[0285] Signaling of the palette
[0286] 13. The compliant bitstream shall satisfy that the number of entries signaled directly for the current block (e.g.,
[0287] ) shall be within the range [0, palette_max_size -
[0288] NumPredictedPaletteEntries], which is a closed range including 0 and
[0289] palette_max_size - NumPredictedPaletteEntries.
[0290] a. How to binarize num_signaled_palette_entries may depend on the allowed range.
[0291] i. Truncated binarization coding may be used instead of EG-0 th .
[0292] b. How to binarize num_signaled_palette_entries may depend on the decoded information (e.g., block dimensions).
[0293] Regarding the wavefront with 1 - CTU
[0294] 14. It is proposed to re-initialize the predictor palette (e.g., entries and / or size) when parsing of the CTU syntax ends (e.g., in clause 7.3.8.2 of VVC), entropy_coding_sync_enabled_flag equals 1, and the current CTB is the first in a new CTU row or the current CTB is not in the same tile as its previous CTB.
[0295] a. Alternatively, in addition, maintain PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp to record the updated size and entries of the predictor palette after completion of the encoding / decoding of the above CTU.
[0296] i. Alternatively, in addition, PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp can be used for encoding / decoding the current block in the current CTU.
[0297] b. In one example, when the parsing of the CTU syntax in Clause 7.3.8.2 is ended, entropy_coding_sync_enabled_flag is equal to 1, and when CtbAddrInRs %
[0298] PicWidthInCtbsY is equal to 0 or BrickId[CtbAddrInBs] is not equal to BrickId[CtbAddrRsToBs[CtbAddrInRs - 1]], call the storage procedure of the context variables specified in Clause 9.3.2.3 with TableStateIdx0Wpp, TableStateIdx1Wpp, TableMpsValWpp, PredictorPaletteSizeWpp, and PredictorPaletteEntriesWpp (when palette_mode_enabled_flag is equal to 1)
[0299] as output.
[0300] Overview
[0301] 15. Whether and / or how to apply the above method may be based on the following:
[0302] a. Video content (such as screen content or natural content)
[0303] b. Messages signaled in DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header / largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit
[0304] c. The position of CU / PU / TU / block / video coding unit
[0305] d. Decoding information of the current block and / or its neighboring blocks
[0306] i. Block dimension / block shape of the current block and / or its neighboring blocks
[0307] e. Indication of color format (e.g., 4:2:0, 4:4:4, RGB or YUV)
[0308] f. Coding tree structure (e.g., dual tree or single tree)
[0309] g. Slice / tile group type and / or picture type
[0310] h. Color component (e.g., can be applied only to the luminance component and / or chrominance component)
[0311] i. Temporal layer ID
[0312] j. Profile / level / hierarchy of the standard
[0313] 5. Additional embodiments
[0314] In the following embodiments, the newly added text is shown in bold and underlined italic text. Any deleted text is marked by [[ ]].
[0315] 5.1. Embodiment #1
[0316] 9.3.1 Overview This procedure is called when parsing a syntax element using the descriptor ae(v) in Clauses 7.3.8.1 to 7.3.8.12.
[0317] The input to this procedure is a request for the value of the syntax element and the values of previously parsed syntax elements.
[0318] The output of this procedure is the value of the syntax element.
[0319] When starting to parse one or more of the following, the initialization procedure specified in Clause 9.3.2 is called:
[0320] 1. The slice segmentation data syntax specified in Clause 7.3.8.1,
[0321] 2. The CTU syntax specified in Clause 7.3.8.2 and the CTU is the first CTU in [[slice]], 3. The CTU syntax specified in Clause 7.3.8.2, [[entropy_coding_sync_enabled_flag is equal to 1 and]] the associated luminance CTB is the first luminance CTB in the CTU row of [[slice]].
[0322] The parsing process of the syntax element is as follows:
[0323] When cabac_bypass_alignment_enabled_flag is equal to 1, for a request for the value of a syntax element targeting syntax element coeff_abs_level_remaining[] or coeff_sign_flag[] and escapeDataPresent is equal to 1, call the calibration process before calibration bypass decoding as specified in Clause 9.3.4.3.6.
[0324] For each requested value of a syntax element, perform the derived binarization as specified in Clause 9.3.3.
[0325] The binarization of the syntax element and the parsed binary bit sequence determine the decoding process, as described in Clause 9.3.4.
[0326] When processing a request for the value of a syntax element for pcm_flag and the decoded value of pcm_flag is equal to 1, initialize the decoding engine after decoding any pcm_alignment_zero_bit and all pcm_sample_luma and pcm_sample_chroma data as specified in Clause 9.3.2.6. The storage process of context variables is applied as follows:
[0327] – When finishing parsing the CTU syntax in Clause 7.3.8.2, entropy_coding_sync_enabled_flag is equal to 1, and CtbAddrInRs %
[0328] PicWidthInCtbsY is equal to 1, or both CtbAddrInRs are greater than 1 and TileId[CtbAddrInTs] is not equal to TileId[CtbAddrRsToTs[CtbAddrInRs - 2]], call the storage process of context variables, Rice parameter initialization status, and palette predictor variables as specified in Clause 9.3.2.4, where TableStateIdxWpp, TableMpsValWpp,
[0329] TableStatCoeffWpp (when persistent_rice_adaptation_enabled_flag is equal to 1),
[0330] PredictorPaletteSizeWpp and PredictorPaletteEntriesWpp (when palette_mode_enabled_flag is equal to 1) are used as outputs.
[0331] – When the parsing of the overall strip segment data syntax in Clause 7.3.8.1 is completed, dependent_slice_segments_enabled_flag is equal to 1 and end_of_slice_segment_flag is equal to 1, call the storage procedures for context variables, Rice parameter initialization status, and palette predictor variables specified in Clause 9.3.2.4, where TableStateIdxDs, TableMpsValDs, TableStatCoeffDs (when persistent_rice_adaptation_enabled_flag is equal to 1),
[0332] PredictorPaletteSizeDs and PredictorPaletteEntriesDs (when palette_mode_enabled_flag is equal to 1) are used as outputs.
[0333] 5.2. Example #2
[0334] 9.3 CABAC Parsing Process of Strip Data
[0335] 9.3.1 Overview
[0336] The input to this process is a request for the value of a syntax element and the values of previously parsed syntax elements.
[0337] The output of this process is the value of the syntax element.
[0338] When starting to parse the CTU syntax specified in Clause 7.3.8.2 and one or more of the following conditions are true, call the initialization process specified in Clause 9.3.2,
[0339] – The CTU is the first CTU in the tile.
[0340] – The value of entropy_coding_sync_enabled_flag is equal to 1, and the CTU is the first CTU in the CTU row of the tile.
[0341] The parsing process of the syntax element is as follows:
[0342] For each requested value of the syntax element, obtain the binarization as specified in Subclause 9.3.3.
[0343] The binarization of the syntax element and the sequence of parsed binary bits determine the decoding process, as described in Subclause 9.3.4.
[0344] The storage procedure for context variables is applied as follows:
[0345] – When parsing the CTU syntax in Clause 7.3.8.2 is completed, entropy_coding_sync_enabled_flag is equal to 1, and CtbAddrInRs % PicWidthInCtbsY is equal to 0 or BrickId[CtbAddrInBs] is not equal to BrickId[CtbAddrRsToBs[CtbAddrInRs - 1]], call the storage procedure of the context variables specified in Clause 9.3.2.3, where TableStateIdx0Wpp, TableStateIdx1Wpp, and TableMpsValWpp are used as outputs.
[0346] 9.3.2 Initialization Process
[0347] 9.3.2.1 Overview
[0348] – The output of this process is the initialized CABAC internal variables.
[0349] – The context variables of the arithmetic decoding engine are initialized as follows:
[0350] – If the CTU is the first CTU in the brick, call the initialization process of the context variables as specified in Clause 9.3.2.2, and the variables PredictorPaletteSize[0 / 1 / 2] are initialized to 0.
[0351] – Otherwise, if entropy_coding_sync_enabled_flag is equal to 1, and CtbAddrInRs % PicWidthInCtbsY is equal to 0 or BrickId[CtbAddrInBs] is not equal to BrickId[CtbAddrRsToBs[CtbAddrInRs - 1]], the following applies:
[0352] – Use the position (x0, y0) of the top - left luma sample of the current CTB to derive the position (xNbT, yNbT) of the top - left luma sample of the spatial neighborhood block T as follows:
[0353] ( xNbT, yNbT ) = ( x0, y0 - CtbSizeY ) (9 - 3)
[0354] – Invoke the derivation process of neighboring block availability specified in Clause 6.4.4, where the position (xCurr, yCurr) is set to be equal to (x0, y0), the neighboring position (xNbY, yNbY) is set to be equal to (xNbT, yNbT), checkPredModeY is set to be equal to FALSE, cIdx is set to be equal to 0 as the input, and the output is assigned to availableFlagT.
[0355] – The synchronization process of context variables is invoked as follows:
[0356] – If availableFlagT is equal to 1, then invoke the synchronization process of context variables specified in Clause 9.3.2.4, where TableStateIdx0Wpp, TableStateIdx1Wpp, TableMpsValWpp are used as inputs, and the variable PredictorPaletteSize is initialized to 0.
[0357] – Otherwise, invoke the initialization process of context variables as specified in Clause 9.3.2.2, and the variable PredictorPaletteSize is initialized to 0.
[0358] – Otherwise, invoke the initialization process of context variables as specified in Clause 9.3.2.2, and the variable PredictorPaletteSize is initialized to 0.
[0359] – Initialize the decoding engine registers ivlCurrRange and ivlOffset with 16-bit register precision by invoking the initialization process of the arithmetic decoding engine specified in Sub-clause 9.3.2.5.
[0360] 9.3.2.3 Storage Process of Context Variables
[0361] The inputs of this process include:
[0362] – CABAC context variables indexed by ctxTable and ctxIdx.
[0363] The outputs of this process include:
[0364] – variables tableStateSync0, tableStateSync1, and tableMPSSync that contain the values of variables pStateIdx0, pStateIdx1, and valMps used during the initialization of context variables, where these context variables are assigned to all syntax elements in Clauses 7.3.8.1 to 7.3.8.11, except for end_of_brick_one_bit and end_of_subset_one_bit.
[0365]
[0366] For each context variable, the corresponding entries pStateIdx0, pStateIdx1, and valMps in tables tableStateSync0, tableStateSync1, and tableMPSSync are initialized to the corresponding pStateIdx0, pStateIdx1, and valMps.
[0367]
[0368] Alternatively, the following may apply:
[0369]
[0370] 5.3. Example #3
[0371]
[0372] Alternatively, the in the above table can be set to another integer value, such as a fixed value or the predictor palette size.
[0373] 6. Example Implementations of the Disclosed Technology
[0374] Figure 9 is a block diagram of a video processing device 900. The device 900 can be used to implement one or more of the methods described herein. The device 900 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 900 can include one or more processors 902, one or more memories 904, and video processing hardware 906. The processor 902 can be configured to implement one or more of the methods described in this document. The memory 904 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 906 can be used to implement some of the techniques described in this document in hardware circuitry and can be partially or fully part of the processor 902 (e.g., a graphics processing unit core GPU or other signal processing circuitry).
[0375] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of a current video block may correspond to bits that are co-located within the bitstream or scattered at different locations as defined by the syntax. For example, a macroblock may be encoded based on transformed and decoded error residual values, and may also be decoded using bits in the header and other fields in the bitstream.
[0376] It will be understood that the disclosed methods and techniques will benefit video encoder and / or decoder embodiments by allowing the use of the techniques disclosed in this document incorporated into video processing devices such as smartphones, laptops, desktops, and similar devices.
[0377] Figure 10 FIG. 8 is a block diagram of an example video processing system 1000 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all components of system 1000. System 1000 may include an input 1002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8- or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0378] System 1000 may include a codec component 1004 that may implement various codec or encoding methods described in this document. The codec component 1004 may reduce the average bitrate of the video from input 1002 to the output of the codec component 1004 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1004 may be stored or transmitted via a connected communication, represented by component 1006. The stored or transmitted (or coded) representation of the video received at input 1002 may be used by component 1008 to generate pixel values or a displayable video that is sent to a display interface 1010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "codec" operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0379] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), a Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, an IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as a mobile phone, a laptop computer, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0380] Figure 11 is a flowchart of an example method 1100 for video processing. At 1110, method 1100 includes performing a conversion between a video block in a video unit and a coded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict current palette information of the video block, and further, wherein the predictor palette is selectively reset before the conversion between the video block and the bitstream representation of the video block.
[0381] Some embodiments may be described using the following clause-based format.
[0382] 1. A method for video processing, comprising:
[0383] performing a conversion between a video block in a video unit and a coded / decoded representation of the video block using a palette mode, wherein, during the conversion, a predictor palette is used to predict current palette information of the video block, and further, wherein the predictor palette is selectively reset before the conversion between the video block and the bitstream representation of the video block.
[0384] 2. The method according to clause 1, wherein the video unit includes one of the following: one or more coding tree units, one or more coding tree blocks, a sub-region of a coding tree unit or a coding tree block, or a view of a coding tree block row / slice / brick / sub-picture / coding tree unit.
[0385] 3. The method according to any one of clauses 1-2, wherein delayed wavefront parallel processing is disabled during the conversion.
[0386] 4. The method according to clause 3, wherein the entropy_coding_sync_enabled_flag is set to be equal to 0.
[0387] 5. The method according to clause 1, wherein the video unit is a row of chrominance coding tree units.
[0388] 6. The method according to clause 5, wherein the predictor palette is reset before decoding the first chrominance coding tree block (CTB) in a new chrominance CTU row.
[0389] 7. The method according to clause 5, wherein when a bi-predictive coding tree is applied and the current split of the bi-predictive coding tree is a chroma coding tree unit, the predictor palette is reset.
[0390] 8. The method according to clause 1, wherein the size of the predictor palette is reset to zero.
[0391] 9. The method according to clause 1, wherein the size of the predictor palette is reset to the number of entries in the sequence palette predictor initializer or the maximum number of allowed entries.
[0392] 10. The method according to clause 9, wherein the sequence palette predictor initializer is used to initialize the palette predictor before it is applied to a video unit.
[0393] 11. The method according to clause 1, wherein when the entropy_coding_sync_enabled_flag is set to 1, the palette predictor applied to a previous video block is re-initialized before it is applied to a video unit.
[0394] 12. The method according to clause 1, wherein updating the predictor palette is prohibited based on the coding information associated with the video unit.
[0395] 13. The method according to clause 12, wherein the coding information includes the dimensions of the video unit.
[0396] 14. The method according to clause 13, wherein updating the predictor palette is prohibited based on the dimensions of the video unit reaching one or more threshold conditions.
[0397] 15. The method according to clause 14, wherein the one or more threshold conditions are predefined.
[0398] 16. The method according to clause 14, wherein the one or more threshold conditions are signaled explicitly or implicitly in the coded representation of the video unit.
[0399] 17. A method for video processing, comprising:
[0400] Performing a conversion between a video block in a video unit and the coded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, wherein when multiple coded units of the video unit have a common shared region, the predictor palette is a shared predictor palette.
[0401] 18. The method according to clause 17, wherein the shared region is associated with any one of the following: TT partition, BT partition, QT partition.
[0402] 19. The method according to clause 17, wherein the shared predictor palette is constructed before being applied to a plurality of coding units.
[0403] 20. The method according to clause 17, wherein an indication of the use of the shared predictor palette is signaled explicitly or implicitly in the coded representation associated with the first palette coding unit of the shared region.
[0404] 21. The method according to clause 17, further comprising:
[0405] After the coding unit applied to a plurality of coding units, skipping updating the shared predictor palette.
[0406] 22. A method for video processing, comprising:
[0407] Performing a conversion between a video block in a video unit and a coded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict current palette information of the video block, and further wherein the size of the predictor palette is adaptively changed according to one or more conditions.
[0408] 23. The method according to clause 22, wherein the one or more conditions are at least associated with the following: the size of the previous palette information, the dimensions of the video unit, the content of the video unit, the color format of the video unit, the color components of the video unit, the coding tree structure of the video block, the relative position of the video block in the coded representation, the temporal layer ID of the video block, the slice / group type and / or picture type of the video block, or the profile / level / tier of the video block.
[0409] 24. A method for video processing, comprising:
[0410] Performing a conversion between a video block in a video unit and a coded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict current palette information of the video block, and further wherein the predictor palette is updated based on the size or number of entries in the predictor palette.
[0411] 25. The method according to clause 24, wherein the size of the predictor palette is updated from a previous video block to a current video block.
[0412] 26. The method according to clause 24, wherein the size of the predictor palette is signaled explicitly or implicitly in the coded representation.
[0413] 27. The method according to clause 24, wherein the size of the predictor palette depends on one or more of the following: the dimensions of the video block, the quantization parameter of the video block, or one or more color components of the video block.
[0414] 28. A method for video processing, comprising:
[0415] Performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, wherein the entries of the predictor palette are reordered or modified.
[0416] 29. The method according to clause 28, wherein when the entropy_coding_sync_enabled_flag is equal to 1, the entries of the predictor palette are reordered or modified.
[0417] 30. The method according to clause 28, wherein when the end of the codec tree unit syntax is encountered, the entries of the predictor palette are reordered or modified.
[0418] 31. The method according to clause 28, wherein when the current CTB is the first CTB in a new CTU row or the current CTB is not in the same tile as the previous CTB, the entries of the predictor palette are reordered or modified.
[0419] 32. A method for video processing, comprising:
[0420] Performing a conversion between a video block in a video unit and an encoded / decoded representation of the video block using a palette mode, wherein during the conversion, a predictor palette is used to predict the current palette information of the video block, and further, wherein the use of the predictor palette is indicated by maintaining a counter that tracks the number of times the predictor palette is used.
[0421] 33. The method according to any of the above clauses, wherein enabling or disabling the predictor palette is associated with at least one of the following: the size of the previous palette information, the dimensions of the video block, the content of the video block, the color format of the video block, the color components of the video block, the codec tree structure of the video block, the relative position of the video block in the encoded / decoded representation, the temporal layer ID of the video block, the slice / tile group type and / or picture type of the video block, or the profile / level / tier of the video block.
[0422] 34. The method according to any of the above clauses, wherein more than one predictor palette is used during the conversion.
[0423] 35. A video decoding device includes a processor configured to implement one or more of the methods recited in clauses 1 to 34.
[0424] 36. A video encoding device includes a processor configured to implement one or more of the methods recited in clauses 1 to 34.
[0425] 37. A computer program product has computer code stored thereon, which, when executed by a processor, causes the processor to implement any one of the methods recited in clauses 1 to 34.
[0426] 38. A method, device, or system described in this document.
[0427] Figure 12 is a flowchart representation of a method 1200 for video processing according to the present technology. At operation 1210, method 1200 includes performing a conversion between a current block of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode / decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and, based on characteristics of the current block, updating of the predictor palette is prohibited according to a rule after conversion of the current block.
[0428] In some embodiments, the characteristics of the current block include encoding / decoding information associated with the current block. In some embodiments, the characteristics of the current block include the dimensions of the current block. In some embodiments, the rule specifies that updating of the predictor palette is prohibited when the width of the current block is less than or equal to a first threshold and the height of the current block is less than or equal to a second threshold. In some embodiments, the rule specifies that updating of the predictor palette is prohibited when the height of the current block is less than or equal to the first threshold. In some embodiments, the rule specifies that updating of the predictor palette is prohibited when the width of the current block is greater than or equal to the first threshold and the height of the current block is greater than or equal to the second threshold. In some embodiments, the rule specifies that updating of the predictor palette is prohibited when the height of the current block is greater than or equal to the first threshold.
[0429] In some embodiments, the first threshold or the second threshold is predefined or signaled in the encoded / decoded representation. In some embodiments, the first threshold is 4, 16, or 1024. In some embodiments, the second threshold is 4, 16, or 1024. In some embodiments, the first threshold or the second threshold is based on a color component of the current block.
[0430] Figure 13is a flowchart representation of a video processing method 1300 according to the present technology. At operation 1310, method 1300 includes performing a conversion between a current block of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode / decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and it is determined whether to perform a change to the predictor palette based on the color components of the current block.
[0431] In some embodiments, the change to the predictor palette includes updating, initializing, or resetting the predictor palette. In some embodiments, the color components include luminance or chrominance components. In some embodiments, the predictor palette includes information corresponding to the color components of the current block. In some embodiments, the predictor palette includes information corresponding to all color components of the current block. In some embodiments, the predictor palette includes information corresponding to two chrominance components of the current block.
[0432] Figure 14 is a flowchart representation of a video processing method 1400 according to the present technology. At operation 1410, method 1400 includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode / decode the current block. During the conversion, multiple predictor palettes are used to predict the palette of representative sample values.
[0433] In some embodiments, the predictor palette of the current block is updated at least based on a global palette and a local palette. In some embodiments, the multiple predictor palettes are associated with K blocks in a video unit that has been encoded / decoded using the palette mode. In some embodiments, palettes for different color components are determined based on different predictor palettes of the multiple predictor palettes. In some embodiments, the multiple predictor palettes are reset or reinitialized before the conversion of the first block in a video unit or after the conversion of the last block in a previously converted video unit. In some embodiments, the index of the predictor palette of the multiple predictor palettes is signaled in an encoded / decoded unit, a prediction unit, a coding tree unit, a coding tree block, a sub-region of a coding tree unit, or a sub-region of a coding tree block in the encoded / decoded representation.
[0434] Figure 15It is a flowchart representation of a video processing method 1500 according to the present technology. At operation 1510, method 1500 includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode / decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. According to a rule, the predictor palette is reset or reinitialized before the conversion of the first block in a video unit or after the conversion of the last video block in a previous video unit.
[0435] In some embodiments, a video unit includes a sub-region of a coding tree unit, a virtual pipeline data unit, one or more coding tree units, a coding tree block, one or more coding units, a coding tree unit row, a slice, a tile, a sub-picture, or a view of a video. In some embodiments, the rule that the predictor palette is reset or reinitialized applies to a video unit regardless of whether wavefront parallel processing of multiple video units is enabled. In some embodiments, a video unit includes a coding tree unit row corresponding to a chrominance component. In some embodiments, the first block includes a first coding tree block corresponding to a chrominance component in a coding tree unit row. In some embodiments, the rule that the predictor palette is reset or reinitialized applies to a video unit in the case where dual-tree splitting is applied and the current split tree is a coding tree corresponding to a chrominance component. In some embodiments, the size of the predictor palette is reset or reinitialized to 0. In some embodiments, the size of the predictor palette is reset or reinitialized to the number of entries in a sequence palette predictor initializer or the maximum number of entries allowed in a predictor palette signaled in the encoded / decoded representation.
[0436] In some embodiments, the predictor palette is further reset or reinitialized before a new video unit is converted. In some embodiments, in the case where wavefront parallel processing of multiple video units is enabled, a predictor palette for converting a current coding tree block or a current coding tree unit is determined based on the converted coding tree blocks or coding tree units.
[0437] Figure 16 It is a flowchart representation of a video processing method 1600 according to the present technology. At operation 1610, method 1600 includes performing a conversion between a video unit of a video and an encoded / decoded representation of the video using a palette mode. The video unit includes a plurality of blocks. During the conversion, a shared predictor palette is used by all of the plurality of blocks to predict the palette of representative sample values for each of the plurality of blocks in the predictor palette mode.
[0438] In some embodiments, the ternary tree segmentation is applied to a video unit, and a shared predictor palette is used for a video unit with dimensions of 16×4 or 4×16. In some embodiments, the binary tree segmentation is applied to a video unit, and a shared predictor palette is used for a video unit with dimensions of 8×4 or 4×8. In some embodiments, the quadtree segmentation is applied to a video unit, and a shared predictor palette is used for a video unit with dimensions of 8×8. In some embodiments, a shared predictor palette is constructed before the transformation of all multiple blocks within a video unit.
[0439] In some embodiments, for a first coded block of multiple blocks in a region, an indication of a prediction entry in the shared predictor palette is signaled in the coded representation. In some embodiments, for the remaining portion of the multiple blocks in the region, the indication of the prediction entry in the shared predictor palette is omitted in the coded representation. In some embodiments, the update of the shared predictor palette is skipped after the transformation of one of the multiple blocks in the region.
[0440] Figure 17 is a flowchart representation of a method 1700 for video processing according to the present technology. At operation 1710, method 1700 includes performing a transformation between a current block of a video and a coded representation of the video using a palette mode, in which a palette of representative sample values is used to code the current block. During the transformation, a predictor palette is used to predict the palette of representative sample values, and a counter indicating the usage frequency of the corresponding entry is maintained for each entry of the predictor palette.
[0441] In some embodiments, for a new entry to be added to the predictor palette, the counter is set to K, where K is an integer. In some embodiments, K = 0. In some embodiments, whenever the corresponding entry is reused during the transformation of the current block, the counter is incremented by N, where N is a positive integer. In some embodiments, N = 1.
[0442] In some embodiments, before the predictor palette is used for the transformation, the entries of the predictor palette are reordered according to a rule. In some embodiments, the rule specifies reordering the entries of the predictor palette according to the coding information of the current sample. In some embodiments, the rule specifies reordering the entries of the predictor palette according to the counter of each corresponding entry in the predictor palette.
[0443] In some embodiments, a second counter is used to indicate the occurrence frequency of sample points. In some embodiments, in a predictor palette, a first sample point with a higher occurrence frequency is located before a second sample point with a lower occurrence frequency. In some embodiments, the predictor palette is updated with escape sample points in the current block according to a rule. In some embodiments, the rule stipulates that the predictor palette is updated with escape sample points when a condition is met. In some embodiments, the condition is met when the predictor palette is not full after inserting the current block of the current block.
[0444] Figure 18 is a flowchart representation of a method 1800 for video processing according to the present technology. At operation 1810, method 1800 includes performing a conversion between a current block of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block to predict the palette of representative sample values of the current block. The number of palette entries signaled in the encoded / decoded representation is in the range of [0, maximum allowed size of the palette - number of palette entries derived during conversion], which is a closed range including 0 and the maximum allowed size of the palette - number of palette entries derived during conversion.
[0445] In some embodiments, the number of palette entries signaled in the encoded / decoded representation is binarized based on this range. In some embodiments, a truncated binary encoding / decoding process is used to binarize the number of entries signaled in the encoded / decoded representation. In some embodiments, the number of entries signaled in the encoded / decoded representation is binarized based on the characteristics of the current block. In some embodiments, the characteristics include the dimensions of the current block.
[0446] Figure 19 is a flowchart representation of a method 1900 for video processing according to the present technology. At operation 1910, method 1900 includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and wherein the size of the predictor palette is adaptively adjusted according to a rule.
[0447] In some embodiments, the size of the predictor palette in the dual-tree segmentation is different from that in the single-tree segmentation. In some embodiments, a video unit includes a block, a codec unit, a codec tree unit, a slice, a tile, or a sub-picture. In some embodiments, the rule specifies that the predictor palette has a first size for a video unit and a different second size for the transformation of a subsequent video unit. In some embodiments, the rule specifies that the size of the predictor palette is adjusted according to the size of the current palette used for the transformation. In some embodiments, the size of the predictor palette is equal to the size of the current palette determined after the transformation of the current block. In some embodiments, the size of the predictor palette is equal to the size of the current palette determined after the transformation of the current block plus or minus an offset, where the offset is an integer. In some embodiments, the offset is signaled in the codec representation. In some embodiments, the offset is derived during the transformation.
[0448] In some embodiments, the rule specifies a predefined size S for the predictor palette of the current block, and the rule further specifies that the size of the predictor palette is adjusted according to the size of the current block. In some embodiments, when the size of the current block is less than or equal to T, the size of the predictor palette is adjusted to be less than the predefined size S, where T and S are integers. In some embodiments, the first K entries in the predictor palette are used for the transformation, where K is an integer and K ≤ S. In some embodiments, a subsampled predictor palette with a size less than the predefined size S is used for the transformation. In some embodiments, when the size of the current block is greater than or equal to T, the size of the predictor palette is adjusted to the predefined size S.
[0449] In some embodiments, K or T is determined based on characteristics of the video. In some embodiments, the characteristics of the video include the content of the video. In some embodiments, the characteristics of the video include information signaled in any of the following: a decoder parameter set, a strip parameter set, a video parameter set, a picture parameter set, an adaptive parameter set, a picture header, a strip header, a slice group header, a largest coding unit (LCU), a codec unit, an LCU row, an LCU group, a transform unit, a picture unit, or a video codec unit in the codec representation. In some embodiments, the characteristics of the video include the position of a codec unit, a picture unit, a transform unit, a block, or a video codec unit within the video. In some embodiments, the characteristics of the video include an indication of the color format of the video. In some embodiments, the characteristics of the video include the codec tree structure applicable to the video. In some embodiments, the characteristics of the video include the strip type, slice group type, or picture type of the video. In some embodiments, the characteristics of the video include the color components of the video. In some embodiments, the characteristics of the video include the temporal layer identifier of the video. In some embodiments, the characteristics of the video include the profile, level, or tier of the video standard.
[0450] In some embodiments, the rule stipulates that the size of the predictor palette is adjusted according to one or more counters for each entry in the predictor palette. In some embodiments, during the conversion, entries with a counter less than a threshold T, where T is an integer, are discarded. In some embodiments, the entry with the smallest counter is discarded until the size of the predictor palette is less than a threshold T, where T is an integer.
[0451] In some embodiments, the rule stipulates that the predictor palette is updated based only on the current palette used for the conversion. In some embodiments, the predictor palette is updated to the current palette of a subsequent block after the conversion.
[0452] Figure 20 is a flowchart of a method 2000 for video processing according to the present technology. Method 2000 includes, at operation 2010, performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values, and the size of the palette of representative samples or the predictor palette is determined according to a rule that allows the size to vary between video units of the video.
[0453] In some embodiments, a video unit includes a block, a coding / decoding unit, a coding / decoding tree unit, a block, a tile, or a sub-picture. In some embodiments, the size of the palette of representative samples or the predictor palette is further determined based on the characteristics of the current block or neighboring blocks of the current block. In some embodiments, the characteristics include the dimensions of the current block or neighboring blocks. In some embodiments, the characteristics at least include the quantization parameter of the current block or neighboring blocks. In some embodiments, the characteristics include the color components of the current block or neighboring blocks.
[0454] In some embodiments, different sizes of the palette of representative samples or the predictor palette are used for different color components. In some embodiments, the size of the palette of representative samples or the predictor palette of the luminance component and the chrominance component is indicated in the encoded / decoded representation. In some embodiments, the size of the palette of representative samples or the predictor palette of each color component is indicated in the encoded / decoded representation. In some embodiments, the signaling of different sizes in the encoded / decoded representation is based on the use of a dual-tree segmentation, a slice type, or a picture type used for the conversion.
[0455] Figure 21It is a flowchart representation of a video processing method 2100 according to the present technology. At operation 2110, method 2100 includes performing a conversion between a current block in a video unit of a video and an encoded / decoded representation of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block. During the conversion, a predictor palette is used to predict the palette of representative sample values. The predictor palette is reinitialized when a condition is met, where the condition is met when the video unit is the first video unit in a video unit row and an indication of enabling wavefront parallel processing for the video unit is included in the encoded / decoded representation.
[0456] In some embodiments, a video unit includes a coding tree unit or a coding tree block. In some embodiments, the condition is met if the current block and a previous block are not in the same tile. In some embodiments, after the conversion of the video unit, at least one syntax element is maintained to record the size of the predictor palette and / or the number of entries in the predictor palette. In some embodiments, at least one syntax element is used for the conversion of the current block.
[0457] In some embodiments, in case (1) the current block is in the first column of a picture or case (2) the current block and a previous block are not in the same tile, a storage procedure of context variables of the video is invoked. In some embodiments, the output of the storage procedure includes at least the size of the predictor palette or the number of predictor palette entries.
[0458] In some embodiments, the applicability of one or more of the above methods is based on characteristics of the video. In some embodiments, characteristics of the video include the content of the video. In some embodiments, characteristics of the video include information signaled in any of the following: decoder parameter sets, slice parameter sets, video parameter sets, picture parameter sets, adaptive parameter sets, picture headers, slice headers, picture group headers, largest coding units (LCUs), coding units, LCU rows, LCU groups, transform units, picture units, or video coding units in a codec representation. In some embodiments, characteristics of the video include the positions of coding units, picture units, transform units, blocks, or video coding units within the video. In some embodiments, characteristics of the video include characteristics of the current block or neighboring blocks of the current block. In some embodiments, characteristics of the current block or neighboring blocks of the current block include the dimensions of the current block or the dimensions of the neighboring blocks of the current block. In some embodiments, characteristics of the video include an indication of the color format of the video. In some embodiments, characteristics of the video include a codec tree structure applicable to the video. In some embodiments, characteristics of the video include the slice type, group type, or picture type of the video. In some embodiments, characteristics of the video include the color components of the video. In some embodiments, characteristics of the video include the temporal layer identifier of the video. In some embodiments, characteristics of the video include the profile, level, or tier of the video standard.
[0459] In some embodiments, the transformation includes encoding the video into a codec representation. In some embodiments, the transformation includes decoding the codec representation to generate pixel values of the video.
[0460] Some embodiments of the disclosed techniques include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the transformation from video blocks to the bitstream representation of the video will use the video processing tool or mode when the video processing tool or mode is enabled based on the decision or determination. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the transformation from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the decision or determination.
[0461] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode when converting video blocks into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode enabled based on the decision or determination.
[0462] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, e.g., including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated for encoding information for transmission to a suitable receiver apparatus.
[0463] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the relevant program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or to execute on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0464] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be executed by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special-purpose logic circuitry.
[0465] By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0466] Although this patent document contains many details, these details should not be construed as limiting the scope of any subject matter or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.
[0467] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In addition, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0468] Only some embodiments and examples are described, and other embodiments, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: For the conversion between a first video unit of a video including a luminance block and two chrominance blocks and the bitstream of the video, determining to apply prediction modes to the luminance block and the two chrominance blocks respectively, wherein, in the prediction modes, reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Constructing a first palette for the luminance block based on a first palette prediction table, wherein the first palette is used to derive the reconstructed samples of the luminance block; Constructing a second palette for the two chrominance blocks based on a second palette prediction table, wherein the second palette is used to derive the reconstructed samples of the two chrominance blocks; and Performing the conversion based on the first palette and the second palette, wherein the luminance block and the two chrominance blocks have a tree type of double tree; wherein an indication of the size of the first palette and an indication of the size of the second palette are signaled explicitly respectively; and wherein the sizes of the first palette prediction table and the second palette prediction table are signaled implicitly.
2. The method according to claim 1, wherein The maximum sizes of the first palette prediction table and the second palette prediction table are fixed integer values.
3. The method according to claim 1, wherein The size of the first palette prediction table and the size of the first palette change for different luminance blocks, and the size of the second palette prediction table and the size of the second palette change for different chrominance blocks.
4. The method according to claim 1, wherein, The sizes of the first palette and the second palette are different.
5. The method according to claim 1, wherein, Syntax elements for deriving the first palette and the second palette are included in the bitstream respectively.
6. The method according to claim 1, wherein The conversion includes encoding the first video unit into the bitstream.
7. The method according to claim 1, wherein, The conversion includes decoding the first video unit from the bitstream.
8. The method according to claim 1, wherein For a video block with a tree type of single tree and a video block with a tree type of double tree, the first palette prediction table or the second palette prediction table has different maximum sizes.
9. The method according to claim 1, wherein The first palette prediction table is updated based on the first palette, and The second palette prediction table is updated based on the second palette.
10. The method for video processing according to claim 1, further comprising: Performing the conversion between a current block in a second video unit of the video and the bitstream of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block, wherein, during the conversion, a predictor palette is used to predict the palette of the representative sample values, and wherein the dimension of the predictor palette is adaptively adjusted according to rules.
11. The method according to claim 10, wherein The dimension of the predictor palette in double tree segmentation is different from the dimension of the predictor palette in single tree segmentation.
12. The method according to claim 10 or 11, wherein, The second video unit includes a block, a codec unit, a codec tree unit, a slice, a tile, or a subpicture.
13. The method according to claim 10 or 11, wherein The rule stipulates that the predictor palette has a first size for the second video unit and a different second size for the transformation of subsequent video units.
14. The method according to claim 10 or 11, wherein The rule stipulates that the size of the predictor palette is adjusted according to the size of the current palette of the transformation.
15. The method according to claim 14, wherein, The size of the predictor palette is equal to the size of the current palette determined after the transformation of the current block.
16. The method according to claim 14, wherein, The size of the predictor palette is equal to the size of the current palette determined after the transformation of the current block plus or minus an offset, where the offset is an integer.
17. The method according to claim 16, wherein, The offset is signaled in the bitstream.
18. The method according to claim 16, wherein, The offset is derived during the transformation.
19. The method according to claim 10 or 11, wherein, The rule stipulates a predefined size of the predictor palette of the current block as S, and wherein the rule further stipulates that the size of the predictor palette is adjusted according to the size of the current block.
20. The method according to claim 19, wherein, In the case where the size of the current block is less than or equal to T, the size of the predictor palette is adjusted to be less than the predefined size S, where T and S are integers.
21. The method according to claim 20, wherein, The first K entries in the predictor palette are used for the transformation, where K is an integer and K ≤ S.
22. The method according to claim 20, wherein, A subsampled predictor palette with a size less than the predefined size S is used for the transformation.
23. The method according to claim 19, wherein, In the case where the size of the current block is greater than or equal to T, the size of the predictor palette is adjusted to the predefined size S.
24. The method according to claim 20, wherein K or T is determined based on the characteristics of the video.
25. The method according to claim 24, wherein, The characteristics of the video include the content of the video.
26. The method according to claim 24, wherein, The characteristics of the video include the information signaled in any of the following: decoder parameter set, slice parameter set, video parameter set, picture parameter set, adaptive parameter set, picture header, slice header, slice group header, largest coding unit (LCU), coding unit, LCU row, LCU group, transform unit, picture unit, or video coding unit in the bitstream.
27. The method according to claim 24, wherein, The characteristics of the video include the position of the coding unit, picture unit, transform unit, block, or video coding unit within the video.
28. The method according to claim 24, wherein The characteristics of the video include an indication of the color format of the video.
29. The method according to claim 24, wherein The characteristics of the video include the coding tree structure applicable to the video.
30. The method according to claim 24, wherein, The characteristics of the video include the slice type, slice group type, or picture type of the video.
31. The method according to claim 24, wherein, The characteristics of the video include the color components of the video.
32. The method according to claim 24, wherein The characteristics of the video include the temporal layer identifier of the video.
33. The method according to claim 24, wherein, The characteristics of the video include the profile, level, or tier of the video standard.
34. The method according to claim 10 or 11, wherein, The rule stipulates that the size of the predictor palette is adjusted according to one or more counters of each entry in the predictor palette.
35. The method according to claim 34, wherein Entries with a counter less than a threshold T are discarded during the transformation, where T is an integer.
36. The method according to claim 34, wherein, Entries with the smallest counters are discarded until the size of the predictor palette is less than a threshold T, where T is an integer.
37. The method according to claim 10 or 11, wherein, The rule stipulates that the predictor palette is updated only based on the current palette for the transformation.
38. The method according to claim 37, wherein, After the transformation, the predictor palette is updated to the current palette for subsequent blocks.
39. The method for video processing according to claim 1 further includes: Performing a conversion between a current block in a third video unit of the video and the bitstream of the video using a palette mode, in which a palette of representative sample values is used to encode and decode the current block, wherein, during the conversion, a predictor palette is used to predict the palette of the representative sample values, and wherein the size of the palette of the representative samples or the predictor palette is determined according to a rule that allows the size to change between third video units of the video.
40. The method according to claim 39, wherein, The third video unit includes a block, a coding unit, a coding tree unit, a slice, a tile, or a sub-picture.
41. The method according to claim 39, wherein, Further determining the size of the palette of the representative samples or the predictor palette based on characteristics of the current block or neighboring blocks of the current block.
42. The method according to claim 41, wherein, The characteristics include the dimensions of the current block or the neighboring block.
43. The method according to claim 41, wherein, The characteristics at least include the quantization parameter of the current block or the neighboring block.
44. The method according to claim 42, wherein, The characteristics include the color component of the current block or the neighboring block.
45. The method according to claim 44, wherein, Different sizes of the palette of the representative samples or the predictor palette are used for different color components.
46. The method according to claim 45, wherein, Indicating in the bitstream the size of the palette of the representative samples or the predictor palette for the luminance component and the chrominance component.
47. The method according to claim 45, wherein, Indicating in the bitstream the size of the palette of the representative samples or the predictor palette for each color component.
48. The method according to any one of claims 45 to 47, wherein, The signaling of different sizes in the bitstream is based on the use of dual-tree partitioning, stripe type, or picture type in the conversion.
49. The method according to any one of claims 10 to 11 or 39 to 47, wherein, The conversion includes encoding the video into the bitstream.
50. The method according to any one of claims 10 to 11 or 39 to 47, wherein, The conversion includes decoding the bitstream to generate pixel values of the video.
51. A device for processing video data, the device comprising a processor and a non-transitory memory storing instructions thereon, wherein, The instructions, when run by the processor, cause the processor to: For the conversion between a first video unit of a video including a luminance block and two chrominance blocks and the bitstream of the video, determine to apply prediction modes to the luminance block and the two chrominance blocks respectively, wherein, in the prediction mode, the reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) escape samples, or 3) palette information included in the bitstream; Constructing a first palette for the luminance block based on a first palette prediction table, wherein the first palette is used to derive the reconstructed samples of the luminance block; Constructing a second palette for the two chrominance blocks based on a second palette prediction table, wherein the second palette is used to derive the reconstructed samples of the two chrominance blocks; and Performing the conversion based on the first palette and the second palette, wherein the luminance block and the two chrominance blocks have a tree type of a dual-tree; wherein an indication of the size of the first palette and an indication of the size of the second palette are signaled explicitly respectively; and wherein the sizes of the first palette prediction table and the second palette prediction table are signaled implicitly.
52. The apparatus according to claim 51, wherein, The maximum sizes of the first palette prediction table and the second palette prediction table are fixed integer values.
53. The apparatus according to claim 51, wherein, The sizes of the first palette prediction table and the first palette change for different luminance blocks, and the sizes of the second palette prediction table and the second palette change for different chrominance blocks.
54. The apparatus according to claim 51, wherein, The sizes of the first palette and the second palette are different.
55. The apparatus according to claim 51, wherein, Syntax elements for deriving the first palette and the second palette are respectively included in the bitstream.
56. The apparatus according to claim 51, wherein, For a video block of a tree type with a single tree and a video block of a tree type with a double tree, the first palette prediction table or the second palette prediction table has different maximum sizes.
57. The apparatus according to claim 51, wherein, The first palette prediction table is updated based on the first palette, and The second palette prediction table is updated based on the second palette.
58. The device according to claim 51, wherein, The transformation includes encoding the first video unit into the bitstream.
59. The apparatus according to claim 51, wherein, The transformation includes decoding the first video unit from the bitstream.
60. A non-transitory computer-readable storage medium having instructions stored thereon that cause a processor to: For the conversion between a first video unit of a video including a luminance block and two chrominance blocks and a bitstream of the video, determine to apply a prediction mode to the luminance block and the two chrominance blocks respectively, wherein, In the prediction mode, reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) an escape sample, or 3) palette information included in the bitstream; Construct a first palette for the luminance block based on a first palette prediction table, where the first palette is used to derive the reconstructed samples of the luminance block; Construct a second palette for the two chrominance blocks based on a second palette prediction table, where the second palette is used to derive the reconstructed samples of the two chrominance blocks; and Perform the transformation based on the first palette and the second palette, where the luminance block and the two chrominance blocks have a tree type with a double tree; where an indication of the size of the first palette and an indication of the size of the second palette are respectively signaled explicitly; and where the sizes of the first palette prediction table and the second palette prediction table are signaled implicitly.
61. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein, The method includes: For a first video unit of a video including a luminance block and two chrominance blocks, determine to apply a prediction mode to the luminance block and the two chrominance blocks respectively, where in the prediction mode, reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) an escape sample, or 3) palette information included in the bitstream; Construct a first palette for the luminance block based on a first palette prediction table, where the first palette is used to derive the reconstructed samples of the luminance block; Construct a second palette for the two chrominance blocks based on a second palette prediction table, where the second palette is used to derive the reconstructed samples of the two chrominance blocks; and Generate the bitstream based on the first palette and the second palette, where the luminance block and the two chrominance blocks have a tree type with a double tree; wherein an indication of the size of the first palette and an indication of the size of the second palette are signaled explicitly respectively; and wherein the sizes of the first palette prediction table and the second palette prediction table are signaled implicitly.
62. A method for storing a bitstream of a video, comprising: For a first video unit of a video including a luminance block and two chrominance blocks, determining to apply a prediction mode to the luminance block and the two chrominance blocks respectively, wherein in the prediction mode, reconstructed samples are represented by a set of representative color values, and the set of representative color values includes at least one of the following: 1) a palette predictor, 2) an escape sample, or 3) palette information included in the bitstream; Constructing a first palette for the luminance block based on a first palette prediction table, wherein the first palette is used to derive the reconstructed samples of the luminance block; Constructing a second palette for the luminance block based on a second palette prediction table, wherein the second palette is used to derive the reconstructed samples of the two chrominance blocks; and Generating the bitstream based on the first palette and the second palette; and Storing the bitstream in a non-transitory computer-readable recording medium, wherein the luminance block and the two chrominance blocks have a tree type of a double tree; wherein an indication of the size of the first palette and an indication of the size of the second palette are signaled explicitly respectively; and wherein the sizes of the first palette prediction table and the second palette prediction table are signaled implicitly.
63. A video processing apparatus, comprising a processor configured to implement the method according to any one of claims 10 to 50.
64. A non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the processor is caused to implement the method according to any one of claims 10 to 50.
Citation Information
Patent Citations
Palette coding for non-4:4:4 screen content video
CN107211147A
Methods and systems for palette table coding
US20160234498A1