Palette mode for local dual trees
By adopting the palette mode and local dual-tree structure in video encoding and decoding, using a palette of representative color values to represent video block samples, and combining adaptive color transformation and transform skip residual codec tools, the problem of low video encoding and decoding efficiency is solved, the bandwidth requirement is reduced, and efficient video data transmission is achieved.
Patent Information
- Application Number
- CN202180013088.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-21
- Filing Date
- 2021-02-05
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Existing video codec technologies suffer from low efficiency and high bandwidth requirements when processing videos, especially in palette mode. This is especially true in the Internet and digital communication networks, where bandwidth requirements continue to grow as the number of user devices increases.
The palette mode is adopted for video processing, and a palette of representative color values is used to represent the samples of the video block. The encoding and decoding process is optimized through the local dual-tree structure and adaptive color transformation mode. The transform skip residual encoding and decoding tool and the context encoding and decoding process are combined to optimize the conversion between video blocks and bitstreams.
It improves video encoding and decoding efficiency, reduces bandwidth requirements, adapts to different video formats and conditions, and achieves more efficient video data transmission.
Smart Images

Figure CN115176460B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of International Patent Application No. PCT / CN2020 / 074316, filed on February 5, 2020, and International Patent Application No. PCT / CN2020 / 091661, filed on May 21, 2020, under the applicable Patent Law and / or Paris Convention. The entire disclosures of the foregoing applications are incorporated herein by reference for all legal purposes and are made a part of the disclosure of this application. Technical Field
[0003] This patent document relates to encoding and decoding technology for images and videos. Background Art
[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for video processing using a palette mode, in which a palette of representative sample values is used for representation of video.
[0006] In one example aspect, a video processing method is disclosed. The method includes converting between a video block of a video and a bitstream of the video using a palette mode, wherein samples of the video block are represented using a palette of representative color values. A size of the palette for the video block is determined based on whether a local dual tree is applied to the video block.
[0007] In another example aspect, a video processing method is disclosed. The method includes converting between a video block of a video and a bitstream of the video using a palette mode in which samples of the video block are represented using a palette of representative color values. A size of a palette predictor for the video block is based on whether a local dual tree is applied to the video block.
[0008] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video using a palette mode in which samples of the video block are represented using a palette of representative color values. The bitstream conforms to a rule that specifies encoding and decoding values of escaped samples in the bitstream using a quantization parameter that is constrained by at least a maximum allowed value or a minimum allowed value.
[0009] In another example aspect, a video processing method is disclosed. The method includes performing conversion between blocks of a video and a bitstream of the video according to a format rule, the format rule specifying whether the presence of parameters associated with a chroma coding tool in an adaptation parameter set of the bitstream is based on a control flag in the adaptation parameter set.
[0010] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a block of video and a bitstream of the video. The bitstream conforms to a format rule that specifies omitting syntax elements associated with quantization parameters in picture headers of the bitstream when the video is monochrome or color components of the video are processed separately.
[0011] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video according to a rule, the rule specifying that, for the conversion, a palette mode that uses a palette of representative color values to represent samples of the video block and an adaptive color transform mode that performs color space conversion in a residual domain are mutually exclusively enabled.
[0012] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video, wherein an adaptive color transform mode that performs color space conversion in a residual domain is applied to a residual block of the video block, regardless of the color space of the residual block.
[0013] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video. The video block is encoded and decoded using a transform skip residual codec tool, wherein residual coefficients of a transform skip codec of the video block are encoded and decoded using a context codec process or a bypass codec process. During the conversion, at the beginning or end of the bypass codec process, an operation is applied to a variable specifying a number of remaining context coded bins allowed in the video block.
[0014] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video using a transform skip residual codec process. During the conversion, an operation is applied to a variable indicating whether a syntax element belongs to a particular scan phase.
[0015] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video. During the conversion, whether a syntax element indicating a sign at a coefficient level is encoded using a bypass codec process or a context codec process is indexed based on a scan phase in which identical syntax elements for one or more coefficients in a region of the video block are encoded in order.
[0016] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video block of a video and a bitstream of the video using a transform skip residual codec process. During the conversion, whether a syntax element indicating a sign of a coefficient level is coded using a bypass codec process or a context codec process is based on whether the syntax element is signaled in the same scan phase as another syntax element.
[0017] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video block of a video and a codec representation of the video, wherein a palette mode is used for the codec representation of the video block, wherein samples of the video block are represented using a palette of representative color values; and wherein samples outside the palette are encoded and decoded using an escape character and a value quantized using a quantization parameter, the quantization parameter being within a range between a minimum allowed value and a maximum allowed value determined by a rule.
[0018] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video block of a video and a codec representation of the video, wherein a palette mode is used for the codec representation of the video block, wherein samples of the video block are represented using a palette of representative color values; and wherein a size of the palette depends on a rule regarding whether a local dual tree is used for conversion between the video block and the codec representation.
[0019] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video block of a video and a codec representation of the video, wherein a palette mode is used for the codec representation of the video block, wherein samples of the video block are represented using a palette of representative color values; and wherein a size of a palette predictor depends on a rule regarding whether a local dual tree is used for conversion between the video block and the codec representation.
[0020] In another example aspect, a video processing method is disclosed. The method includes, for converting between a video block of a video region of a video and a codec representation of the video, determining, based on a codec condition, whether a syntax element identifying a deblocking offset for a chroma component of the video is included in the codec representation at the video region level; and performing the conversion based on the determination; wherein the deblocking offset is used to selectively enable a deblocking operation on the video block.
[0021] In another example aspect, a video processing method is disclosed. The method includes, for converting between a video block of a video region of a video and a codec representation of the video, determining, based on a codec condition, whether a syntax element identifying use of a chroma codec tool is included in the codec representation at the video region level; and performing the conversion based on the determination; wherein a deblocking offset is used to selectively enable a deblocking operation on the video block.
[0022] In another example aspect, a video processing method is disclosed. The method includes performing conversion between video blocks of a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format; wherein the format specifies whether a first flag indicating a deblocking offset for a chroma component of the video is included in the codec representation based on whether a second flag indicating a quantization parameter offset for the chroma component is included in the codec representation.
[0023] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video block of a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule; wherein the format rule specifies whether syntax elements in the codec representation control one or more parameters indicating applicability of one or more chroma codec tools to be included in the codec representation at the video region or video block level.
[0024] In yet another exemplary aspect, a video encoder apparatus is disclosed, wherein the video encoder includes a processor configured to implement the above method.
[0025] In yet another exemplary aspect, a video decoder apparatus is disclosed, wherein the video decoder includes a processor configured to implement the above method.
[0026] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code is in the form of processor-executable code for implementing one of the methods described herein.
[0027] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 An example of a block encoded and decoded in palette mode is shown.
[0029] Figure 2 The use of a palette predictor to signal palette entries is illustrated.
[0030] Figure 3 Examples of horizontal and vertical traversal scans are shown.
[0031] Figure 4 Illustrate example encoding and decoding of palette indexes.
[0032] Figure 5A-5B An example of a minimum chroma inter prediction unit (SCIPU) is shown.
[0033] Figure 6 It is a schematic diagram of the decoding process using ACT.
[0034] Figure 7 is a block diagram of an example video processing system.
[0035] Figure 8 It is a block diagram of a video processing device.
[0036] Figure 9 is a flow chart of an example method of video processing.
[0037] Figure 10 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0038] Figure 11 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0039] Figure 12 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0040] Figure 13 is a flowchart representation of a video processing method according to the present technology.
[0041] Figure 14 is a flowchart representation of another video processing method according to the present technology.
[0042] Figure 15 is a flowchart representation of another video processing method according to the present technology.
[0043] Figure 16 is a flowchart representation of another video processing method according to the present technology.
[0044] Figure 17 is a flowchart representation of another video processing method according to the present technology.
[0045] Figure 18 is a flowchart representation of another video processing method according to the present technology.
[0046] Figure 19 is a flowchart representation of another video processing method according to the present technology.
[0047] Figure 20 is a flowchart representation of another video processing method according to the present technology.
[0048] Figure 21 is a flowchart representation of another video processing method according to the present technology.
[0049] Figure 22 is a flowchart representation of another video processing method according to the present technology.
[0050] Figure 23 is a flowchart representation of yet another video processing method according to the present technology. DETAILED DESCRIPTION
[0051] The use of section headers in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0052] 1. Overview
[0053] This document relates to video codec technology. Specifically, it covers index and escape symbol coding in palette coding, chroma format signaling, and residual coding. It can be applied to existing video codec standards, such as HEVC, or to the upcoming standard (Versatile Video Codec). It can also be applied to future video codec standards or video codecs.
[0054] 2. Video codec standards
[0055] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methodologies and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, VCEG (Q6 / 16) and ISO / IEC JTC1SC29 / WG11 (MPEG) established the Joint Video Experts Team (JVET), which is working on the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC.
[0056] 2.1 Palette Mode in HEVC Screen Content Codec Extension (HEVC-SCC)
[0057] 2.1.1 Concept of Palette Mode
[0058] The basic idea behind palette mode is that pixels in a CU are represented by a small set of representative color values. This set is called the palette. Samples outside the palette can also be indicated by signaling an escape character followed by a (possibly quantized) component value. Such pixels are called escape pixels. Figure 1 As shown. Figure 1 As shown, for each pixel having three color components (luminance and two chrominance components), an index into a palette is established, and a block can be reconstructed based on the values established in the palette.
[0059] 2.1.2 Encoding and decoding of palette entries
[0060] For encoding and decoding of palette entries, a palette predictor is maintained. The maximum size of the palette and palette predictor is signaled in the SPS. In HEVC-SCC, palette_predictor_initializer_present_flag is introduced in the PPS. When this flag is 1, an entry for initializing the palette predictor is signaled in the bitstream. The palette predictor is initialized at the beginning of each CTU row, each slice, and each slice. Depending on the value of palette_predictor_initializer_present_flag, the palette predictor is reset to 0 or initialized using the palette predictor initializer entry signaled in the PPS. In HEVC-SCC, a palette predictor initializer of size 0 is enabled to allow palette predictor initialization to be explicitly disabled at the PPS level.
[0061] For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. Figure 2 The reuse flag is conveyed using run-length encoding of zero. Subsequently, the number of new palette entries is signaled using an exponential Golomb (EG) code of order 0, i.e., EG-0. Finally, the component values of the new palette entries are signaled.
[0062] 2.1.3 Encoding and decoding of palette index
[0063] like Figure 3 As shown, the palette index is encoded using horizontal and vertical traversal scans. The scan order is explicitly signaled in the bitstream using palette_transpose_flag. For the rest of this subsection, it is assumed that the scan is horizontal.
[0064] The palette index is encoded and decoded using two palette sampling modes: "COPY_LEFT" and "COPY_ABOVE". In "COPY_LEFT" mode, the palette index is assigned to the decoded index. In "COPY_ABOVE" mode, the palette index of the sample in the previous row is copied. For both "COPY_LEFT" and "COPY_ABOVE" modes, a run length value is signaled that specifies the number of subsequent samples that are also encoded and decoded using the same mode.
[0065] In palette mode, the index value of the escape symbol is the number of palette entries. Also, when the escape symbol is part of a downstream in "COPY_LEFT" or "COPY_ABOVE" mode, the escape component value is signaled for each escape symbol. The palette index is encoded and decoded as follows: Figure 4 shown.
[0066] This syntax order is done as follows. First, the number of index values for the CU is signaled. Next, the actual index values for the entire CU are signaled using truncated binary encoding. Both the number of indices and the index values are encoded and decoded in bypass mode. This groups the bypass binaries associated with the indices together. Then, the palette sample mode (if necessary) and the run length are signaled in an interleaved manner. Finally, the component escape values corresponding to the escape samples for the entire CU are grouped together and encoded and decoded in bypass mode. The binarization of the escape samples is EG encoding with three orders, i.e. EG-3.
[0067] An additional syntax element, last_run_type_flag, is signaled after the index value is signaled. This syntax element, combined with the number of indices, eliminates the need to signal the run value corresponding to the last run in the block.
[0068] In HEVC-SCC, palette modes also support 4:2:2, 4:2:0 and monochrome chroma formats. The signaling of palette entries and palette indices is almost the same for all chroma formats. In the case of non-monochrome formats, each palette entry consists of 3 components. For monochrome formats, each palette entry consists of one component. For subsampled chroma direction, chroma samples are associated with luma sample indices that are divisible by 2. After reconstructing the palette index for a CU, if a sample has only one component associated with it, only the first component of the palette entry is used. The only difference in the signaling is the escape component values. For each escape symbol, the number of signaled escape component values may be different depending on the number of components associated with that symbol.
[0069] Figure 4 Illustrate example encoding and decoding of palette indexes.
[0070] In addition, the palette index codec includes an index adjustment process. When signaling a palette index, the left-neighboring or top-neighboring index should be different from the current index. Therefore, by removing one possibility, the range of the current palette index can be reduced by 1. The index is then signaled using truncated binary (TB) binarization.
[0071] The text related to this section is as follows, where CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the predicted index.
[0072] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is an index into the array represented by CurrentPaletteEntries. The array index xC, yC specifies the position of the sample relative to the top-left luma sample of the picture (xC, yC). The value of PaletteIndexMap[xC][yC] should be in the range of 0 to MaxPaletteIndex, inclusive.
[0073] The variable adjustedRefPaletteIndex is derived as follows:
[0074]
[0075] When CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows:
[0076] if(CurrPaletteIndex>=adjustedRefPaletteIndex)
[0077] CurrPaletteIndex++
[0078] In addition, the run length element in the palette mode is context-encoded. The relevant context derivation process described in JVET-O2011-vE is as follows.
[0079] Derivation of ctxInc for the syntax element palette_run_prefix
[0080] The input to this process is the binary index binIdx and the syntax elements copy_above_palette_indexes_flag and palette_idx_idc.
[0081] The output of this process is the variable ctxInc.
[0082] The variable ctxInc is derived as follows:
[0083] – If copy_above_palette_indexes_flag is equal to 0 and binIdx is equal to 0,
[0084] Then the derivation of ctxInc is as follows:
[0085] ctxInc=(palette_idx_idc<1)? 0:((palette_idx_idc<3)?1:2)(9-69)
[0086] – Otherwise, ctxInc is provided by Table 1:
[0087] Table 1 – Specification of ctxIdxMap[copy_above_palette_indices_flag][binIdx]
[0088]
[0089] 2.2 Palette Mode in VCC
[0090] 2.2.1 Palette in Dual Tree
[0091] In VVC, a dual-tree codec structure is used to encode and decode intra-frame slices, so the luma component and the two chroma components can have different palettes and palette indices. In addition, the two chroma components share the same palette and palette index.
[0092] 2.2.2 Palette as a Separate Mode
[0093] In some embodiments, the prediction modes for the codec unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. The binarization of the prediction mode changes accordingly.
[0094] When IBC is disabled, for I slices, the first binary number is used to indicate whether the current prediction mode is MODE_PLT. For P / B slices, the first binary number is used to indicate whether the current prediction mode is MODE_INTRA. If not, an additional binary number is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTER.
[0095] When IBC is enabled, for I slices, the first binary number is used to indicate whether the current prediction mode is MODE_IBC. If not, the second binary number is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. For P / B slices, the first binary number is used to indicate whether the current prediction mode is MODE_INTRA. If the current prediction mode is intra mode, the second binary number is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. If not, the second binary number is used to indicate whether the current prediction mode is MODE_IBC or MODE_INTER.
[0096] The sample text is shown below.
[0097]
[0098]
[0099] 2.2.3 Palette Mode Syntax
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107] 2.2.4 Palette Mode Semantics
[0108] In the following semantics, array indices x0 and y0 specify the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture. Array indices xC and yC specify the position (xC, yC) of that sample relative to the top-left luma sample of the picture. Array index startComp specifies the first color component of the current palette table. startComp equal to 0 specifies the Y component; startComp equal to 1 specifies the Cb component; and startComp equal to 2 specifies the Cr component. numComps specifies the number of color components in the current palette table.
[0109] The predictor palette includes palette entries from the previous codec unit that are used to predict entries in the current palette.
[0110] The variable PredictorPaletteSize[startComp] specifies the size of the predictor palette for the first color component startComp of the current palette table. PredictorPaletteSize is derived as specified in clause 8.4.5.3.
[0111] The variable PalettePredictorEntryReuseFlags[i] equal to 1 specifies that the i-th entry in the predictor palette is reused in the current palette. PalettePredictorEntryReuseFlags[i] equal to 0 specifies that the i-th entry in the predictor palette is not an entry in the current palette. All elements of the array PalettePredictorEntryReuseFlags[i] are initialized to 0.
[0112] palette_predictor_run is used to determine the number of zeros preceding the non-zero entries in the array PalettePredictorEntryReuseFlags.
[0113] A bitstream conformance requirement is that the value of palette_predictor_run shall be in the range of 0 to (PredictorPaletteSize - predictorEntryIdx), inclusive, where predictorEntryIdx corresponds to the current position in the array PalettePredictorEntryReuseFlags. The variable NumPredictedPaletteEntries specifies the number of entries in the current palette that are reused from the predictor palette. The value of NumPredictedPaletteEntries shall be in the range of 0 to palette_max_size, inclusive.
[0114] num_signalled_palette_entries specifies the number of entries in the current palette that are explicitly signaled for the first color component startComp of the current palette table.
[0115] When num_signalled_palette_entries is not present, it is inferred to be equal to 0.
[0116] The variable CurrentPaletteSize[startComp] specifies the size of the current palette for the first color component startComp of the current palette table and is derived as follows:
[0117] CurrentPaletteSize[startComp]=NumPredictedPaletteEntries+num_signalled_palette_entries(7-155)
[0118] The value of CurrentPaletteSize[startComp] should be in the range of 0 to palette_max_size, inclusive.
[0119] new_palette_entries[cIdx][i] specifies the value of the i-th signaled palette entry for color component cIdx.
[0120] The variable PredictorPaletteEntries[cIdx][i] specifies the i-th element in the predictor palette for color component cIdx.
[0121] The variable CurrentPaletteEntries[cIdx][i] specifies the i-th element in the current palette for color component cIdx and is derived as follows:
[0122]
[0123] palette_escape_val_present_flag equal to 1 specifies that the current codec contains at least one escape codec sample. escape_val_present_flag equal to 0 specifies that there are no escape codec samples in the current codec. When not present, the value of palette_escape_val_present_flag is inferred to be equal to 1.
[0124] The variable MaxPaletteIndex specifies the maximum possible value of the palette index of the current codec unit. The value of MaxPaletteIndex is set equal to CurrentPaletteSize[startComp]-1+palette_escape_val_present_flag.
[0125] num_palette_indices_minus1 plus 1 is the number of palette indices explicitly signaled or inferred for the current block.
[0126] When num_palette_indices_minus1 is not present, it is inferred to be equal to 0.
[0127] palette_idx_idc is an indication of the index CurrentPaletteEntries into the palette table. For the first index in the block, the value of palette_idx_idc shall be in the range of 0 to MaxPaletteIndex, inclusive, and for the remaining indices in the block, the value of palette_idx_idc shall be in the range of 0 to MaxPaletteIndex-1, inclusive.
[0128] When palette_idx_idc is not present, it is inferred to be equal to 0.
[0129] The variable PaletteIndexIdc[i] stores the i-th palette_idx_idc, either explicitly signaled or inferred. All elements of the array PaletteIndexIdc[i] are initialized to 0.
[0130] copy_above_indices_for_final_run_flag equal to 1 specifies that the palette index of the last position in the codec unit is copied from the palette index in the previous row if horizontal traversal scanning is used, or from the palette index in the left column if vertical traversal scanning is used. copy_above_indices_for_final_run_flag equal to 0 specifies that the palette index of the last position in the codec unit is copied from PaletteIndexIdc[num_palette_indices_minus1].
[0131] When copy_above_indices_for_final_run_flag is not present, it is inferred to be equal to 0.
[0132] palette_transpose_flag equal to 1 specifies that vertical traversal scanning should be applied to scan the indices of the samples in the current codec unit. palette_transpose_flag equal to 0 specifies that horizontal traversal scanning should be applied to scan the indices of the samples in the current codec unit. When not present, the value of palette_transpose_flag is inferred to be equal to 0.
[0133] The array TraverseScanOrder specifies the scan order array for palette encoding and decoding. If palette_transpose_flag is equal to 0, TraverseScanOrder is assigned the horizontal scan order HorTravScanOrder, and if palette_transpose_flag is equal to 1, TraverseScanOrder is assigned the vertical scan order VerTravScanOrder.
[0134] copy_above_palette_indices_flag equal to 1 specifies that the palette index is equal to the palette index at the same position in the previous row if horizontal traversal scan is used, or the palette index is equal to the palette index at the same position in the left column if vertical traversal scan is used. copy_above_palette_indices_flag equal to 0 specifies that the indication of the palette index of the sample is encoded or inferred in the bitstream.
[0135] The variable CopyAboveIndicesFlag[xC][yC] equal to 1 specifies that the palette index is copied from the palette index in the previous row (horizontally scanned) or left column (vertically scanned). CopyAboveIndicesFlag[xC][yC] equal to 0 specifies that the palette index is explicitly encoded or inferred in the bitstream. The array indices xC, yC specify the position of the sample relative to the top-left luma sample of the picture (xC, yC). The value of PaletteIndexMap[xC][yC] shall be in the range of 0 to (MaxPaletteIndex–1), inclusive.
[0136] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is an index into the array represented by CurrentPaletteEntries. The array index xC, yC specifies the position of the sample relative to the top-left luma sample of the picture (xC, yC). The value of PaletteIndexMap[xC][yC] should be in the range of 0 to MaxPaletteIndex, inclusive.
[0137] The variable adjustedRefPaletteIndex is derived as follows:
[0138]
[0139]
[0140] When CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows:
[0141] if(CurrPaletteIndex>=adjustedRefPaletteIndex)
[0142] CurrPaletteIndex++ (7-158)
[0143] palette_run_prefix, when present, specifies the prefix portion of the binarization of PaletteRunMinus1.
[0144] palette_run_suffix is used in the derivation of the variable PaletteRunMinus1. When not present, the value of palette_run_suffix is inferred to be equal to 0.
[0145] When RunToEnd is equal to 0, the variable PaletteRunMinus1 is derived as follows:
[0146] – If PaletteMaxRunMinus1 is equal to 0, PaletteRunMinus1 is set equal to 0.
[0147] – Otherwise (PaletteMaxRunMinus1 is greater than 0), the following applies:
[0148] – If palette_run_prefix is less than 2, the following applies:
[0149] PaletteRunMinus1=palette_run_prefix (7-159)
[0150] – Otherwise (palette_run_prefix is greater than or equal to 2), the following applies:
[0151] PrefixOffset=1<<(palette_run_prefix-1)
[0152] PaletteRunMinus1=PrefixOffset+palette_run_suffix (7-160)
[0153] The variable PaletteRunMinus1 is used as follows:
[0154] – If CopyAboveIndicesFlag[xC][yC] is equal to 0, PaletteRunMinus1 specifies the number of consecutive locations with the same palette index minus 1.
[0155] – Otherwise, if palette_transpose_flag is equal to 0, then PaletteRunMinus1 specifies the number of consecutive positions that have the same palette index as used in the corresponding position in the previous row, minus 1.
[0156] – Otherwise, PaletteRunMinus1 specifies the number of consecutive positions whose palette index is the same as the palette index used in the corresponding position in the left column, minus one.
[0157] When RunToEnd is equal to 0, the variable PaletteMaxRunMinus1 represents the maximum possible value of PaletteMaxRunMinus1, and the requirement for bitstream conformance is that the value of PaletteMaxRunMinus1 should be greater than or equal to 0.
[0158] palette_escape_val specifies the quantization escape codec sample value for the component.
[0159] The variable PaletteEscapeVal[cIdx][xC][yC] specifies the escape value for the sample, where PaletteIndexMap[xC][yC] is equal to MaxPaletteIndex and palette_escape_val_present_flag is equal to 1. Array index cIdx specifies the color component. Array indices xC and yC specify the position (xC, yC) of the sample relative to the top-left luma sample of the picture.
[0160] The bitstream conformance requirement is that for cIdx equal to 0, PaletteEscapeVal[cIdx][xC][yC] shall be between 0 and (1<<(BitDepth Y +1))-1, including the end value, if cIdx is not equal to 0, then it is in the range of 0 to (1<<(BitDepth C +1))-1, inclusive.
[0161] 2.2.5 Line-based CG palette mode
[0162] VVC adopts a line-based CG palette mode. In this method, each CU of the palette mode is divided into multiple segments consisting of m samples (m=16 in this test) according to the traversal scan mode. The encoding order of the palette run-length encoding and decoding in each segment is as follows: For each pixel, a context-coded binary number run_copy_flag=0 is signaled to indicate whether the pixel has the same mode as the previous pixel, that is, whether the previous scanned pixel and the current pixel are both run-length type COPY_ABOVE, or whether the previous scanned pixel and the current pixel are both run-length type INDEX and the same index value. Otherwise, run_copy_flag=1 is signaled. If the pixel and the previous pixel are different modes, a context-coded binary number copy_above_palette_indices_flag is signaled to indicate the run-length type of the pixel, that is, INDEX or COPY_ABOVE. As in palette mode in VTM6.0, if the sample is located in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to parse the run type because the index mode is used by default. In addition, if the previously parsed run type is COPY_ABOVE, the decoder does not have to parse the run type. After palette run-length encoding and decoding of pixels in a segment, the index value (for INDEX mode) and the quantized escape color are bypass-encoded and grouped separately from the encoding / parsing of the context-encoded binary number to improve the throughput within each line CG. Because the index value is now encoded / parsed after run-length encoding and decoding, instead of being processed before palette run-length encoding and decoding as in VTM, the encoder does not have to signal the number of index values num_palette_indices_minus1 and the final run type copy_above_indices_for_final_run_flag.
[0163] The text of the line-based CG palette mode in some embodiments is shown below.
[0164] Palette encoding and decoding syntax
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171] 7.4.9.6 Palette encoding and decoding semantics
[0172] In the following semantics, array indices x0 and y0 specify the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture. Array indices xC and yC specify the position (xC, yC) of that sample relative to the top-left luma sample of the picture. Array index startComp specifies the first color component of the current palette table. startComp equal to 0 specifies the Y component; startComp equal to 1 specifies the Cb component; and startComp equal to 2 specifies the Cr component. numComps specifies the number of color components in the current palette table.
[0173] The predictor palette consists of palette entries from the previous codec unit that was used to predict the entries in the current palette.
[0174] The variable PredictorPaletteSize[startComp] specifies the size of the predictor palette for the first color component of the current palette table startComp. PredictorPaletteSize is derived as specified in clause 8.4.5.3.
[0175] The variable PalettePredictorEntryReuseFlags[i] equal to 1 specifies that the i-th entry in the predictor palette is reused in the current palette. PalettePredictorEntryReuseFlags[i] equal to 0 specifies that the i-th entry in the predictor palette is not an entry in the current palette. All elements of the array PalettePredictorEntryReuseFlags[i] are initialized to 0.
[0176] palette_predictor_run is used to determine the number of zeros preceding the non-zero entries in the array PalettePredictorEntryReuseFlags.
[0177] A bitstream conformance requirement is that the value of palette_predictor_run shall be in the range of 0 to (PredictorPaletteSize - predictorEntryIdx), inclusive, where predictorEntryIdx corresponds to the current position in the array PalettePredictorEntryReuseFlags. The variable NumPredictedPaletteEntries specifies the number of entries in the current palette that are reused from the predictor palette. The value of NumPredictedPaletteEntries shall be in the range of 0 to palette_max_size, inclusive.
[0178] num_signalled_palette_entries specifies the number of entries in the current palette that are explicitly signaled for the first color component of the current palette table startComp.
[0179] When num_signalled_palette_entries is not present, it is inferred to be equal to 0.
[0180] The variable CurrentPaletteSize[startComp] specifies the size of the current palette for the first color component of the current palette table startComp and is derived as follows:
[0181] CurrentPaletteSize[startComp]=NumPredictedPaletteEntries+
[0182] num_signalled_palette_entries(7-155)
[0183] The value of CurrentPaletteSize[startComp] should be in the range of 0 to palette_max_size, inclusive.
[0184] new_palette_entries[cIdx][i] specifies the value of the i-th signaled palette entry for color component cIdx.
[0185] The variable PredictorPaletteEntries[cIdx][i] specifies the i-th element in the predictor palette for color component cIdx.
[0186] The variable CurrentPaletteEntries[cIdx][i] specifies the i-th element in the current palette for color component cIdx and is derived as follows:
[0187]
[0188]
[0189] palette_escape_val_present_flag equal to 1 specifies that the current codec contains at least one escaped coded sample. escape_val_present_flag equal to 0 specifies that there are no escaped coded samples in the current codec. When not present, the value of palette_escape_val_present_flag is inferred to be equal to 1.
[0190] The variable MaxPaletteIndex specifies the maximum possible value of the palette index of the current codec unit. The value of MaxPaletteIndex is set equal to CurrentPaletteSize[startComp]-1+palette_escape_val_present_flag.
[0191] palette_idx_idc is an indication of the index into the palette table CurrentPaletteEntries. For the first index in the block, the value of palette_idx_idc shall be in the range of 0 to MaxPaletteIndex, inclusive, and for the remaining indices in the block, the value of palette_idx_idc shall be in the range of 0 to MaxPaletteIndex-1, inclusive.
[0192] When palette_idx_idc is not present, it is inferred to be equal to 0.
[0193] palette_transpose_flag equal to 1 specifies that vertical traversal scanning should be applied to scan the indices of the samples in the current codec unit. palette_transpose_flag equal to 0 specifies that horizontal traversal scanning should be applied to scan the indices of the samples in the current codec unit. When not present, the value of palette_transpose_flag is inferred to be equal to 0.
[0194] The array TraverseScanOrder specifies the scan order array used for palette encoding and decoding. If palette_transpose_flag is equal to 0, TraverseScanOrder is assigned the horizontal scan order HorTravScanOrder, and if palette_transpose_flag is equal to 1, TraverseScanOrder is assigned the vertical scan order VerTravScanOrder.
[0195] run_copy_flag equal to 1 specifies that if copy_above_palette_indices_flag is equal to 0, the palette run type is the same as the run type of the previous scan position and the palette run index is the same as the index of the previous position. Otherwise, run_copy_flag is equal to 0
[0196] copy_above_palette_indices_flag equal to 1 specifies that the palette index is equal to the palette index at the same position in the previous row if horizontal traversal scanning is used, or the palette index is equal to the palette index at the same position in the left column if vertical traversal scanning is used. copy_above_palette_indices_flag equal to 0 specifies that the indication of the palette index of the sample is coded or inferred in the bitstream.
[0197] CopyAboveIndicesFlag[xC][yC] equal to 1 specifies that the palette index is copied from the palette index in the previous row (horizontally scanned) or left column (vertically scanned). CopyAboveIndicesFlag[xC][yC] equal to 0 specifies that the palette index is explicitly coded or inferred in the bitstream. The array indices xC, yC specify the position of the sample relative to the top-left luma sample (xC, yC) of the picture.
[0198] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is an index into the array represented by CurrentPaletteEntries. The array index xC, yC specifies the position of the sample relative to the top-left luma sample of the picture (xC, yC). The value of PaletteIndexMap[xC][yC] should be in the range of 0 to MaxPaletteIndex, inclusive.
[0199] The variable adjustedRefPaletteIndex is derived as follows:
[0200]
[0201] When CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows:
[0202] if(CurrPaletteIndex>=adjustedRefPaletteIndex)
[0203] CurrPaletteIndex++ (7-158)
[0204] palette_escape_val specifies the quantization escape codec sample value for the component.
[0205] The variable PaletteEscapeVal[cIdx][xC][yC] specifies the escape value for the sample, where PaletteIndexMap[xC][yC] is equal to MaxPaletteIndex and palette_escape_val_present_flag is equal to 1. Array index cIdx specifies the color component. Array indices xC and yC specify the position (xC, yC) of the sample relative to the top-left luma sample of the picture.
[0206] The bitstream conformance requirement is that for cIdx equal to 0, PaletteEscapeVal[cIdx][xC][yC] shall be between 0 and (1<<(BitDepth Y +1))-1, including the end value, if cIdx is not equal to 0, then it is in the range of 0 to (1<<(BitDepth C +1))-1, inclusive.
[0207] 2.3 Local Dual Tree in VVC
[0208] In typical hardware video encoders and decoders, processing throughput decreases when a picture has more small intra blocks due to data dependencies in sample processing between adjacent intra blocks. Predictor generation for an intra block requires reconstructed samples from the top and left boundaries of adjacent blocks. Therefore, intra prediction must be processed sequentially, block by block.
[0209] In HEVC, the smallest intra CU is 8×8 luma samples. The luma component of the smallest intra CU can be further divided into four 4×4 luma intra prediction units (PUs), but the chroma components of the smallest intra CU cannot be further split. Therefore, the hardware processing throughput is the worst when processing 4×4 chroma intra blocks or 4×4 luma intra blocks.
[0210] In VTM5.0, in a single codec tree, since chroma partitioning always follows luma and the minimum intra CU is 4×4 luma samples, the minimum chroma intra CB is 2×2. Therefore, in VTM5.0, the minimum chroma intra CB in a single codec tree is 2×2. The worst-case hardware processing throughput of VVC decoding is only 1 / 4 of that of HEVC decoding. In addition, after adopting tools including cross-component linear model (CCLM), 4-tap interpolation filter, position-dependent intra prediction combination (PDPC), and combined inter-intra prediction (CIIP), the reconstruction process of chroma intra CB becomes much more complex than in HEVC. Achieving high processing throughput in hardware decoders is challenging. In this section, we propose a method to improve the worst-case hardware processing throughput.
[0211] The goal of this method is to prohibit chroma intra CBs with less than 16 chroma samples by constraining the segmentation of chroma intra CBs.
[0212] In a single codec tree, a SCIPU is defined as a codec tree node whose chroma block size is greater than or equal to TH chroma samples and has at least one sub-luminance block less than 4TH luminance samples, where TH is set to 16 in this proposal. It is required that in each SCIPU, all CBs are inter-frame or all CBs are non-inter-frame, i.e., intra-frame or IBC. In the case of non-inter SCIPU, it is further required that the chroma of the non-inter SCIPU should not be further divided, and the luminance of the SCIPU is allowed to be further divided. In this way, the minimum chroma intra-frame CB size is 16 chroma samples, and 2×2, 2×4 and 4×2 chroma CBs are removed. In addition, in the case of non-inter SCIPU, chroma scaling is not applied. In addition, when the luminance block is further divided and the chroma block is not divided, a local dual-tree codec structure is constructed.
[0213] Figure 5A-5B Two SCIPU examples are shown in . Figure 5A In , one chroma CB and three luma CBs (4×8, 8×8, 4×8 luma CB) of 8×4 chroma samples form one SCIPU, because the ternary tree (TT) partitioned from 8×4 chroma samples will result in a chroma CB with less than 16 chroma samples. Figure 5B , one chroma CB of 4×4 chroma samples (on the left side of 8×4 chroma samples) and three luma CBs (8×4, 4×4, 4×4 luma CB) form one SCIPU, and another chroma CB of 4×4 samples (on the right side of 8×4 chroma samples) and two luma CBs (8×4, 8×4 luma CB) form one SCIPU, because a binary tree (BT) partitioned from 4×4 chroma samples will result in a chroma CB of less than 16 chroma samples.
[0214] In the proposed method, if the current slice is an I slice or the current SCIPU has a 4×4 luma split in it after being further divided once (because inter 4×4 is not allowed in VVC), the type of SCIPU is inferred as non-inter; otherwise, the type of SCIPU (inter or non-inter) is indicated by a signaled flag before parsing the CU in the SCIPU.
[0215] By applying the above method, the worst-case hardware processing throughput is achieved when processing 4×4, 2×8, or 8×2 chroma blocks instead of 2×2 chroma blocks. The worst-case hardware processing throughput is the same as HEVC and 4 times that of VTM5.0.
[0216] 2.4 Transform Skip (TS)
[0217] As in HEVC, the residual of a block can be encoded and decoded in transform skip mode. To avoid syntax encoding and decoding redundancy, the transform skip flag is not signaled when the CU level MTS_CU_flag is not equal to zero. The block size restrictions for transform skip are the same as those for MTS in JEM4, which indicates that transform skip applies to a CU when both the block width and height are equal to or less than 32. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, implicit MTS can still be enabled.
[0218] Furthermore, for transform skip blocks, the minimum allowed quantization parameter (QP) is defined as 6*(internalBitDepth−inputBitDepth)+4.
[0219] 2.5 Alternative Luma Half-Pixel Interpolation Filter
[0220] In some embodiments, an alternative half-pixel interpolation filter is proposed.
[0221] The switching of the half-pixel luma interpolation filter depends on the motion vector accuracy. In addition to the existing quarter-pixel, full-pixel, and 4-pixel AMVR modes, a new half-pixel precision AMVR mode has been introduced. The alternative half-pixel luma interpolation filter can only be selected when the motion vector accuracy is half-pixel.
[0222] For non-affine non-merge inter-coded CUs using half-pixel motion vector accuracy (i.e., half-pixel AMVR mode), the HEVC / VVC half-pixel luma interpolation filter and one or more alternative half-pixel interpolation filters are switched based on the value of the new syntax element hpelIfIdx. The syntax element hpelIfIdx is signaled only in half-pixel AMVR mode. In the case of skip / merge mode with spatial merging candidates, the value of the syntax element hpelIfIdx is inherited from the neighboring block.
[0223] 2.6 Adaptive Color Transformation (ACT)
[0224] Figure 6 The figure shows the decoding flow chart of the application ACT. Figure 6 As shown, the color space conversion is performed in the residual domain. Specifically, an additional decoding module, namely inverse ACT, is introduced after the inverse transform to convert the residual from the YCgCo domain back to the original domain.
[0225] In VVC, a CU leaf node is also used as the unit of transform processing unless the maximum transform size is less than the width or height of a codec unit (CU). Therefore, in the proposed implementation, the ACT flag is signaled for a CU to select the color space used for encoding and decoding its residual. In addition, following the HEVC ACT design, for inter and IBC CUs, ACT is only enabled when there is at least one non-zero coefficient in the CU. For intra CUs, ACT is only enabled when the chroma components select the same intra prediction mode (i.e., DM mode) as the luma component.
[0226] The core transform for color space conversion remains the same as that used for HEVC. Specifically, the following forward and inverse YCgCo color transform matrices are applied as described below.
[0227]
[0228] In addition, to compensate for the dynamic range change of the residual signal before and after color transformation, a QP adjustment of (-5, -5, -3) is applied to the transformed residual.
[0229] On the other hand, the forward and inverse color transforms require access to the residuals of all three components. Accordingly, in the proposed implementation, ACT is disabled in the following two cases, where not all residuals of the three components are available.
[0230] 1. Split tree partitioning: When split tree is applied, the luma and chroma samples within a CTU are partitioned by different structures. This results in the CU in the luma tree containing only luma components, while the CU in the chroma tree contains only two chroma components.
[0231] 2. Intra-frame sub-partition prediction (ISP): ISP sub-partitioning is only applied to luma, while chroma signals are encoded and decoded without splitting. In the current ISP design, except for the last ISP sub-partitioning, other sub-partitions contain only luma components.
[0232] 2.7 Binarization of Outliers Using EG(k)
[0233] When using EG(k) for binarization of escaped values, when the base Qp is large enough (or the symbols to be encoded and decoded are small enough), the bit length of EG(k) cannot be reduced any further. For example, when the base Qp>=23, the bit length reaches 6 for EG(5), which is the minimum bit length of EG5. Similarly, when the base Qp>=35, the bit length reaches the minimum value for EG3. When the base Qp>=29, the bit length reaches the minimum value for EG4. In this case, further increasing Qp does not reduce the bit rate, but increases distortion. This is a waste of bits.
[0234] 2.8 Coefficient Encoding and Decoding in Transform Skip Mode
[0235] In the current VVC draft, several modifications are proposed to the coefficient coding in transform skip (TS) mode compared to non-TS coefficient coding in order to adapt the residual coding to the statistical and signal characteristics of the transform skip stage.
[0236] In the current VVC, three scanning stages are used to encode and decode coefficients in the transform skip residual encoding process. The first scanning stage is used to encode and decode the syntax elements indicating whether the transform coefficient level is greater than 0 and other related syntax elements (e.g., sig_coeff_flag, coeff_sign_flag, and par_level_flag). The second / greater than X scanning stage is used to encode and decode the syntax elements indicating whether the transform coefficient level is greater than X (e.g., X=1, 2, 3, 4, 5). The third / remaining scanning stage is used to encode and decode the remaining syntax elements (e.g., abs_remainder and coeff_sign_flag).
[0237] In the current VVC, as shown in Table 131, whether the syntax element indicating the sign of the transform coefficient level (e.g., coeff_sign_flag) is coded in bypass mode or context coding mode depends on the syntax element indicating whether the transform is applied to the associated transform block (e.g., transform_skip_flag), the number of remaining allowed context coding bins (e.g., RemCcbs), and the syntax element indicating whether the residual_coding() syntax structure is used to parse the residual samples of the transform skipped blocks of the current slice (e.g., sh_ts_residual_coding_disabled_flag).
[0238] 7.3.10.11 Residual Codec Syntax
[0239]
[0240]
[0241]
[0242]
[0243]
[0244]
[0245]
[0246]
[0247] 2.8.1 Context Modeling and Context Index Offset Derivation for the Sign Flag coeff_sign_flag
[0248] Table 51 – Association of ctxIdx and syntax elements for each initializationType during initialization
[0249]
[0250] Table 51 – Association of ctxIdx and syntax elements for each initializationType during initialization
[0251]
[0252] Table 125 - Specification of initValue and shiftIdx of ctxInc of coeff_sign_flag
[0253]
[0254] Table 131 – Assignment of ctxInc to syntax elements with context-coded bins
[0255]
[0256]
[0257] 9.3.4.2.10 Derivation of ctxInc for the syntax element coeff_sign_flag for transform skip mode
[0258] The input of this process is the color component index cIdx, the brightness position (x0, y0) of the upper left sample point of the current transform block relative to the upper left sample point of the current picture, and the current coefficient scanning position (xC, yC). The output of this process is the variable ctxInc.
[0259] The variables leftSign and aboveSign are derived as follows:
[0260] leftSign=(xC==0)? 0:CoeffSignLevel[xC-1][yC] (1594)
[0261] aboveSign=(yC==0)? 0:CoeffSignLevel[xC][yC-1] (1595)
[0262] The variable ctxInc is derived as follows:
[0263] – If leftSign is equal to 0 and aboveSign is equal to 0, or if leftSign is equal to
[0264] -aboveSign, the following applies:
[0265] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?0:3) (1596)
[0266] – Otherwise, if leftSign is greater than or equal to 0 and aboveSign is greater than or equal to 0, then the following applies:
[0267] ctxInc=(BdpcmFlag[x0][y0][cIdx]?1:4) (1597)
[0268] – Otherwise, the following applies:
[0269] ctxInc=(BdpcmFlag[x0][y0][cIdx]?2:5) (1598)
[0270] 3. Example technical problems solved by the technical solutions described in this article
[0271] 1. When Qp is greater than a threshold, EG(k), which is a binarization method of outliers, may waste bits.
[0272] 2. The palette size may be too large for the local double tree.
[0273] 3. When the chroma tool is not applied, no chroma parameters need to be signaled.
[0274] 4. Although the coefficient codec in JVET-R2001-vA can achieve codec advantages over screen content codec, the coefficient codec and TS mode may still have some disadvantages:
[0275] a. For the following cases, it is unclear whether to use bypass codec or context codec for the symbol flag:
[0276] i. The number of remaining bins for which context coding and decoding is allowed (denoted by RemCcbs) is equal to 0.
[0277] ii. The current block is encoded and decoded in TS mode.
[0278] iii.sh_ts_residual_coding_disabled_flag is no.
[0279] 4. Example Embodiments and Techniques
[0280] The following list of items should be considered as examples to explain general concepts. These items should not be interpreted in a narrow sense.
[0281] Furthermore, these items can be combined in any way.
[0282] The following examples can be applied to the palette scheme in VVC and all other palette-related schemes.
[0283] 1. Qp used for outlier reconstruction may have a maximum and / or minimum allowed value.
[0284] a. In one example, the QP may be clipped to be no greater than a maximum allowed value and / or no less than a minimum allowed value.
[0285] b. In one example, the maximum allowed Qp for outlier reconstruction may depend on the binarization method.
[0286] c. In one example, the maximum allowed Qp for outlier reconstruction may be (T+B), where B is based on the bit depth.
[0287] i. In one example, T can be a constant.
[0288] 1. In one example, T can be 23.
[0289] 2. In one example, T may be a number less than 23.
[0290] 3. In one example, T may be 35.
[0291] 4. In one example, T may be a number less than 35.
[0292] 5. In one example, T may be 29.
[0293] 6. In one example, T may be a number less than 29.
[0294] ii. In one example, T may be indicated in a video region (eg, sequence, picture, slice / slice / sub-picture).
[0295] 1. In one example, T may be indicated in VPS / SPS / PPS / PH / SH.
[0296] iii. In one example, B can be set to QpBdOffset (e.g., ).
[0297] d. In one example, the maximum allowed Qp for outlier reconstruction may be (23 + QpBdOffset).
[0298] i. Additionally, alternatively, EG5 is used to encode and decode escaped values.
[0299] ii. Alternatively, the maximum allowed Qp for outlier reconstruction may be (K+QpBdOffset), where K is a number smaller than 23.
[0300] e. In one example, the maximum allowed Qp for outlier reconstruction may be (35 + QpBdOffset).
[0301] i. Additionally, alternatively, EG3 is used to encode and decode escaped values.
[0302] ii. Alternatively, the maximum allowed Qp for outlier reconstruction may be (K+QpBdOffset), where K is a number less than 35.
[0303] f. Alternatively, the maximum allowed Qp for outlier reconstruction may be (29 + QpBdOffset).
[0304] i. Additionally, alternatively, EG4 is used to encode and decode escaped values.
[0305] ii. Alternatively, the maximum allowed Qp for outlier reconstruction may be (K+QpBdOffset), where K is a number smaller than 29.
[0306] Palette size related
[0307] 2. It is proposed that the palette size can be different when applying or not applying the local double tree.
[0308] a. In one example, for a local dual tree, it is proposed that the palette size can be reduced.
[0309] b. In one example, when local dual-tree is applied, the palette sizes of luma CU and chroma CU can be different.
[0310] c. In one example, the palette size of the chroma CU can be reduced compared to the palette size of the luma CU in the local dual tree, or compared to the palette size when the local dual tree is not applied.
[0311] i. In one example, the palette size for chroma can be halved.
[0312] 3. It is proposed that the size of the palette predictor can be different when applying or not applying the local dual tree.
[0313] a. In one example, for a local dual tree, it is proposed that the size of the palette predictor can be reduced.
[0314] b. In one example, when local dual-tree is applied, the palette predictor sizes of luma CU and chroma CU may be different.
[0315] c. In one example, the palette predictor size of the chroma CU can be reduced compared to the palette predictor size of the luma CU in the local dual tree, or compared to the palette predictor size when the local dual tree is not applied.
[0316] i. In one example, the size of the palette predictor for chroma can be halved.
[0317] Chroma Deblocking
[0318] 4. Whether the chroma deblocking offset is signaled / parsed at the slice level and / or higher level (i.e., region size is larger than the slice) (e.g., in the PPS or picture header) may depend on the color format and / or the separate plane codec enable flag and / or the ChromaArrayType and / or a flag indicating whether the chroma deblocking offset is present and / or a flag indicating whether the chroma deblocking offset or some other chroma tool parameter is present.
[0319] a. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane codec is applied or a flag indicates that chroma deblocking offset is not present, the signaling / parsing of chroma deblocking offset at slice level and / or higher level (i.e., region size is larger than slice) can always be skipped.
[0320] b. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied or the flag indicates that chroma deblocking offset does not exist, the signaling / parsing of pps_cb_beta_offset_div2, pps_cb_tc_offset_div2, pps_cr_beta_offset_div2, pps_cr_tc_offset_div2 can always be skipped.
[0321] c. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied or the flag indicates that chroma deblocking offset does not exist, the signaling / parsing of ph_cb_beta_offset_div2, ph_cb_tc_offset_div2, ph_cr_beta_offset_div2, ph_cr_tc_offset_div2 can always be skipped.
[0322] d. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied or the flag indicates that chroma deblocking offset does not exist, the signaling / parsing of slice_cb_beta_offset_div2, slice_cb_tc_offset_div2, slice_cr_beta_offset_div2, slice_cr_tc_offset_div2 can always be skipped.
[0323] e. Alternatively, the conforming bitstream shall satisfy that when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied, pps_cb_beta_offset_div2, pps_cb_tc_offset_div2, pps_cr_beta_offset_div2, pps_cr_tc_offset_div2 shall be equal to 0.
[0324] f. In one example, when chroma_format_idc is equal to 0 and separate_colour_plane_flag is not equal to 1 or the flag indicates that chroma deblocking offset is not present, signaling / parsing of pps_cb_beta_offset_div2, pps_cb_tc_offset_div2, pps_cr_beta_offset_div2, pps_cr_tc_offset_div2 may always be skipped.
[0325] g. In one example, when chroma_format_idc is equal to 0 and separate_colour_plane_flag is not equal to 1 or the flag indicates that chroma deblocking offset is not present, signaling / parsing of pps_cb_beta_offset_div2, ph_cb_tc_offset_div2, ph_cr_beta_offset_div2, ph_cr_tc_offset_div2 may always be skipped.
[0326] h. In one example, when chroma_format_idc is equal to 0 and separate_colour_plane_flag is not equal to 1 or the flag indicates that chroma deblocking offset is not present, signaling / parsing of slice_cb_beta_offset_div2, slice_cb_tc_offset_div2, slice_cr_beta_offset_div2, slice_cr_tc_offset_div2 may always be skipped.
[0327] i. Alternatively, when the signaling of a syntax element is skipped, the value of the syntax element is inferred to be equal to 0.
[0328] 5. The color format and / or separate plane codec enable flag and / or ChromaArrayType and / or a flag indicating whether chroma deblocking offset exists and / or a flag indicating whether chroma deblocking offset or some other chroma tool parameters exist (e.g., pps_chroma_tool_params_present_flag) can be indicated in PPS and / or SPS and / or APS.
[0329] a. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 and / or the flag is false, the signaling / parsing of pps_cb_beta_offset_div2, pps_cb_tc_offset_div2, pps_cr_beta_offset_div2, pps_cr_tc_offset_div2 can always be skipped.
[0330] b. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 and / or the flag is false, the signaling / parsing of pps_cb_beta_offset_div2, pps_cb_tc_offset_div2, pps_cr_beta_offset_div2, pps_cr_tc_offset_div2 can always be skipped.
[0331] c. In one example, under the condition that ChromaArrayType is not equal to 0 and / or the flag is false, chroma tool offset related syntax elements (e.g., pps_cb_qp_offset, pps_cr_qp_offset, pps_joint_cbcr_qp_offset_present_flag, pps_slice_chroma_qp_offsets_present_flag, pps_cu_chroma_qp_offset_list_enabled_flag) are signaled.
[0332] d. In a conforming bitstream, the color format and / or separate plane codec enable flag and / or ChromaArrayType signaled in the PPS shall be the same as the corresponding information signaled in the associated SPS.
[0333] 6. Propose a flag to control whether chroma qp offset should be signaled / resolved. It is also possible to control whether chroma deblocking offset should be signaled / resolved.
[0334] a. In one example, the flag pps_chroma_tool_params_present_flag can be used to control whether chroma qp offset should be signaled / resolved, and whether chroma deblocking offset should be signaled / resolved.
[0335] 7. A control flag may be added in the PPS, such as pps_chroma_deblocking_params_present_flag, to control whether the chroma deblocking offset should be signaled / resolved.
[0336] a. In one example, when the flag is equal to 0, the signaling / parsing of pps_cb_beta_offset_div2, pps_cb_tc_offset_div2, pps_cr_beta_offset_div2, pps_cr_tc_offset_div2 may always be skipped.
[0337] b. In one example, when the flag is equal to 0, the signaling / parsing of ph_cb_beta_offset_div2, ph_cb_tc_offset_div2, ph_cr_beta_offset_div2, ph_cr_tc_offset_div2 may always be skipped.
[0338] c. In one example, when the flag is equal to 0, the signaling / parsing of slice_cb_beta_offset_div2, slice_cb_tc_offset_div2, slice_cr_beta_offset_div2, slice_cr_tc_offset_div2 may always be skipped.
[0339] d. Additionally, alternatively, in a conforming bitstream, when ChromaArrayType is equal to 0, the requirement flag should be equal to 0.
[0340] Chroma tool related parameters in APS
[0341] 8. A control flag may be added in APS, such as aps_chroma_tool_params_present_flag, to control whether chroma tool related parameters should be signaled / parsed in APS.
[0342] a. In one example, when aps_chroma_tool_params_present_flag is equal to 0, alf_chroma_filter_signal_flag, alf_cc_cb_filter_signal_flag, and alf_cc_cr_filter_signal_flag may always be skipped and inferred to be equal to 0.
[0343] b. In one example, when aps_chroma_tool_params_present_flag is equal to 0, scaling_list_chroma_present_flag may always be skipped and inferred to be equal to 0.
[0344] Other Chroma Tool related parameters in the image header
[0345] 9. In one example, signaling / parsing of ph_log2_diff_min_qt_min_cb_intra_slice_luma may always be skipped when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied or a flag indicates that the syntax element ph_log2_diff_min_qt_min_cb_intra_slice_luma (and possibly other syntax elements) are not present.
[0346] 10. In one example, when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied or the flag indicates that the syntax element ph_log2_diff_min_qt_min_cb_intra_slice_chroma (and possibly other syntax elements) are not present, the signaling / parsing of ph_log2_diff_min_qt_min_cb_intra_slice_chroma can always be skipped.
[0347] 11. In one example, signaling / parsing of ph_log2_diff_min_qt_min_cb_inter_slice may always be skipped when ChromaArrayType is equal to 0 or the color format is 4:0:0 or separate plane coding is applied or a flag indicates that the syntax element ph_log2_diff_min_qt_min_cb_inter_slice (and possibly other syntax elements) are not present.
[0348] Adaptive Color Transform (ACT) related
[0349] 12. Palette mode and adaptive color transform can be applied exclusively to blocks.
[0350] a. In one example, when palette mode is used for a block, adaptive color transform is not applied to the block.
[0351] i. In one example, when palette mode is applied to a block, signaling of ACT usage can be skipped.
[0352] 1. Alternatively, furthermore, the use of ACT is inferred to be NO.
[0353] b. In one example, when adaptive color transform is used for a block, palette mode is not applied to the block.
[0354] i. In one example, when ACT is applied to a block, signaling of palette mode usage can be skipped.
[0355] 1. Alternatively, furthermore, the use of palette mode is inferred to be no.
[0356] c. Whether to signal the indication of the ACT on / off flag may depend on whether the prediction mode is not equal to MODE_PLT.
[0357] d. Whether to signal the indication of the ACT on / off flag may depend on whether the indication of palette mode is not used (eg, !pred_mode_plt_flag).
[0358] 13. Adaptive color transform can be applied to the residual block of the codec unit regardless of its color space.
[0359] e. In one example, adaptive color transform can be applied to the residual block of the codec unit in GBR color space.
[0360] f. In one example, adaptive color transform can be applied to the residual block of the codec unit in the YCbCr color space.
[0361] How to use bypass coding or context coding for coefficient sign flags
[0362] 14. At the beginning (or / and end) of bypass coding of the remaining syntax elements (e.g., syntax elements abs_remainder and coeff_sign_flag) in the third / remaining coefficient scan stage of the transform skip residual coding process, an operation may be applied to a variable specifying the number of remaining allowed context coding bins (e.g., RemCcbs).
[0363] g. In one example, the operation may use a temporary variable (eg, tempRemCcbs) to save RemCcbs and set RemCcbs to a certain value (such as 0). When the bypass codec ends, RemCcbs is set to be equal to tempRemCcbs.
[0364] i. In one example, the operation may be to set RemCcbs equal to some value N, where the value is an integer and less than M.
[0365] 1. In one example, M is equal to 3.
[0366] h. The syntax element indicating the sign of the coefficient level (eg, coeff_sign_flag) Whether to encode in bypass mode or context coding mode may depend on the number of remaining allowed context coding bins (eg, RemCcbs).
[0367] i. In one example, when the number of remaining allowed context coding bins (eg, RemCcbs) is equal to N (eg, N=0), the coefficient-level sign (eg, coeff_sign_flag) is coded in bypass mode.
[0368] ii. In one example, when RemCcbs is greater than or equal to M, the symbol flag is encoded and decoded using the context encoding and decoding mode.
[0369] i. In one example, the operation may be to set RemCcbs equal to a value that depends on at least one variable or syntax element other than RemCcbs.
[0370] 15. In the transform skip residual coding process, operations can be applied to the variable indicating whether the syntax element (eg, coeff_sign_flag) belongs to a certain scanning stage (eg, the first scanning stage or / and the third / residual coefficient scanning stage).
[0371] j. In one example, the operation may use a variable (eg, remScanPass) to indicate whether the current scan pass is the third / remaining scan pass.
[0372] i. In one example, at the beginning of the first scanning phase, remScanPass is set equal to A, and at the beginning of the third / remaining scanning phases, remScanPass is set equal to B. Where A is not equal to B.
[0373] ii. Whether the syntax element indicating the sign of the coefficient level (eg, coeff_sign_flag) is coded in the bypass mode or the context coding mode may depend on remScanPass.
[0374] 1. In one example, when remScanPass is equal to B, the coefficient-level signs are coded in bypass mode.
[0375] 2. In one example, when remScanPass is equal to A, the coefficient-level signs are coded using the context coding mode.
[0376] k. Alternatively, the operation may use a variable to indicate whether the current scan phase is the first scan pass phase.
[0377] 1. Alternatively, the operation may use a variable to indicate whether the current scan stage is the second / greater than X scan stages.
[0378] 16. Whether the syntax element (SE) indicating the sign of the coefficient level is coded in bypass mode or context coding mode may depend on the index of the scanning stage, where within the scanning stage, the same syntax element of one or more coefficients in a region of a block is coded in sequence.
[0379] m. In one example, when the SE is signaled in the first scanning phase, it can be encoded and decoded using the context encoding and decoding mode.
[0380] n. In one example, when the SE is signaled in the third / remaining scanning phase, it can be coded in bypass mode.
[0381] o. In one example, the above method is applicable to a transform skip (TS, including or excluding BDPCM / QR-BDPCM) residual coding process and / or a coefficient coding process for a non-TS coded block.
[0382] 17. Whether the syntax element (SE) indicating the sign of the coefficient level is coded in bypass mode or context coding mode may depend on whether it is signaled in the same scanning phase with another syntax element (e.g., sig_coeff_flag, par_level_flag, abs_remainder) during transform skip residual coding.
[0383] p. In one example, when SE is signaled in the same scanning phase as sig_coeff_flag or / and par_level_flag, it can be coded in context coding mode.
[0384] q. In one example, when SE is signaled in the same scan phase as abs_remainder, it can be coded in bypass mode.
[0385] General Features
[0386] 18. Whether and / or how the above approach should be applied may be based on:
[0387] a. Video content (e.g., screen content or natural content)
[0388] b. Messages signaled in DPS / SPS / VPS / PPS / APS / picture header / slice header / slice group header / largest codec unit (LCU) / codec unit (CU) / LCU row / LCU group / TU / PU block / video codec unit
[0389] c. Location of CU / PU / TU / block / video codec unit
[0390] d. Block dimensions of the current block and / or its neighboring blocks
[0391] e. Block shape of the current block and / or its neighboring blocks
[0392] f. Quantization parameter of the current block
[0393] g. Indication of the color format (such as 4:2:0, 4:4:4, RGB, or YUV)
[0394] h. Codec tree structure (such as dual tree or single tree)
[0395] i. Slice / slice group type and / or picture type
[0396] j. Color components (e.g., may only be applied to luma components and / or chroma components)
[0397] k. Temporal layer ID
[0398] l. Standard profiles / levels / tiers
[0399] m. Whether the current block has an escape sample point.
[0400] i. In one example, the above method can be applied only when the current block has at least one escaped sample.
[0401] n. Whether the current block is encoded or decoded in lossless mode (for example, cu_transquant_bypass_flag)
[0402] ii. In one example, the above method can be applied only when the current block is not encoded or decoded in lossless mode.
[0403] o. Whether lossless codec is enabled (e.g., transquant_bypass_enabled, cu_transquant_bypass_flag)
[0404] 5. Examples
[0405] In the following examples, added parts are marked with bold, underlined and italic text. Deleted parts are marked with [[ ]].
[0406] 5.1 Example #1
[0407] 8.4.5.3 Palette Mode Decoding Process
[0408] – If bIsEscapeSample is equal to 0, the following applies:
[0409] recSamples[x][y]=
[0410] CurrentPaletteEntries[cIdx][PaletteIndexMap[xCbL+xL][yCbL+yL]](443)
[0411] Otherwise (bIsEscapeSample is equal to 1), the following ordered steps are applied:
[0412] 1. The quantization parameter qP is derived as follows:
[0413] – If cIdx is equal to 0,
[0414]
[0415] – Otherwise, if cIdx is equal to 1,
[0416]
[0417] – Otherwise (cIdx is equal to 2),
[0418]
[0419] 2. The list levelScale[] is specified as levelScale[k] = {40, 45, 51, 57, 64, 72}, where k = 0..5.
[0420] 3. The following applies:
[0421] tmpVal=(PaletteEscapeVal[cIdx][xCbL+xL][yCbL+yL]*levelScale[qP%6])<<(qP / 6)+32)>>6 (447)
[0422] recSamples[x][y]=Clip3(0,(1< <BitDepth)-1,tmpVal) (448)
[0423] 5.2 Example #2
[0424] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0425]
[0426] 7.3.2.7 Picture header structure syntax
[0427]
[0428]
[0429] 7.3.7.1 General Strip Header Syntax
[0430]
[0431]
[0432] 5.3 Example #3
[0433] 7.4.2.4 Picture Parameter Set RBSP Syntax
[0434]
[0435]
[0436] 7.3.2.5 Adaptation Parameter Set RBSP Syntax
[0437]
[0438] 7.3.2.7 Picture header structure syntax
[0439]
[0440]
[0441]
[0442]
[0443]
[0444] 7.3.2.19 Adaptive Loop Filter Data Syntax
[0445]
[0446] 7.3.2.21 Zoom List Data Syntax
[0447]
[0448] 7.3.7.1 General Strip Header Syntax
[0449]
[0450]
[0451] 7.4.3.4 Picture Parameter Set RBSP Semantics
[0452] …
[0453] [[offsets]] Equal to 1 specifies the presence of syntax elements related to the chroma tool [[offset]] in the PPS RBSP syntax structure. [[offsets]]_present_flag equal to 0 specifies that there are no syntax elements related to the chroma tool [[offsets]] in the PPS RBSP syntax structure. When ChromaArrayType is equal to 0, pps_chroma_tool_ The value of [[offsets]]_present_flag shall be equal to 0.
[0454] …
[0455] 7.4.3.5 Adaptation parameter set semantics
[0456] …
[0457]
[0458] …
[0459] 5.4 Example #4
[0460] 7.4.2.4 Picture Parameter Set RBSP Syntax
[0461]
[0462]
[0463] 7.3.2.7 Picture header structure syntax
[0464]
[0465]
[0466] 7.3.7.1 General Strip Header Syntax
[0467]
[0468]
[0469] 7.4.3.4 Picture Parameter Set RBSP Semantics
[0470] …
[0471] 1 specifies that chroma deblocking related syntax elements are present in the PPS RBSP syntax structure. pps_chroma_deblocking_params_present_flag equal to 0 specifies that chroma deblocking related syntax elements are not present in the PPS RBSP syntax structure. When ChromaArrayType is equal to 0, the value of pps_chroma_deblocking_params_present_flag shall be equal to 0.
[0472] …
[0473] 5.5 Example #5
[0474]
[0475]
[0476]
[0477] 5.5.1 Example #5.1
[0478] 7.3.10.5 Codec unit syntax
[0479]
[0480]
[0481]
[0482]
[0483] Alternatively, the following may apply:
[0484]
[0485] 5.6 Example #6
[0486] Codec unit semantics
[0487] …
[0488] Equal to 1 specifies The residual of the current codec unit [[in YC g C o The color space is encoded and decoded]]. cu_act_enabled_flag is equal to 0 to specify The residual of the current codec unit [[encoded in the original color space]]. When cu_act_enabled_flag is not present, it is inferred to be equal to 0.
[0489] 5.7 Example #7
[0490] 7.3.10.11 Residual Codec Syntax
[0491]
[0492]
[0493]
[0494]
[0495]
[0496] 9.3.4 Decoding Process Flow
[0497] 9.3.4.2 Derivation of ctxTable, ctxIdx, and bypassFlag
[0498] 9.3.4.2.1 General
[0499] Table 131 – Syntax element assigning ctxInc to bins with context codec
[0500]
[0501] 5.8 Example #8
[0502] 7.3.10.11 Residual Codec Syntax
[0503]
[0504]
[0505]
[0506]
[0507]
[0508]
[0509] Where A is not equal to B. For example, A=0, B=1.
[0510] Alternatively, A=1, B=0.
[0511] Alternatively, A=-1, B=0.
[0512] Alternatively, A=0, B=-1.
[0513] 9.3.4 Decoding Process Flow
[0514] 9.3.4.2 Derivation of ctxTable, ctxIdx, and bypassFlag
[0515] 9.3.4.2.1 General
[0516] Table 131 - Syntax elements assigning ctxInc to bins with context codec
[0517]
[0518]
[0519] 5.9 Example #9
[0520] 7.3.10.11 Residual Codec Syntax
[0521]
[0522]
[0523]
[0524]
[0525]
[0526] 7.4.11.11 Residual Encoding and Decoding Semantics
[0527] Specifies the sign of the transform coefficient level at scan position n as follows:
[0528] – If coeff_sign_flag[n] is equal to 0, the corresponding transform coefficient level has a positive value.
[0529] – Otherwise (coeff_sign_flag[n] is equal to 1), the corresponding transform coefficient level has a negative value.
[0530] When coeff_sign_flag[n] is not present, it is inferred to be equal to 0.
[0531]
[0532] The value of CoeffSignLevel[xC][yC] specifies the sign of the transform coefficient level at location (xC, yC) as follows:
[0533] – If CoeffSignLevel[xC][yC] is equal to 0, the corresponding transform coefficient level is equal to 0
[0534] – Otherwise, if CoeffSignLevel[xC][yC] is equal to 1, the corresponding transform coefficient level has a positive value.
[0535] – Otherwise (CoeffSignLevel[xC][yC] is equal to -1), the corresponding transform coefficient level has a negative value.
[0536] 9.3.2 Initialization Process
[0537] 9.3.2.2 Context variable initialization process
[0538] Table 125–coeff_sign _ctx_coding_ Specification of initValue and shiftIdx of flag's ctxInc
[0539]
[0540] 9.3.3 Binarization Process
[0541] 9.3.3.1 General
[0542] Table 126 - Syntax elements and associated binarization
[0543]
[0544] 9.3.4 Decoding Process Flow
[0545] 9.3.4.2 Derivation of ctxTable, ctxIdx, and bypassFlag
[0546] 9.3.4.2.1 General
[0547] Table 131 – Syntax element assigning ctxInc to bins with context codec
[0548]
[0549]
[0550] 9.3.4.2.10 Syntax element coeff_sign for transform skip mode _ctx_coding_ The derivation process of flag's ctxInc
[0551] The input of this process is the color component index cIdx, the brightness position (x0, y0) of the upper left sample of the current transform block relative to the upper left sample of the current picture, and the current coefficient scanning position (xC, yC).
[0552] The output of this process is the variable ctxInc.
[0553] The variables leftSign and aboveSign are derived as follows:
[0554] leftSign=(xC==0)? 0:CoeffSignLevel[xC-1][yC] (1594)
[0555] aboveSign=(yC==0)? 0:CoeffSignLevel[xC][yC-1] (1595)
[0556] The variable ctxInc is derived as follows:
[0557] – If leftSign is equal to 0 and aboveSign is equal to 0, or if leftSign is equal to aboveSign, then the following applies:
[0558] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?0:3) (1596)
[0559] – Otherwise, if leftSign is greater than or equal to 0 and aboveSign is greater than or equal to 0, then the following applies:
[0560] ctxInc=(BdpcmFlag[x0][y0][cIdx]?1:4) (1597)
[0561] – Otherwise, the following applies:
[0562] ctxInc=(BdpcmFlag[x0][y0][cIdx]?2:5) (1598)
[0563] 5.10 Example #10
[0564] 9.3.4.2.10 Derivation of ctxInc for the syntax element coeff_sign_flag for transform skip mode
[0565] The input of this process is the color component index cIdx, the brightness position (x0, y0) of the upper left sample of the current transform block relative to the upper left sample of the current picture, and the current coefficient scanning position (xC, yC).
[0566] The output of this process is the variable ctxInc.
[0567] The variables leftSign and aboveSign are derived as follows:
[0568] leftSign=(xC==0)? 0:CoeffSignLevel[xC-1][yC] (1594)
[0569] aboveSign=(yC==0)? 0:CoeffSignLevel[xC][yC-1] (1595)
[0570] The variable ctxInc is derived as follows:
[0571] – If leftSign is equal to 0 and aboveSign is equal to 0, or if leftSign is equal to aboveSign, then the following applies:
[0572] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?0:3) (1596)
[0573] – Otherwise, if leftSign is greater than or equal to 0 and aboveSign is greater than or equal to 0, then the following applies:
[0574]
[0575] – Otherwise, the following applies:
[0576]
[0577] Figure 7 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces, such as Ethernet, a passive optical network (PON), and the like, as well as wireless interfaces, such as Wi-Fi or cellular interfaces.
[0578] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to generate a coded representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. As represented by component 1906, the output of codec component 1904 can be stored or sent via a connected communication. Component 1908 can use the stored or transmitted bitstream (or coded) representation of the video received at input 1902 to generate pixel values or displayable video sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations opposite to the encoding results will be performed by the decoder.
[0579] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0580] Figure 8 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor(s) 3602 can be configured to implement one or more methods described in this document. Memory(s) 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0581] Figure 10 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.
[0582] like Figure 10 As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.
[0583] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0584] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. Associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 via the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0585] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0586] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain coded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the coded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, with the destination device 120 being configured to interface with an external display device.
[0587] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.
[0588] Figure 11 is a block diagram illustrating an example of a video encoder 200, which may be Figure 10 The video encoder 114 in the system 100 is shown.
[0589] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 11In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0590] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0591] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.
[0592] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated but are not shown here for explanation purposes. Figure 11 In the example, they are represented separately.
[0593] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0594] The mode selection unit 203 may, for example, select a codec mode - intra or inter - based on the error result, and provide the resulting intra or inter codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the codec block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[0595] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from the buffer 213.
[0596] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0597] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0598] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0599] In some examples, motion estimation unit 204 may output complete motion information for use in the decoding process of a decoder.
[0600] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for the neighboring video block.
[0601] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0602] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0603] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0604] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0605] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0606] In other examples, the current video block may not have residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform the subtraction operation.
[0607] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0608] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0609] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0610] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0611] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0612] Figure 12 is a block diagram illustrating an example of a video decoder 300, which may be Figure 10 The video decoder 114 in the system 100 is shown.
[0613] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 12 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0614] exist Figure 12 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those generally performed for the video encoder 200 ( Figure 11 ) is a decoding process that is the inverse of the encoding process described.
[0615] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information from the entropy-decoded video data. The motion compensation unit 302 can determine this information, for example, by performing AMVP and merge modes.
[0616] The motion compensation unit 302 may generate a motion compensated block and may perform interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0617] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used by video encoder 20 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to produce a prediction block.
[0618] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and a list of reference frames) for each inter-coded block, and other information for decoding the encoded video sequence.
[0619] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0620] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0621] The following provides a list of preferred solutions for some embodiments.
[0622] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, item 1).
[0623] 1. A video processing method (e.g., Figure 9 ), comprising: performing (902) a conversion between a video block of a video and a codec representation of the video, wherein a palette mode is applied to the codec representation of the video block, in which the samples of the video block are represented using a palette of representative color values; and wherein the samples outside the palette are encoded and decoded using an escape character and a value quantized using a quantization parameter, the quantization parameter being within a range between a minimum allowed value and a maximum allowed value determined by a rule.
[0624] 2. The method according to solution 1, wherein the maximum allowed value depends on the binarization method used for the codec representation of the video block.
[0625] 3. The method according to solution 1, wherein the maximum allowed value is expressed as T+B, where B is a number based on the bit depth of the representation of samples of the video block and T is a predefined number.
[0626] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 2).
[0627] 4. A video processing method comprising: performing a conversion between a video block of a video and a codec representation of the video, wherein a palette mode is applied to the codec representation of the video block, in which a palette of representative color values is used to represent samples of the video block; and wherein a size of the palette depends on a rule regarding whether a local dual tree is applied to the conversion between the video block and the codec representation.
[0628] 5. The method of solution 4, wherein the size of the palette depends on the color components of the video due to the use of local dual trees.
[0629] 6. The method of solution 5, wherein the rule specifies that a smaller palette size is used when the video block is a chroma block than when the video block is a luma block.
[0630] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 3).
[0631] 7. A video processing method comprising: performing conversion between a video block of a video and a codec representation of the video, wherein a palette mode is used for the codec representation of the video block, in which a palette of representative color values is used to represent samples of the video block; and wherein the size of a palette predictor depends on a rule regarding whether a local dual tree is used for conversion between the video block and the codec representation.
[0632] 8. The method of solution 7, wherein the size of the palette predictor depends on the color components of the video block due to the use of a local dual tree.
[0633] 9. The method of solution 8, wherein the rule specifies that a smaller palette size is used when the video block is a chroma block than when the video block is a luma block.
[0634] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 4).
[0635] 10. A video processing method, comprising: for conversion between a video block of a video region of a video and a codec representation of the video, determining, based on a codec condition, whether a syntax element identifying a deblocking offset for a chroma component of the video is included in the codec representation at the video region level; and performing the conversion based on the determination; wherein the deblocking offset is used to selectively enable a deblocking operation on the video block.
[0636] 11. The method according to solution 10, wherein the video area is a video strip or a video picture.
[0637] 12. The method according to any one of solutions 10-11, wherein the encoding and decoding conditions include the color format of the video.
[0638] 13. The method according to any of solutions 10-12, wherein the codec condition is based on whether separate plane codec is enabled for the conversion.
[0639] 14. A method according to any of solutions 10-13, wherein the encoding solution condition is based on whether the chroma array type is included in the codec representation.
[0640] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 5).
[0641] 15. A video processing method comprising: for conversion between video blocks of a video region of a video and a codec representation of the video, determining, based on a codec condition, whether a syntax element identifying use of a chroma codec tool is included in the codec representation at the video region level; and performing the conversion based on the determination; wherein a deblocking offset is used to selectively enable deblocking operations on the video blocks.
[0642] 16. The method of solution 15, wherein the video region is a video strip or a video picture.
[0643] 17. The method according to any of solutions 15-16, wherein the coding condition corresponds to including syntax elements in an adaptation parameter set.
[0644] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 6, 7).
[0645] 18. A video processing method, comprising: performing conversion between video blocks of a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format; wherein the format specifies whether a first flag indicating a deblocking offset for a chroma component of the video is included in the codec representation based on whether a second flag indicating a quantization parameter offset for the chroma component is included in the codec representation.
[0646] 19. The method of solution 1, wherein the format rule specifies that the codec representation includes a third flag that indicates whether the first flag and the second flag are included in the codec representation.
[0647] 20. The method according to any of solutions 1-2, wherein the third flag is included in the codec representation in the picture parameter set.
[0648] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 8-12).
[0649] 21. A video processing method, comprising: performing conversion between a video block of a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule; wherein the format rule specifies whether syntax elements in the codec representation control one or more parameters indicating the applicability of one or more chroma codec tools to be included in the codec representation of the video region or video block.
[0650] 22. The method of solution 1, wherein the syntax element is included in an adaptation parameter set.
[0651] 23. The method of any of solutions 1-2, wherein the format rule specifies that a first value of a syntax element indicates that one or more parameters are excluded from the codec representation and are skipped during parsing of the codec representation.
[0652] 24. The method according to any of the preceding solutions, wherein the conversion uses the method because the video meets a condition.
[0653] 25. The method of solution 24, wherein the condition includes a codec indicating the type or profile or level or tier of video content used.
[0654] 26. The method according to any of the preceding solutions, wherein the conditions include block dimensions of the video block and / or neighboring video blocks or color format of the video or codec tree structure for video block conversion or type of video region.
[0655] 27. A method according to any of solutions 1 to 26, wherein converting includes encoding the video into a codec representation.
[0656] 28. A method according to any of solutions 1 to 26, wherein converting includes decoding the codec representation to generate pixel values of the video.
[0657] 29. A video decoding device comprising a processor configured to implement the method described in one or more of solutions 1 to 28.
[0658] 30. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions 1 to 28.
[0659] 31. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions 1 to 28.
[0660] 32. The methods, apparatus, or systems described in this document.
[0661] Figure 13 13 is a flowchart representation of a video processing method 1300 according to the present technology. The method 1300 includes, at operation 1310, performing conversion between a video block of a video and a bitstream of the video using a palette mode in which samples of the video block are represented using a palette of representative color values. A size of the palette for the video block is determined based on whether a local dual tree is applied to the video block.
[0662] In some embodiments, where a local dual tree is applied to a codec unit, based on a segmentation parameter of the codec unit, the segmentation operation is applied only to the luma block of the codec unit, and based on a mode type of the codec unit, the segmentation operation is not applied to at least one chroma block of the codec unit. In some embodiments, where a local dual tree is applied to a video block, the size of the palette is reduced. In some embodiments, where a local dual tree is applied to a video block, the size of the palette is reduced compared to the size of the palette when the local dual tree is not applied to the video block. In some embodiments, where a local dual tree is applied to a video block, the size of the palette is reduced compared to the size of the palette when a conventional single tree is applied to the video block. In some embodiments, the size of the palette is halved.
[0663] In some embodiments, when the local dual tree is applied, the size of the palette of the video block of the chroma component is different from the size of the palette of the video block of the luma component. In some embodiments, the size of the palette of the video block of the chroma component is smaller than the size of the palette of the video block of the luma component. In some embodiments, the size of the palette of the video block of the chroma component is half the size of the palette of the video block of the luma component.
[0664] Figure 141 is a flowchart representation of a video processing method 1400 according to the present technology. The method 1400 includes, at operation 1410, performing conversion between a video block of a video and a bitstream of the video using a palette mode in which samples of the video block are represented using a palette of representative color values. The size of the palette predictor for the video block is based on whether a local dual tree is applied to the video block.
[0665] In some embodiments, when a local dual tree is applied to the conversion, the size of the palette predictor is reduced. In some embodiments, when a local dual tree is applied, the size of the palette predictor for the video block of the chroma component is different from the size of the palette predictor for the video block of the luma component. In some embodiments, the size of the palette predictor for the video block of the chroma component is smaller than the size of the palette for the video block of the luma component. In some embodiments, the size of the palette predictor for the video block of the chroma component is halved compared to the size of the palette for the video block of the luma component.
[0666] Figure 15 1 is a flow chart representation of a video processing method 1500 according to the present technology. The method 1500 includes, at operation 1510, performing conversion between a video block of a video and a bitstream of the video using a palette mode in which samples of the video block are represented using a palette of representative color values. The conversion complies with a rule that specifies encoding values of escaped samples in the bitstream using a quantization parameter that is constrained by at least a maximum allowed value or a minimum allowed value.
[0667] In some embodiments, the escape samples include a subset of samples of representative color values that are not classified into the palette, and the quantization parameter is constrained to be no greater than a maximum allowed value or no less than a minimum allowed value. In some embodiments, the maximum allowed value is determined based on a binarization method used for conversion. In some embodiments, the maximum allowed value is expressed as (T+B), where B represents the bit depth. In some embodiments, T is indicated in the video region of the bitstream. In some embodiments, T is signaled in a video parameter set, a sequence parameter set, a picture parameter set, a picture header, or a slice header. In some embodiments, B is equal to the bit depth offset associated with the quantization parameter. In some embodiments, the maximum allowed value is equal to the bit depth offset associated with the quantization parameter plus T. In some embodiments, T is a constant. In some embodiments, T is equal to 23, 35, or 39. In some embodiments, T is a constant less than 23, less than 35, or less than 29. In some embodiments, the value of the escape samples is encoded using the bit length of EG5, EG3, or EG4.
[0668] Figure 161 is a flow chart representation of a video processing method 1600 according to the present technology. The method 1600 includes, at operation 1610, performing conversion between a block of video and a bitstream of the video. The conversion complies with format rules that specify whether parameters associated with a chroma codec tool are present in an adaptation parameter set in the bitstream based on a control flag in the adaptation parameter set.
[0669] In some embodiments, when the control flag is equal to 0, the parameter is omitted from the adaptive parameter set. In some embodiments, the parameter includes at least a signal flag of a chroma component of the adaptive loop filter, a signal flag of a Cb component of the adaptive loop filter, a signal flag of a Cr component of the adaptive loop filter, or a signal flag indicating whether a scaling list for the chroma component exists.
[0670] Figure 17 is a flowchart representation of a video processing method 1700 according to the present technology. The method 1700 includes, at operation 1710, performing conversion between blocks of a video and a bitstream of the video. The bitstream conforms to a format rule that specifies that syntax elements associated with quantization parameters are omitted in a picture header of the bitstream when the video is monochrome or the color components of the video are processed separately. In some embodiments, the video is determined to be monochrome based on (1) the syntax element ChromaArrayType being equal to 0 and (2) the video color format being equal to 4:0:0. In some embodiments, the syntax element includes ph_log2_diff_min_qt_min_cb_intra_slice_luma, ph_log2_diff_min_qt_min_cb_intra_slice_chroma, or ph_log2_diff_min_qt_min_cb_inter_slice.
[0671] Figure 18 is a flowchart representation of a video processing method 1800 according to the present technology. The method 1800 includes, at operation 1810, performing conversion between a video block of a video and a bitstream of the video according to a rule, the rule specifying that for the conversion, a palette mode for representing samples of the video block using a palette of representative color values and an adaptive color transform mode for performing color space conversion in the residual domain are mutually exclusively enabled.
[0672] In some embodiments, in the case where palette mode is applied to a conversion, the adaptive color transform mode is disabled. In some embodiments, signaling of information of the adaptive color transform mode is omitted for the conversion. In some embodiments, use of the adaptive color transform mode is inferred to be disabled. In some embodiments, in the case where adaptive color transform mode is applied to a conversion, the palette mode is disabled. In some embodiments, signaling of information of the palette mode is omitted for the conversion. In some embodiments, use of the palette mode is inferred to be disabled.
[0673] Figure 19 1 is a flowchart representation of a video processing method 1900 according to the present technology. Method 1900 includes, at operation 1910, performing conversion between a video block of a video and a bitstream of the video. An adaptive color transform mode that performs color space conversion in the residual domain is applied to a residual block of the video block, regardless of the color space of the residual block. In some embodiments, the color space of the residual block includes a green-blue-red (GBR) color space or a YCbCr color space.
[0674] Figure 20 2 is a flowchart representation of a video processing method 2000 according to the present technology. The method 2000 includes, at operation 2010, performing conversion between a video block of a video and a bitstream of the video. The video block is encoded using a transform skip residual codec tool, wherein residual coefficients of transform skip codecs of the video block are encoded using a context codec process or a bypass codec process. During the conversion, an operation is applied to a variable specifying the number of remaining context codec bins allowed in the video block at the beginning or end of the bypass codec process.
[0675] In some embodiments, the operation includes storing the number of bins of remaining context coding allowed in the video block in a temporary variable; and setting the variable based on the temporary variable. In some embodiments, the operation includes setting the variable to a value N, where N is an integer. In some embodiments, N is less than M, and M is equal to 3. In some embodiments, N is based on another variable or another syntax element. In some embodiments, the syntax element indicating the sign of the coefficient level encodes the number of bins of remaining context coding allowed in the video block using a bypass codec process or a context codec process.
[0676] In some embodiments, the sign at the coefficient level is encoded using a bypass codec process when the number of remaining context coding bins allowed in the video block is equal to N, where N is an integer equal to or greater than 0. In some embodiments, the sign at the coefficient level is encoded using a context codec process when the number of remaining context coding bins allowed in the video block is greater than or equal to M, where M is an integer.
[0677] Figure 21 2 is a flowchart representation of a video processing method 2100 according to the present technology. The method 2100 includes, at operation 2110, performing conversion between a video block of a video and a bitstream of the video using a transform skip residual codec process. During the conversion, an operation is applied to a variable indicating whether a syntax element belongs to a particular scan phase.
[0678] In some embodiments, the operations include assigning a value to a variable indicating whether the current scan phase is the third or remaining scan phase. In some embodiments, the operations include assigning a value to a variable indicating whether the current scan phase is the first scan phase. In some embodiments, the operations include assigning a value to a variable indicating whether the current scan phase is the second or a scan phase greater than X. In some embodiments, the operations include assigning a first value to the variable at the beginning of the first scan phase and assigning a second value to the variable at the beginning of the third or remaining scan phase, wherein the first value is not equal to the second value.
[0679] In some embodiments, whether a syntax element indicating a coefficient-level symbol is encoded using a bypass codec process or a context codec process is based on a variable. In some embodiments, if the variable indicates that the particular scan stage is the third or remaining scan stage, the coefficient-level symbol is encoded using the bypass codec process. In some embodiments, if the variable indicates that the particular scan stage is the first scan stage, the coefficient-level symbol is encoded using the context codec process.
[0680] Figure 22 2 is a flowchart representation of a video processing method 2200 according to the present technology. The method 2200 includes, at operation 2210, performing conversion between a video block of a video and a bitstream of the video. During the conversion, whether a syntax element indicating a sign of a coefficient level is encoded using a bypass codec process or a context codec process is indexed based on a scan phase in which identical syntax elements for one or more coefficients in a region of the video block are encoded sequentially. In some embodiments, the conversion is performed using a transform skip residual codec process or a coefficient codec process, where the video block is non-transform skip codec.
[0681] Figure 23 23 is a flowchart representation of a video processing method 2300 according to the present technology. The method 2300 includes, at operation 2310, performing conversion between a video block of a video and a bitstream of the video using a transform skip residual codec process. During conversion, whether a syntax element indicating a sign of a coefficient level is coded using a bypass codec process or a context codec process is based on whether the syntax element is signaled in the same scan phase as another syntax element.
[0682] In some embodiments, the syntax element is signaled as sig_coeff_flag or par_level_flag.In some embodiments, where the syntax element is signaled in the same scanning phase as abs_remainder, the syntax element is encoded using a bypass encoding process.
[0683] In some embodiments, the applicability of one or more of the above methods is based on characteristics of the video. In some embodiments, the characteristics include the content of the video. In some embodiments, the characteristics include messages signaled in a decoder parameter set, sequence parameter set, video parameter set, picture parameter set, adaptation parameter set, picture header, slice header, slice group header, largest codec unit (LCU), codec unit (CU), LCU row, LCU group, transform unit (TU), picture unit (PU) block, or video codec unit. In some embodiments, the characteristics include the position of the codec unit, picture unit, transform unit, or block. In some embodiments, the characteristics include the dimensions or shape of the video block and / or its neighboring blocks. In some embodiments, the characteristics include quantization parameters of the video block. In some embodiments, the characteristics include the color format of the video. In some embodiments, the characteristics include the codec tree structure of the video. In some embodiments, the characteristics include the type of slice, slice group, or picture. In some embodiments, the characteristics include the color components of the video block. In some embodiments, the characteristics include a temporal layer identifier. In some embodiments, the characteristics include a profile, level, or tier of a video standard. In some embodiments, the characteristics include whether the video block includes escape samples. In some embodiments, the method is only applicable when the video block includes at least one escape sample. In some embodiments, the characteristic includes whether the video block is encoded using a lossless mode. In some embodiments, the method is only applicable when the video block is not encoded or decoded using a lossless mode.
[0684] In some embodiments, converting includes encoding the video into a bitstream. In some embodiments, converting includes decoding the video from the bitstream.
[0685] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. As defined by the syntax, the bitstream representation of the current video block may, for example, correspond to bits that are co-located or interspersed at different locations within the bitstream. For example, a macroblock may be encoded based on a transformed and coded error residual value, and also using bits from a header and other fields in the bitstream.
[0686] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this application document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or any combination thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible, non-volatile computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. A computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or any combination thereof. The term "data processing unit" or "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or a plurality of processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or any combination thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0687] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one or more computers, located at one site or distributed across multiple sites and interconnected by a communications network.
[0688] The processes and logic flows described in this application document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and the apparatus can also be implemented as, special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0689] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a processor that executes instructions and one or more memory devices that store instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0690] While this patent document contains many specifics, they should not be construed as limitations on the scope of any invention or the claims, but rather as descriptions of features for particular embodiments of particular inventions. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment can also be implemented separately in multiple embodiments, or in any suitable subcombination. Furthermore, while the features described above may be described as functioning in certain combinations, or even initially claimed to be so, in some cases one or more features in a claim combination may be removed from the combination, and a claim combination may be directed to a subcombination or variations of a subcombination.
[0691] Likewise, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0692] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing video data, comprising: For conversion between a first video block of a video and a bitstream of the video, determining that a first prediction mode is applied to the first video block; Maintain predictor palette table; constructing, in the first prediction mode, a first palette comprising one or more palette predictors for the first video block based on the predictor palette table; as well as performing said converting based on said first prediction mode, wherein, in the first prediction mode, reconstructed samples of the first video block are represented by a set of representative color values, the set of representative color values comprising at least one of 1) a palette predictor derived from the first palette, 2) escaped samples, or 3) palette information included in the bitstream, wherein the maximum number of entries in the first palette is based on a tree type of the first video block, wherein the first prediction mode is applied to a second video block of the video, and a second palette of the second video block is constructed based on the predictor palette table, and Wherein, when a single tree is applied to the first video block and a local dual tree is applied to the second video block, a maximum number of entries in the second palette is half of a maximum number of entries in the first palette.
2. The method according to claim 1, wherein The predictor palette table is updated based on the first palette.
3. The method according to claim 1, wherein The second video block is a luma block, and the second video block has a tree type not equal to single-tree and has a slice type not equal to I-slice.
4. The method according to claim 1, wherein The second video block is a luma block, and the second video block is obtained by splitting a luma parent block of a codec tree node, wherein splitting a chroma parent block in the codec tree node is not allowed.
5. The method according to claim 4, wherein The first video block is a luma block, and a palette size of a chroma block corresponding to the first video block is larger than a palette size of the chroma parent block.
6. The method according to claim 4, wherein: The size of the second palette is larger than the palette size of the chroma parent block.
7. The method according to claim 1, wherein The first prediction mode is applied to a second video block of the video, and a second palette for the second video block is constructed based on the predictor palette table; as well as When a dual-tree is applied to the first video block and a local dual-tree is applied to the second video block, a maximum number of entries in the predictor palette table is different for the first video block and the second video block.
8. The method according to claim 7, wherein: The first video block is a chroma block, wherein the second video block is a luma block, and the second video block is obtained by splitting a luma parent block of a codec tree node, wherein splitting a chroma parent block of the codec tree node is not allowed, and The size of the first palette is larger than the palette size of the chroma parent block.
9. The method according to claim 1, wherein The converting includes encoding the video into the bitstream.
10. The method according to claim 1, wherein The converting includes decoding the video from the bitstream.
11. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For conversion between a first video block of a video and a bitstream of the video, determining that a first prediction mode is applied to the first video block; Maintain predictor palette table; In the first prediction mode, constructing a first palette comprising one or more palette predictors for the first video block based on the predictor palette table; and performing said converting based on said first prediction mode, in, In the first prediction mode, reconstructed samples of the first video block are represented by a set of representative color values, the set of representative color values comprising at least one of 1) a palette predictor derived from the first palette, 2) escape samples, or 3) palette information included in the bitstream, wherein the maximum number of entries in the first palette is based on a tree type of the first video block, wherein the first prediction mode is applied to a second video block of the video, and a second palette of the second video block is constructed based on the predictor palette table, and Wherein, when a single tree is applied to the first video block and a local dual tree is applied to the second video block, a maximum number of entries in the second palette is half of a maximum number of entries in the first palette.
12. A non-transitory computer-readable storage medium having stored therein instructions that cause a processor to: For conversion between a first video block of a video and a bitstream of the video, determining that a first prediction mode is applied to the first video block; Maintain predictor palette table; In the first prediction mode, constructing a first palette comprising one or more palette predictors for the first video block based on the predictor palette table; and performing said converting based on said first prediction mode, in, In the first prediction mode, reconstructed samples of the first video block are represented by a set of representative color values, the set of representative color values comprising at least one of 1) a palette predictor derived from the first palette, 2) escape samples, or 3) palette information included in the bitstream, wherein the maximum number of entries in the first palette is based on a tree type of the first video block, wherein the first prediction mode is applied to a second video block of the video, and a second palette of the second video block is constructed based on the predictor palette table, and Wherein, when a single tree is applied to the first video block and a local dual tree is applied to the second video block, a maximum number of entries in the second palette is half of a maximum number of entries in the first palette.
13. A non-transitory computer-readable recording medium storing a video bitstream generated by a method performed by a video processing device, wherein the method comprises: For a first video block of a video, determining that a first prediction mode is applied to the first video block; Maintain predictor palette table; constructing, in the first prediction mode, a first palette comprising one or more palette predictors for the first video block based on the predictor palette table; as well as generating a bitstream of the video based on the first prediction mode, wherein, in the first prediction mode, reconstructed samples of the first video block are represented by a set of representative color values, the set of representative color values comprising at least one of 1) a palette predictor derived from the first palette, 2) escaped samples, or 3) palette information included in the bitstream, wherein the maximum number of entries in the first palette is based on a tree type of the first video block, wherein the first prediction mode is applied to a second video block of the video, and a second palette of the second video block is constructed based on the predictor palette table, and Wherein, when a single tree is applied to the first video block and a local dual tree is applied to the second video block, a maximum number of entries in the second palette is half of a maximum number of entries in the first palette.
14. A method for storing a video bitstream, comprising: For a first video block of a video, determining that a first prediction mode is applied to the first video block; Maintain predictor palette table; constructing, in the first prediction mode, a first palette comprising one or more palette predictors for the first video block based on the predictor palette table; generating a bitstream of the video based on the first prediction mode; as well as storing the bitstream in a non-transitory computer-readable recording medium, wherein, in the first prediction mode, reconstructed samples of the first video block are represented by a set of representative color values, the set of representative color values comprising at least one of 1) a palette predictor derived from the first palette, 2) escaped samples, or 3) palette information included in the bitstream, wherein the maximum number of entries in the first palette is based on a tree type of the first video block, wherein the first prediction mode is applied to a second video block of the video, and a second palette of the second video block is constructed based on the predictor palette table, and Wherein, when a single tree is applied to the first video block and a local dual tree is applied to the second video block, a maximum number of entries in the second palette is half of a maximum number of entries in the first palette.
Citation Information
Patent Citations
Methods for palette mode context coding and binarization in video coding
CN107431824A
Motion vector prediction based on lookup table is extended with temporal information
CN110719463A