Derivation of quantization parameters for palette modes
Fixed-length and variable-length coding techniques, combined with optimized local dual-tree structures and interpolation filters, address inefficiencies in palette mode encoding, improving decoding quality and throughput in VVC standards.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-09-18
- Publication Date
- 2026-03-17
AI Technical Summary
Current video encoding technologies face challenges in efficient handling of palette mode encoding, including variable-length coding of escape symbols, parsing dependencies, and hardware processing throughput issues due to local dual trees and chroma QP table constraints, which affect decoding quality and efficiency.
Implement fixed-length coding for escape symbols, disable escape symbols in certain conditions, apply variable-length coding with truncated binary or exponential Golomb methods, and use a local dual-tree coding structure to optimize chroma block sizes, along with alternative interpolation filters for improved hardware processing.
Enhances decoding quality and efficiency by stabilizing escape symbol coding, reducing parsing dependencies, and optimizing hardware throughput, particularly in versatile video coding (VVC) standards.
Smart Images

Figure 0007832105000015 
Figure 0007832105000016 
Figure 0007832105000017
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application is a national phase entry of the international patent application PCT / US2020 / 051540, filed on September 18, 2020. This is done in a timely manner, claiming priority and interest in International Patent Application No. PCT / CN2019 / 106700, filed on 29 September 2019, and PCT / CN2019 / 108736, filed on 27 September 2019. For all purposes under the law, the entire disclosure of the aforementioned applications is incorporated by reference as part of the disclosure of this application.
[0002] Technical field This paper concerns video and image encoding and decoding techniques. [Background technology]
[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to rise. [Overview of the Initiative] [Means for solving the problem]
[0004] The disclosed techniques may be used by embodiments of video or image decoders or encoders for quantization parameter derivation in palette-mode coding and decoding.
[0005] In one exemplary aspect, a method for video processing is disclosed. The method includes the steps of determining that, for a conversion between a current block of video and a bitstream representation of the video, the current block is encoded using palette mode and escape symbol values are signaled in the bitstream representation, and that, based on the determination, the conversion is performed, wherein in the conversion, clipped quantization parameters for the current block are used, the palette mode encoding tool represents the current video block using a palette of representative color values, the escape values are used for a sample of the current video block that is encoded without using the representative color values, and the clipped quantization parameters used for the chroma components of the video are derived based on quantization parameters after a mapping operation of a quantization or dequantization process.
[0006] In another exemplary aspect, a method for video processing is disclosed. The method includes the steps of determining that, for a conversion between a current block of video and a bitstream representation of the video, the current block is encoded using palette mode and escape symbol values are signaled in the bitstream representation, and that, based on the determination, the conversion is performed, wherein, in the conversion, clipped quantization parameters for the current block are used, the palette mode encoding tool represents the current video block using a palette of representative color values, the escape values are used for a sample of the current video block that is encoded without using the representative color values, and the clipped quantization parameters used for the chroma components of the video are derived based on quantization parameters prior to a mapping operation of a quantization or dequantization process.
[0007] In yet another illustrative aspect, the method described above may be implemented by a video encoder device equipped with a processor.
[0008] In yet another exemplary aspect, the methods described above may be implemented by a video decoder device comprising a processor.
[0009] In yet another exemplary aspect, these methods may be embodied in the form of processor-executable instructions and stored in a computer-readable program medium.
[0010] These and other aspects are further described herein.
Brief Description of the Drawings
[0011] [Figure 1] It is a diagram showing an example of a block encoded in palette mode.
[0012] [Figure 2] An example of the use of a palette predictor for signaling palette entries is shown.
[0013] [Figure 3] Examples of horizontal and vertical cross-scans are shown.
[0014] [Figure 4] An exemplary encoding of a palette index is shown.
[0015] [Figure 5] A and B show examples of minimum chroma intra prediction units (SCIPUs).
[0016] [Figure 6] Examples of repeated palette entries in the local dual tree case are shown.
[0017] [Figure 7] Examples of left and upper blocks in the process of context derivation are shown.
[0018] [Figure 8]This is a block diagram of an example hardware platform used to implement the techniques described in this paper.
[0019] [Figure 9] This is a block diagram of an exemplary video processing system in which the disclosed techniques may be implemented.
[0020] [Figure 10] This block diagram shows a video encoding system according to some embodiments of the present disclosure.
[0021] [Figure 11] Block diagrams show encoders according to some embodiments of this disclosure.
[0022] [Figure 12] Block diagram of a decoder according to some embodiments of the present disclosure.
[0023] [Figure 13] A flowchart illustrating an exemplary method of video processing is shown.
[0024] [Figure 14] A flowchart illustrating another exemplary method of video processing is shown. [Modes for carrying out the invention]
[0025] This paper provides various techniques that can be used by image or video bitstream decoders to improve the quality of decompressed or decoded digital video or images. For brevity, the term “video” is used herein to include both sequences of pictures (traditionally called video) and individual images. Furthermore, video encoders may also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0026] Section headings are used in this paper for ease of understanding and do not limit embodiments and techniques to the corresponding sections. Therefore, embodiments from one section can be combined with embodiments from other sections.
[0027] 1. Overview This paper discusses video encoding 〔coding〕 Regarding technology. Specifically, palette coding. 〔coding〕 Index and escape symbol coding in 〔coding〕 This concerns existing video encodings like HEVC. 〔coding〕 Standards, or standards that will be finalized in the future (Multipurpose Video Coding) 〔coding〕 This may be applied to future video encoding. 〔coding〕 It may also be applicable to standards or video codecs.
[0028] 2. Background Video coding standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on hybrid video coding structures that utilize temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, a joint video expert team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), working on the VVC standard, which aims to reduce the bitrate by 50% compared to HEVC.
[0029] The latest version of the VVC draft, namely Versatile Video Coding (Draft 6), can be found at the following location: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 15_Gothenburg / wg11 / JVET-O2001-v14.zip
[0030] The latest VVC reference software (VTM) can be found at the following location: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0
[0031] 2.1 Palette Mode in HEVC Screen Content Encoding Extension (HEVC-SCC) 2.1.1 The concept of Palette Mode
[0032] The fundamental idea behind palette mode is that pixels within a CU are represented by a small set of representative color values. This set is called the palette. It is also possible to indicate samples outside the palette by signaling escape symbols followed by (potentially quantized) component values. These types of pixels are called escape pixels. Palette mode is shown in Figure 1. As shown in Figure 1, for each pixel with three color components (luma and two chroma components), an index to the palette is established, and the block can be reconstructed based on the established values in the palette.
[0033] 2.1.2 Encoding of Palette Items Palette predictors are maintained for encoding palette items. The maximum palette size and palette predictors are signaled in the SPS. In HEVC-SCC, palette_predictor_initializer_present_flag is introduced in the PPS. If this flag is 1, an item for initializing palette predictors is signaled in the bitstream. Palette predictors are initialized at the beginning of each CTU row, each slice, and each tile. Depending on the value of palette_predictor_initializer_present_flag, palette predictors are either reset to 0 or initialized using palette predictor initializer entries signaled in the PPS. In HEVC-SCC, a palette predictor initializer of size 0 allows for explicit disabling of palette predictor initialization at the PPS level.
[0034] For each item in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. This is shown in Figure 2. The reuse flag is transmitted using zero run-length coding. Then, the number of new palette items is signaled using exponential Golomb (EG) coding of order zero, i.e., EG-0. Finally, the component values for the new palette items are signaled.
[0035] 2.1.3 Encoding of Palette Indexes The palette index is encoded using horizontal and vertical transverse scanning as shown in Figure 3. The scanning order is explicitly signaled in the bitstream using palette_transpose_flag. For the remainder of the subsection, the scanning is assumed to be horizontal.
[0036] The palette index is encoded using two palette sample modes: 'COPY_LEFT' and 'COPY_ABOVE'. In 'COPY_LEFT' mode, the palette index is assigned to the decoded index. In 'COPY_ABOVE' mode, the palette index of the sample in the row above is copied. For both 'COPY_LEFT' and 'COPY_ABOVE' modes, a run value is signaled that specifies the number of subsequent samples to be encoded in the same mode.
[0037] In palette mode, the index value for an escape symbol is the number of palette items. Additionally, if an escape symbol is part of a run in 'COPY_LEFT' or 'COPY_ABOVE' mode, the escape component value is signaled for each escape symbol. The encoding of the palette index is shown in Figure 4.
[0038] This syntax order is achieved as follows: First, the number of index values for the CU is signaled. Subsequently, the actual index values for the entire CU are signaled using truncated binary coding. Both the number of indices and the index values are coded in bypass mode. This groups the bypass bins associated with the indices. Then, palette sample mode (if necessary) and run are signaled in an interleaved manner. Finally, the component escape values for the various escape symbols for the entire CU are grouped together and coded in bypass mode. The binarization of the escape symbols is third-order EG coding, i.e., EG-3.
[0039] The additional syntactic element `last_run_type_flag` is signaled after the index value has been signaled. This syntactic element eliminates the need to signal the run value corresponding to the last run in the block, along with the index number.
[0040] In HEVC-SCC, palette mode is also enabled for 4:2:2, 4:2:0, and monochromatic chroma formats. Signaling of palette items and palette indices is nearly identical for all chroma formats. For non-monochromatic formats, each palette item consists of three components. For monochromatic formats, each palette item consists of a single component. For subsampled chroma directions, chroma samples are associated with chroma sample indices that are divisible by 2. After reconstructing the palette index for CU, if a sample has only one associated component, only the first component of the palette item is used. The only difference in signaling is regarding escape component values. For each escape symbol, the number of escape component values that are signaled may vary depending on the number of components associated with that symbol.
[0041] Furthermore, Palette index coding involves an index adjustment process. When a Palette index is signaled, the left neighbor index or upper neighbor index must be different from the current index. Therefore, by eliminating one possibility, the range of the current Palette index can be reduced by 1. The index is then signaled in truncated binary (TB) binary format.
[0042] The text related to this section is shown as follows: Here, CurrPaletteIndex is the current palette index and adjustedRefPaletteIndex is the forecast index.
[0043] The variable PaletteIndexMap[xC][yC] specifies the palette index. This is the CurrentPaletteEntries This is an index to an array represented by . Array indices xC and yC specify the sample position (xC, yC) relative to the top-left luma sample of the picture. The value of PaletteIndexMap[xC][yC] is in the range of 0 to MaxPaletteIndex (inclusive). The variable adjustedRefPaletteIndex is derived as follows: adjustedRefPaletteIndex=MaxPaletteIndex+1 if(PaletteScanPos>0){ xcPrev=x0+TraverseScanOrder[log2CbWidth][log2bHeight][PaletteScanPos-1][0] ycPrev=y0+TraverseScanOrder[log2CbWidth][log2bHeight][PaletteScanPos-1][1] if(CopyAboveIndicesFlag[xcPrev][ycPrev]==0){ adjustedRefPaletteIndex=PaletteIndexMap[xcPrev][ycPrev]{ (7-157) else if(!palette_transpose_flag) adjustedRefPaletteIndex=PaletteIndexMap[xC][yC-1] else adjustedRefPaletteIndex=PaletteIndexMap[xC-1][yC] } }
[0044] If CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows: if(CurrPaletteIndex>=adjustedRefPaletteIndex) CurrPaletteIndex++
[0045] Furthermore, run-length elements in palette mode are context-encoded. The relevant context derivation process described in JVET-O2011-vE is shown below. The derivation process of ctxInc for the syntax element palette_run_prefix The inputs to this process are the bin index binIdx and the syntax elements copy_above_palette_indices_flag and palette_idx_idc. The output of this process is the variable ctxInc. The variable ctxInc is derived as follows: If copy_above_palette_indices_flag is equal to 0 and binIdx is equal to 0, then ctxInc is derived as follows: ctxInc=(palette_idx_idc<1)? 0:((palette_idx_idc<3)?1:2) (9-69) Otherwise, ctxInc is given by Table 1: [Table 1]
[0046] 2.2 Palette Mode in VVC 2.2.1 Palette in a Dual Tree In VVC, because a dual-tree coding structure is used for coding the intraslice, the chroma component and the two chroma components may have different palettes and palette indices. Furthermore, the two chroma components may share the same palette and palette indices.
[0047] 2.2.2 Palette as a separate mode In JVET-N0258 and the current VTM, the prediction mode for the coding unit can be MODE_INTRA, MODE_INTER, MODE_IBC, or MODE_PLT. The binary representation of the prediction mode is changed accordingly.
[0048] When IBC is turned off, on the I-tile, the first bin is used to indicate whether the current prediction mode is MODE_PLT. On the P / B tile, the first bin is used to indicate whether the current prediction mode is MODE_INTRA. If not, an additional bin is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTER. When IBC is turned on, on the I-tile, the first bin is used to indicate whether the current prediction mode is MODE_IBC. If not, a second bin is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. On the P / B tile, the first bin is used to indicate whether the current prediction mode is MODE_INTRA. If it is intra-mode, a second bin is used to indicate whether the current prediction mode is MODE_PLT or MODE_INTRA. If not, a second bin is used to indicate whether the current prediction mode is MODE_IBC or MODE_INTER.
[0049] The relevant text from JVET-O2001-vE is shown below.
[0050] Encoding Unit Syntax [Table 2]
[0051] 2.2.3 Palette Mode Syntax [Table 3] TIFF0007832105000004.tif220170TIFF0007832105000005.tif225170TIFF0007832105000006.tif218170TIFF0007832105000007.tif94170
[0052] 2.2.4 Semantics of Palette Mode In the following semantics, array indices x0 and y0 specify the position (x0, y0) of the top-left lumen sample in the coded block being considered, relative to the top-left lumen sample of the picture. Array indices xC and yC specify the position (xC, yC) of the sample relative to the top-left lumen sample of the picture. Array index startComp specifies the first color component in the current palette table. A startComp equal to 0 indicates the Y component, a startComp equal to 1 indicates the Cb component, and a startComp equal to 2 indicates the Cr component. numComps specifies the number of color components in the current palette table.
[0053] The predictor palette consists of palette items from previous coding units that are used to predict items in the current palette. The variable PredictorPaletteSize[startComp] specifies the size of the predictor palette for the first color component startComp in the current palette table. PredictorPaletteSize is derived as defined in Section 8.4.5.3. The variable PalettePredictorEntryReuseFlags[i] being equal to 1 indicates that the i-th item in the predictor palette will be reused in the current palette. PalettePredictorEntryReuseFlags[i] being equal to 0 indicates that the i-th item in the predictor palette is not an item in the current palette. All elements of the array PalettePredictorEntryReuseFlags[i] are initialized to 0.
[0054] palette_predictor_run is used to determine the number of zeros preceding a non-zero item in the array PalettePredictorEntryReuseFlags. The value of palette_predictor_run must be within the range of 0 to (PredictorPaletteSize - predictorEntryIdx) (inclusive) as a requirement for bitstream compatibility. Here, predictorEntryIdx corresponds to the current position in the array PalettePredictorEntryReuseFlags. The variable NumPredictedPaletteEntries specifies the number of items in the current palette to be reused from the predictor palette. The value of NumPredictedPaletteEntries must be within the range of 0 to palette_max_size (inclusive).
[0055] num_signalled_palette_entries specifies the number of items in the current palette that are explicitly signaled for the first color component, startComp, in the current palette table. If num_signalled_palette_entries does not exist, it is assumed to be equal to 0. The variable CurrentPaletteSize[startComp] specifies the size of the current palette for the first color component startComp in the current palette table, and is derived as follows: CurrentPaletteSize[startComp]=NumPredictedPaletteEntries+num_signalled_palette_entries (7-155) The value of CurrentPaletteSize[startComp] is within the range of palette_max_size (including both ends), from 0.
[0056] new_palette_entries[cIdx][i] specifies the value for the i-th signal-transmitted palette item for the color component cIdx. The variable PredictorPaletteEntries[cIdx][i] specifies the i-th element in the predictor palette for the color component cIdx. The variable CurrentPaletteEntries[cIdx][i] specifies the i-th element in the current palette for the color component cIdx, and is derived as follows: [Table 4]
[0057] A palette_escape_val_present_flag equal to 1 indicates that the current coding unit contains at least one escaped coded sample. A escape_val_present_flag equal to 0 indicates that there are no escaped coded samples in the current coding unit. If none exist, the value of palette_escape_val_present_flag is assumed to be equal to 1. The variable MaxPaletteIndex specifies the maximum possible palette index for the current encoding unit. The value of MaxPaletteIndex is CurrentPaletteSize[startComp]-1+palette_escape_val_presen_flag It is set to be equal to.
[0058] num_palette_indices_minus1 plus 1 is the number of palette indices that are explicitly signaled or inferred for the current block. If num_palette_indices_minus1 does not exist, it is inferred to be equal to 0. palette_idx_idc is an index instruction to the palette table, CurrentPaletteEntries. The value of palette_idx_idc is from 0 to MaxPaletteIndex (inclusive) for the first index in a block, and from 0 to (MaxPaletteIndex-1) (inclusive) for the remaining indices in a block. If palette_idx_idc does not exist, it is assumed to be equal to 0. The variable PaletteIndexIdc[i] stores the i-th palette_idx_idc that is explicitly signaled or estimated. All elements of the array PaletteIndexIdc[i] are initialized to 0.
[0059] A value of 1 for `copy_above_indices_for_final_run_flag` indicates that the palette indices for the last positions of the coding unit are copied from the palette indices in the row above if horizontal cross-sectional scanning is used, or from the palette indices in the column to the left if vertical cross-sectional scanning is used. A value of 0 for `copy_above_indices_for_final_run_flag` indicates that the palette indices for the last positions in the coding unit are copied from `PaletteIndexIdc[num_palette_indices_minus1]`. If copy_above_indices_for_final_run_flag does not exist, it is presumed to be equal to 0.
[0060] A palette_transpose_flag equal to 1 specifies that a vertical transverse scan is applied to scan the index for samples within the current coding unit. A palette_transpose_flag equal to 0 specifies that a horizontal transverse scan is applied to scan the index for samples within the current coding unit. If it does not exist, the value of palette_transpose_flag is assumed to be equal to 0. The array `TraverseScanOrder` specifies the scan order array for palette coding. If `palette_transpose_flag` is equal to 0, `TraverseScanOrder` is assigned the horizontal scan order `HorTravScanOrder`, and if `palette_transpose_flag` is equal to 1, `TraverseScanOrder` is assigned the vertical scan order `VerTravScanOrder`.
[0061] A copy_above_palette_indices_flag equal to 1 specifies that the palette index is equal to the palette index at the same position in the row above if horizontal cross-sectional scanning is used, or at the same position in the column to the left if vertical cross-sectional scanning is used. A copy_above_palette_indices_flag equal to 0 specifies that the palette index indication for the sample is encoded or inferred in the bitstream.
[0062] The variable CopyAboveIndicesFlag[xC][yC] being equal to 1 specifies that the palette index is copied from the palette index in the row above (horizontal scan) or the column to the left (vertical scan). A CopyAboveIndicesFlag[xC][yC] equal to 0 specifies that the palette index is either explicitly encoded or inferred in the bitstream. The array indices xC and yC specify the sample position (xC, yC) relative to the top-left luma sample of the picture. The value of PaletteIndexMap[xC][yC] must be within the range of 0 to (MaxPaletteIndex-1) (including both ends).
[0063] The variable PaletteIndexMap[xC][yC] specifies the palette index, which is the index to the array represented by CurrentPaletteEntries. The array indices xC and yC specify the sample position (xC, yC) relative to the top-left luma sample of the picture. The value of PaletteIndexMap[xC][yC] is within the range of 0 to MaxPaletteIndex (including both ends).
[0064] The variable adjustedRefPaletteIndex is derived as follows: [Table 5] If CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows: if(CurrPaletteIndex>=adjustedRefPaletteIndex) CurrPaletteIndex++ (7-158)
[0065] palette_run_prefix, if present, specifies the prefix portion used in the binary creation of PaletteRunMinus1.
[0066] The `palette_run_suffix` variable is used in the derivation of the variable `PaletteRunMinus1`. If it does not exist, the value of `palette_run_suffix` is assumed to be equal to 0.
[0067] If RunToEnd is equal to 0, the variable PaletteRunMinus1 is derived as follows: If PaletteMaxRunMinus1 is equal to 0, PaletteRunMinus1 will be set to equal to 0. Otherwise (if PaletteMaxRunMinus1 is greater than 0), the following applies: If palette_run_prefix is less than 2, the following applies: PaletteRunMinus1=palette_run_prefix (7-159) • Otherwise (palette_run_prefix is 2 or greater), the following applies: PrefixOffset=1<<(palette_run_prefix-1) PaletteRunMinus1=PrefixOffset+palette_run_suffix (7-160)
[0068] The variable PaletteRunMinus1 is used as follows: If CopyAboveIndicesFlag[xC][yC] is equal to 0, PaletteRunMinus1 specifies the number of consecutive positions with the same palette index minus 1. Otherwise, if palette_transpose_flag is equal to 0, PaletteRunMinus1 specifies the number of consecutive positions minus 1 that have the same palette index as the corresponding position in the row above. Otherwise, PaletteRunMinus1 specifies the number of consecutive positions minus 1 that have the same palette index used in the corresponding position in the left column. If RunToEnd is equal to 0, the variable PaletteMaxRunMinus1 represents the maximum possible value for PaletteRunMinus1, and the value of PaletteMaxRunMinus1 being 0 or greater is a requirement for bitstream compatibility.
[0069] palette_escape_val specifies the quantized escape-encoded sample value for a particular component. The variable PaletteEscapeVal[cIdx][xC][yC] specifies the escape value of a sample where PaletteIndexMap[xC][yC] is equal to MaxPaletteIndex and palette_escape_val_present_flag is equal to 1. The array index cIdx specifies the color component. The array indices xC and yC specify the sample's position (xC, yC) relative to the top-left lumen sample of the picture. The bitstream compatibility requirement is that for cIdx equal to 0, PaletteEscapeVal[cIdx][xC][yC] is within the range of 0 to (1 << (BitDepthY+1) - 1 (including both ends), and for cIdx not equal to 0, it is within the range of 0 to (1 << ((BitDepthC+1)) - 1 (including both ends).
[0070] 2.3 Local Dual Tree in VVC In typical hardware video encoders and decoders, if a picture has more small intrablocks, processing throughput decreases due to the dependence of sample processing data between neighboring intrablocks. Predictor generation for intrablocks requires reconstructed samples from the upper and left boundaries of neighboring blocks. Therefore, intra-predictions must be processed sequentially block by block.
[0071] In HEVC, the smallest CU is an 8x8 lumen sample. The lumen component of the smallest intraCU can be further divided into four 4x4 lumen intraprediction units (PUs), but the chromen component of the smallest intraCU cannot be further divided. Therefore, the worst-case hardware processing throughput occurs when processing 4x4 chromen intrablocks or 4x4 lumen intrablocks.
[0072] In VTM5.0, in a single coding tree, the chroma partition always follows the lumens, and the minimum intra-CU is 4x4 lumens samples, so the minimum chroma intra-CB is 2x2. Therefore, in VTM5.0, the minimum chroma intra-CB in a single coding tree is 2x2. The worst-case hardware processing throughput for VVC decoding is only 1 / 4 of that for HEVC decoding. Furthermore, the chroma intra-CB reconstruction process becomes much more complex than in the HEVC case after employing tools including a cross-component linear model (CCLM), a 4-tap interpolation filter, a position-dependent intra-prediction combination (PDPC), and combined inter-intra-prediction (CIIP). Achieving high processing throughput in hardware decoders is difficult. This section proposes a method to improve the worst-case hardware processing throughput.
[0073] The goal of this method is to restrict the partitioning of chromatic intra-CBs, thereby prohibiting chromatic intra-CBs smaller than 16 chromatic samples.
[0074] In a single coding tree, a SCIPU is defined as a coding tree node with a chroma block size greater than or equal to TH chroma samples and at least one child lumen block smaller than 4TH lumen samples, where TH is set to 16 in this paper. Each SCIPU is required to have either all CBs be inter, or all CBs be non-inter, i.e., intra or IBC. In the case of a non-inter SCIPU, it is further required that the chroma of the non-inter SCIPU is not further subdivided, and the lumen of the SCIPU is further subdivided. Thus, the minimum chroma-intra CB size is 16 chroma samples, and 2x2, 2x4, and 4x2 chroma CBs are removed. Furthermore, chroma scaling is not applied in the case of a non-inter SCIPU. Furthermore, if the lumen block is further subdivided and the chroma block is not subdivided, a local dual-tree coding structure is constructed.
[0075] Two examples of SCIPUs are shown in Figures 5A and 5B. In Figure 5A, one chroma CB and three luma CBs (4×8, 8×8, and 4×8 luma CBs) from an 8×4 chroma sample form one SCIPU because a ternary tree (TT) split from an 8×4 chroma sample results in a smaller chroma CB than 16 chroma samples. In Figure 5B, one chroma CB from a 4×4 chroma sample (the left side of the 8×4 chroma sample) and three luma CBs (8×4, 4×4, and 4×4 luma CBs) form one SCIPU, while the other chroma CB from a 4×4 sample (the right side of the 8×4 chroma sample) and two luma CBs (8×4 and 8×4 luma CBs) form one SCIPU. This is because a binary tree (BT) partition from 4x4 chroma samples results in a smaller chroma CB than 16 chroma samples.
[0076] In the proposed method, if the current slice is an I-slice, or if the current SCIPU has a 4x4 lumen partition within it after one more division, the type of SCIPU is presumed to be non-interface (because interface 4x4 is not allowed in VVC); otherwise, the type of SCIPU (interface or non-interface) is indicated by a signaled flag before parsing the CU in the SCIPU.
[0077] By applying the above method, the worst-case hardware processing throughput occurs when 4x4, 2x8, or 8x2 chroma blocks are processed instead of 2x2 chroma blocks. The worst-case hardware processing throughput is the same as for HEVC and four times that of VTM5.0.
[0078] 2.4 Conversion Skip (TS) As with HEVC, block residuals can be encoded in trans-skip mode. To avoid syntactic coding redundancy, the trans-skip flag is not signaled if the CU level MTS_CU_flag is not equal to zero. The block size limit for trans-skip is the same as for MTS in JEM4. This means that trans-skip can be applied to a CU if both the block width and height are 32 or less. Note that if LFNST or MIP is activated for the current CU, the implicit MTS trans is set to DCT2. Also, the implicit MTS can be enabled even if MTS is enabled for the inter-encoded block.
[0079] Furthermore, for the conversion skip block, the minimum allowable quantization parameter (QP) is defined as 6*(internalBitDepth-inputBitDepth)+4.
[0080] 2.5 Alternative Luma Half-Pixel Interpolation Filter JVET-N0309 proposes an alternative half-pixel interpolation filter.
[0081] The switching of the half-pixel lumer interpolation filter is dependent on the motion vector accuracy. In addition to the existing 1 / 4 pixel, full pixel, and 4-pixel AMVR modes, a new half-pixel accuracy AMVR mode is introduced. An alternative half-pixel lumer interpolation filter can only be selected when half-pixel motion vector accuracy is used.
[0082] For non-affine, non-merged intercoded CUs using half-pixel motion vector precision (i.e., half-pixel AMVR mode), switching between the HEVC / VVC half-pixel lumer interpolation filter and one or more alternative half-pixel interpolations is based on the value of a new syntactic element hpelIfIdx. The syntactic element hpelIfIdx is only transmitted in the case of half-pixel AMVR mode. In skip / merge modes using spatial merge candidates, the value of the syntactic element hpelIfIdx is inherited from the neighboring block.
[0083] 3. Technical problems solved by the technical solutions and embodiments described herein. 1. The current binary representation of escape symbols is not fixed-length, which may be suitable for sources with a uniform distribution. 2. Current palette coding designs perform an index adjustment process to eliminate potential redundancy, which can introduce parsing dependencies, for example, if escape value indices are incorrectly derived. 3. The reference index currently used to derive the index may require encoder constraints, which are not considered in the current design and are undesirable for codec design. 4. When local dual trees are enabled, the palette items of the previous block and the current block may have different numbers of color components. The method for handling such cases is unknown. 5. Local dual-tree and PLT could not be applied simultaneously because some palette items may be repeated when encoding from a single-tree region to a dual-tree region. An example is shown in Figure 6. 6. The chroma QP table for joint_cbcr mode may be constrained.
[0084] 4. List of Embodiments and Solutions The following list should be considered as examples to illustrate general concepts. These items should not be interpreted strictly. Furthermore, these items can be combined in any way.
[0085] The following examples can be applied to the pallet system in VVC and all other pallet-related systems.
[0086] Modulo(x,M) is defined as (x%M) if x is a positive integer, and as M-(-x)%M otherwise.
[0087] The following examples can be applied to the pallet system in VVC and all other pallet-related systems.
[0088] 1. Fixed-length coding may be applied to encode escape symbols. a. In one example, escape symbols may be transmitted via fixed-length binary conversion. b. In one example, escape symbols may be transmitted in fixed-length binary format using N bits. c. For example, the code length for signaling an escape symbol (e.g., N as mentioned in item 1.b) may depend on the internal bit depth. i. Alternatively, the code length for signaling escape symbols may depend on the input bit depth. ii. Alternatively, the code length for signaling escape symbols may depend on the difference between the internal bit depth and the input bit depth. iii. In one example, N is set to be equal to the input / internal bit depth. d. In one example, the code length for signaling escape symbols (e.g., N as mentioned in item 1.b) may depend on the quantization parameter, i.e., Qp. i. In one example, the code length for signaling an escape symbol may be a function of the quantization parameter, such that f(Qp) is denoted by the same parameter. 1. In one example, the function f may be defined as (internal bit depth - g(Qp)). 2. In one example, N may be set to (internal bit depth - max(16, (Qp-4) / 6)). 3. In one example, N may be set to (internal bit depth - max(QpPrimeTsMin, (Qp-4) / 6)), where qP is the decoded quantization parameter and QpPrimeTsMin is the minimum allowable quantization parameter for the transform skip mode. 4. Alternatively, the code length N may be set to max(A, internal bit depth - (Max(QpPrimeTsMin, Qp) - 4) / 6), where A is a non-negative integer value such as 0 or 1. e. In the above examples, N may be 0 or greater.
[0089] 2. It is proposed to disable the use of escape symbols within a single video unit (e.g., CU). a. Alternatively, the signaling of the presence of an escape symbol is skipped. b. In one example, whether to enable or disable the use of escape symbols may depend on the quantization parameters and / or bit depth. i. For example, if (internal bit depth - (Max(QpPrimeTsMin,Qp) - 4) / 6) is not greater than 0, the use of escape symbols may be disabled.
[0090] 3.3 Variable-length coding, excluding the next EG, may be applied to encode escape symbols. a. In one example, the binaryization of an escape symbol may be a truncated binary (TB) with input parameter K. b. In one example, the binaryization of an escape symbol may be K-th order EG, where K is not equal to 3. i. In one example, the binaryization of an escape symbol may be a zero-order EG. 1. Alternatively, in one example, the binaryization of escape symbols may be a first-order EG. 2. Alternatively, in one example, the binaryization of escape symbols may be a secondary EG. c. In the above examples, K may be an integer and may depend on the following: i.Messages transmitted in SPS / VPS / PPS / Picture Header / Slice Header / Tile Group Header / LCU Row / LCU Group / Brick. ii. Internal bit depth iii. Input bit depth iv. Difference between internal bit depth and input depth v. Current block dimensions vi. Current quantization parameters of the current block vii. Color format instructions (4:2:0, 4:4:4, RGB, or YUV, etc.) viii. Encoding structure (single tree or dual tree, etc.) ix. Color components (such as lumens and / or chroma components)
[0091] 4. Multiple binary encoding methods for encoding escape symbols may be applied to video units (e.g., sequences / pictures / slices / tiles / bricks / subpictures / CTU lines / CTU / CTB / CB / CU / subregions within a picture) and / or one or more values of escape symbols. a. In one example, how to select one of several binaryization methods may be signaled for one or more values of video units and / or escape symbols. b. In one example, how to select one of several binarison methods may be derived for one or more values of video units and / or escape symbols. c. In one example, two or more binary conversion methods may be applied to one or more values of a single video unit and / or escape symbol. i. In one example, an index or flag may be encoded / decoded to indicate the selected binary formatting method.
[0092] In the following terms, p may represent the symbolic value of the color component, bd may represent the bit depth (e.g., internal bit depth or input bit depth), ibd may represent the input bit depth, and Qp may represent the quantization parameter for the transform skip block or transform block. Furthermore, the QP for the lumen and chroma components may be different or the same. The bit depth may be associated with a given color component.
[0093] 5. How the quantization and / or dequantization process is applied may depend on whether the blocks are encoded in palette mode or not. a. In one example, the quantization and / or dequantization process for escape symbols may differ from that used for ordinary intra / intercoded blocks to which quantization is applied.
[0094] 6. The quantization and / or dequantization process for escape symbols may use bit shifts. a. In one example, a right bit shift may be used to quantize an escape symbol. i. In one example, the escape symbol may be transmitted as f(p,Qp), where p is the input symbol value (e.g., the input lumen / chroma sample value) and Qp is the derived quantization parameter for the corresponding color component. 1. In one example, the function f may be defined as p >> g(Qp). 2. In one example, the function f may be defined as (p + (1 << (g(QP) - 1)) >> g(Qp)). 3. In one example, the function f may be defined as (0, (1 << bd) - 1, (p + (1 << (g(QP) - 1))) > g(Qp)). ii. In one example, the escape symbol may be signaled as h(p). 1. In one example, the function h may be defined as p >> N. 2. In one example, the function h may be defined as (p + (1 << (N - 1))) >> N. 3. In one example, when cu_transquant_bypass_flag is equal to 1, N may be set to 0. 4. In one example, when cu_transquant_bypass_flag is equal to 1, N may be equal to (bd - ibd). Here, bd is the internal bit depth and ibd is the input bit depth. 5. In one example, the function h may be defined as clip(0, (1 << (bd - N) - 1, p >> N), where bd is the internal bit depth for the current color component. 6. In one example, the function h may be defined as clip(0, (1 << (bd - N) - 1, (p + (1 << (N - 1))) >> N), where bd is the internal bit depth for the current color component. 7. In the above example, N may be in the range of [0, (bd - 1)]. b. In one example, a left bit shift may be used to inverse - quantize the escape symbol. i. In one example, the escape symbol may be de - quantized as f(p, Qp). Here, p is the decoded escape symbol and Qp is the derived quantization parameter for the corresponding color component. 1. In one example, f may be defined as p << g(Qp). 2. In one example, f may be defined as (p << g(Qp))+(1 << (g(Qp) - 1)). ii. In one example, the escape symbol may be reconstructed as f(p, Qp). Here, p is the decoded escape symbol. 1. In one example, f may be defined as clip(0, (1 << bd) - 1, p << g(Qp)). 2. In one example, f may be defined as clip(0, (1 << bd) - 1, (p << g(Qp)) + (1 << (g(Qp) - 1))). iii. In one example, the escape symbol may be reconstructed as h(p). 1. In one example, the function h may be defined as p << N. 2. In one example, the function h may be defined as (p << N) + (1 << (N - 1)). 3. In one example, when cu_transquant_bypass_flag is equal to 1, N may be set to 0. 4. In one example, when cu_transquant_bypass_flag is equal to 1, N may be equal to (bd - ibd), where bd is the internal bit depth and ibd is the input bit depth. 5. In one example, N is set to (max(QpPrimeTsMin, qP) - 4) / 6, where qP is the decoded quantization parameter and QpPrimeTsMin is the minimum allowable quantization parameter for the transform skip mode. a) In the above example, when both luma and chroma have the transform skip mode, different minimum allowable quantization parameters for the transform skip mode may be applied for different color components. 6. Alternatively, for the above examples, N may be further clipped, such as min(bd - 1, N). 7. In the above example, N may be in the range of [0, (bd - 1)].
[0095] 7. When applying a left shift as dequantization, the reconstruction offset of the escape symbol p may depend on the bit depth information. a. In one example, it may depend on the difference between the internal bit depth and the input bit depth, i.e., deltaBD = internal bit depth - input bit depth. If bK is less than or equal to deltaBD, the reconstructed value is p< <Kであってもよい。 If cK is greater than deltaBD, the reconstructed value (p< <K)+(1<<(K-1))であってもよい。 If dK is less than or equal to T0 (for example, T0=2), the reconstructed value is p< <Kであってもよい。 If eK is greater than T1 (for example, T1=2), the reconstruction value is (p< <K)+(1<<(K-1))であってもよい。 f. In one example, T0 and T1 in items d and e may be signaled in the bitstream at sequence / picture / slice / tile / brick / subpicture levels, etc. g. In one example, the reconstruction value is (p< <K)+((1<<(K-1)> >deltaBD< <deltaBD))であってもよい。 h. In one example, the reconstruction value is ((p<<(K+1))+(1<<K))> >(deltaBD+1)< <deltaBD)であってもよい。 i. For example, deltaBD may be transmitted in the bitstream at the sequence / picture / slice / tile / brick / subpicture level, etc. j. In one example, which reconstruction value is used (for example, items b through e) may depend on the quantization parameters of the current block. k. In one example, which reconstruction value is used (for example, items b through e) may depend on the value of deltaBD. l. In one example, K may be set to g(Qp).
[0096] 8. In the above examples, the following may apply: a. In one example, escape symbols may be context-encoded. b. In one example, escape symbols may be bypass-encoded. c. In one example, g(Qp) may be defined as (Qp-4) / 6 or QP / 8. i. Alternatively, g(Qp) may also be defined as Qp / 6 or QP / 8. ii. Alternatively, g(Qp) may also be defined as max(16, Qp / 6). iii. Alternatively, g(Qp) may also be defined as max(16, (Qp - 4) / 6). iv. Alternatively, g(Qp) may also be defined as max((bd - ibd)*6 + 4, (Qp - 4) / 6). <000^562>v. Alternatively, g(Qp) may also be defined as max(M, (Qp - 4) / 6). 1. In one example, M may be signaled to the decoder. vi. Alternatively, g(Qp) may also be defined as max((M, Qp) - 4) / 6. 1. In one example, M may be indicated in the SPS. [[ID=^6]]2. In one example, the same M or different Ms may be applied to the luma component and the chroma component. 3. In one example, M may be equal to (bd - ibd)*6 + 4. vii. Alternatively, g(Qp) may also be defined as Qp / 6 or QP / 8. viii. Alternatively, g(Qp) may also be defined as (max(16, Qp) / 6). ix. Alternatively, g(Qp) may also be defined as (max(16, Qp) - 4) / 6. d. In one example, the value of g(Qp) may be in the range of [0, (bd - 1)]. e. In one example, the maximum function max(a, i) may also be defined as (i <= a? a : i). i. Alternatively, in one example, the maximum function max(a, i) may also be defined as (i < a? a : i). f. In one example, N may be an integer (e.g., 8 or 10) and may also depend on the following. i. Messages signaled in SPS / VPS / PPS / Picture Header / Slice Header / Tile Group Header / LCU row / group of LCUs / brick ii. Internal bit depth iii. Input bit depth iv. Difference between internal bit depth and input depth v. Current block dimensions vi. Current quantization parameters of the current block vii. Color format instructions (e.g., 4:2:0, 4:4:4, RGB, or YUV) viii. Encoding structure (single tree or dual tree, etc.) ix. Color components (such as lumens and / or chroma components) x. Slice / tile group type and / or picture type g. In one example, N may be transmitted to the decoder.
[0097] 9. Qp for escape values may be clipped. a. In one example, the minimum Qp applied to the escape value may be equal to min_qp_prime_ts_minus4. b. In one example, the minimum Qp applied to the escape value may be related to min_qp_prime_ts_minus4. i. In one example, the minimum Qp applied to an escape value may be equal to min_qp_prime_ts_minus4+4. c. In one example, the minimum Qp for each color component may be shown in the SPS / PPS / VPD / DPS / tile / slice header. d. In one example, the minimum Qp applied to the escape value may be (bd-ibd)*6+4, where bd is the internal bit depth and ibd is the input bit depth for a given color component. e. In one example, the above examples may be applied to a certain color component.
[0098] 10. In the above examples, the chroma Qp for escape values may be used before or after the mapping.
[0099] 11. It is suggested that the reference index should not be used when deriving the current palette index in palette mode. a. In one example, the palette index may be signaled directly without ruling out the possibility of a reference index (e.g., adjustedRefPaletteIndex). i. Alternatively, in one example, the encoder may be constrained to allow the reference index to always be different from the current index. In such a case, the palette index may be signaled by eliminating the possibility of the reference index being different. b. In one example, the binaryization of the palette index may be truncated binary (TB) with the maximum palette index used as the binaryization input parameter. c. In one example, the binary representation of the palette index may be of a fixed length. d. In one example, the binary representation of the palette index may be a K-th order EG. i. In one example, K may be an integer (for example, 1, 2, or 3) and may depend on the following: 1. Messages transmitted in SPS / VPS / PPS / Picture Header / Slice Header / Tile Group Header / LCU Row / LCU Group / Brick 2. Internal bit depth 3. Input bit depth 4. Difference between internal bit depth and input depth 5. Current block dimensions 6. Current quantization parameters of the current block 7. Color format instructions (e.g., 4:2:0, 4:4:4, RGB, or YUV) 8. Encoding structure (single tree or dual tree, etc.) 9. Color components (such as lumens and / or chroma components) e. For example, the above examples may only apply if the current block has at least one escape sample.
[0100] 12. The current pallet index may be transmitted independently of previous pallet indices. a. In one example, whether and / or how previous palette indices are used may depend on whether there are escape samples (one or more) in the current block.
[0101] 13. Deriving an index of non-escaped symbols from an index of escaped symbols may not be permitted. a. For example, if escape symbols are applied and the palette index is not equal to the index for the escape symbols, decoding those symbols as escape symbols may not be permitted.
[0102] 14. Deriving an index of escape symbols from an index of non-escape symbols may not be permitted. a. For example, if escape symbols are applied and the palette index is equal to the index for the escape symbols, decoding those symbols as non-escape symbols may not be permitted.
[0103] 15. The derived pallet index may be capped by the current pallet table size. a. For example, if the palette index is greater than MaxPaletteIndex, it may be modified to be equal to MaxPaletteIndex.
[0104] 16. The derived palette index may be capped by the current palette table size, except for the index for escape symbols. a. In one example, if escape symbols are not applied and the palette index is greater than MaxPaletteIndex, it may be modified to be equal to MaxPaletteIndex. b. In one example, if an escape symbol is applied and the palette index is greater than (MaxPaletteIndex-1), it may be modified to be equal to (MaxPaletteIndex-1).
[0105] 17. Indexes indicating escape symbols may not be permitted to be modified. a. In one example, an index equal to MaxPaletteIndex may always indicate an escape symbol if one exists in the current block. b. In one example, an index that is not equal to MaxPaletteIndex cannot be decoded as an index indicating an escape symbol.
[0106] 18. It is proposed to encode the difference between the reference index and the current index. a. In one example, a difference equal to 0 may not be allowed to be encoded. b. Alternatively, for the first index within a palette-encoded block, the index may be directly encoded.
[0107] 19. It is proposed to encode the modulo of the difference between the reference index (denoted as R) and the current index (denoted as C). a. In one example, I = Modulo(CR, MaxPaletteIndex) may be encoded. i. In one example, the index may be reconstructed as Modulo(I+R,MaxPaletteIndex). ii. In one example, a Modulo(CR,MaxPaletteIndex) equal to 0 may not be allowed in the bitstream. iii. In one example, a truncated binary code where cMax = MaxPaletteIndex may be used to encode the value. iv. Alternatively, for the first index within a palette-encoded block, the index may be directly encoded. b. In one example, I = Modulo(CR, MaxPaletteIndex) - 1 may be encoded. i. In one example, the index may be reconstructed as Modulo(I+1+R,MaxPaletteIndex). ii. In one example, Modulo(CR,MaxPaletteIndex)-1, which is less than 0, may not be allowed in the bitstream. iii. In one example, a truncated binary code where cMax = (MaxPaletteIndex - 1) may be used to encode the value I. iv. Alternatively, Modulo(CR,MaxPaletteIndex) may be encoded for the first index within the palette-encoded block. v. Alternatively, for the first index within a palette-encoded block, the index may be directly encoded.
[0108] 20. At the start of decoding a palette block, the reference index R may be set to equal to -1. a. Alternatively, the reference index R may be set to equal to 0.
[0109] 21. It is proposed to enable palette mode and local dual tree exclusively. a. For example, if palette mode is enabled, local dual trees may not be allowed. i. Alternatively, for example, palette mode may not be allowed if local dual tree is enabled. b. For example, local dual-tree may not be enabled for certain color formats such as 4:4:4. c. In one example, palette mode may not be allowed if the encoding tree is MODE_TYPE_INTRA.
[0110] 22. When a local dual tree is applied, it is proposed to remove repeated palette items in the palette prediction table. a. In one example, the palette prediction table may be reset when a local dual tree is applied. i. Alternatively, in one example, the decoder may check all palette items in the prediction table when a local dual tree is applied and remove any repeated items. ii. Alternatively, in one example, the encoder may add a constraint that considers two palette items to be different if the three components of the item are different.
[0111] 23. If the current palette item has a different number of color components than the items in the palette prediction table, the palette prediction table may not be allowed to be used. a. For example, the reuse flag may be marked as true for all items in the palette prediction table, but it may not be used for the current block if the current palette item has a different number of color components than predicted. b. For example, the reuse flag for all items in the palette prediction table may be marked as false if the current palette item has a different number of color components than predicted.
[0112] 24. If the prediction table and the current palette table have different color components (one or more), the palette prediction table may not be permitted to be used. a. For example, the reuse flag may be marked as true for all items in the palette prediction table, but it may not be used for the current block if the prediction table and the current palette table have different color components. b. For example, the reuse flag for all items in the palette prediction table may be marked as false if the prediction table and the current palette table have different color components.
[0113] 25. Escape symbols may be predictively encoded, for example, based on previously encoded escape symbols. a. In one example, the escape symbol for a component may be predicted by the encoded value in the same color component. i. In one example, an escape symbol may use a previously encoded escape symbol in the same component as a predictor, and the residual between them may be transmitted in the signal. ii. Alternatively, the escape symbol may use the Kth previously encoded escape symbol in the same component as a predictor, and the residual between them may be transmitted. iii. Alternatively, an escape symbol may be predicted from multiple (e.g., K) encoded escape symbols in the same component. 1. In one example, K may be an integer (for example, 1, 2, or 3) and may depend on the following: a) Messages transmitted in SPS / VPS / PPS / Picture Header / Slice Header / Tile Group Header / LCU Row / LCU Group / Brick b) Internal bit depth c) Input bit depth d) Difference between internal bit depth and input depth e) Current block dimensions f) Current quantization parameters of the current block g) Specify the color format (e.g., 4:2:0, 4:4:4, RGB, or YUV) h) Encoding structure (single tree or dual tree, etc.) i) Color components (such as lumens and / or chromens) b. In one example, the escape symbol for one component may be predicted by the encoded value of another component. c. In one example, a pixel may have multiple color components, and if the pixel is treated as an escape symbol, the value of one component may be predicted by the sample values of the other components. i. In one example, the U component of an escape symbol may be predicted by the V component of that symbol. d. In one example, the above methods may be applied only to a certain color component (for example, to a lumen component or a chroma component), or under certain conditions based on encoded information, etc.
[0114] 26. The context for run-length encoding in palette mode may depend on a palette index for indexing palette items. a. For example, the palette index after the index adjustment process in the decoder (described in Section 2.1.3) may be used to derive the context for the prefix of the length element (e.g., palette_run_prefix). b. Alternatively, in one example, I as defined in item 13 may replace the palette index to derive a context for the length element prefix (e.g., palette_run_prefix).
[0115] 27. It is proposed to align the positions of the left neighbor block and / or upper neighbor block used in the derivation process for the quantization parameter predictor with the positions of the neighboring left block and / or upper neighbor block used in the mode / MV (e.g., MPM) derivation. a. The positions of the left neighbor block and / or upper neighbor block used in the derivation process for the quantization parameters may be aligned with the positions used in the merge / AMVP candidate list derivation process. b. In one example, the positions of the neighboring left block and / or upper block used in the derivation process for the quantization parameters may be the adjacent left / upper blocks shown in Figure 7.
[0116] 28. Block-level QP differences may be sent regardless of whether an escape sample is present in the current block. a. For example, whether and / or how to transmit block-level QP differences can depend on the blocks encoded in a mode other than palette. b. For example, block-level QP differences may not always be sent for a given pallet block. c. For example, if the block width is greater than the threshold, a QP difference at the block level may be sent for the pallet block. d. In one example, if the block height is greater than the threshold, the QP difference of the block level may be sent for the pallet block. e. For example, if the block size is greater than the threshold, a QP difference at the block level may be sent for the pallet block. f. In one example, the above examples may apply only to rumor blocks or chroma blocks.
[0117] 29. One or more of the coded block flags (CBFs) for a palette block (e.g., cbf_luma, cbf_cb, cbf_cr) may be set to 1. a. In one example, the CBF for a pallet block may always be set to equal 1. b. One or more of the CBFs for a palette block may depend on whether an escape pixel exists in the current block. i. For example, if a palette block has an escape sample, its cbf may be set to 1. ii. Alternatively, if a palette block does not have escape samples, its cbf may be set to 0. c. Alternatively, when accessing a neighboring palette-encoded block, it may be treated as an intra-encoded block with a CBF equal to 1.
[0118] 30. Luma and / or chroma QP applied to a pallet block, and the QP derived for that block (e.g., Qp in the JVET-O2001-vE specification) Y or Qp' Y The difference between ) and may be set to be equal to a fixed value for each palette block. a. In one example, the lumens and / or chroma QP offsets may be set to 0. b. In one example, the chroma QP offsets for Cb and Cr may be different. c. In one example, the lumens QP offset and the chromens QP offset may be different. d. In one example, the chroma QP offset(s) may be indicated in the DPS / VPS / SPS / PPS / slice / brick / tile header.
[0119] 31. Num PltIdx The number of palette indices that are explicitly signaled or inferred for the current block, as shown by (e.g., num_palette_indices_minus1+1), may be constrained to K or greater. a. In one example, K may be determined based on the current palette size, escape flags, and / or other information about the palette-encoded block. Let S be the current palette size of the current block, and E be the value of the escape presence flag (e.g., palette_escape_val_present_flag). Let BlkS be the current block size. i. In one example, K may be set equal to S. ii. Alternatively, in one example, K may be set equal to S+E. iii. Alternatively, in one example, K may be set equal to (the number of predicted palette items + the number of signalled palette items + palette_escape_val_present_flag) (e.g., NumPredictedPaletteEntries + num_signalled_palette_entries + palette_escape_val_present_flag). b. In one example, instead of num_palette_indices_minus1, (Num PltIdx -K) may be signalled / parsed. i. Alternatively, further, it may be signalled only if (S+E) is not less than 1. ii. In one example, the value of (Num PltIdx -K) may be signalled in a binarization method where the binary string may have a prefix (e.g., truncated unary) and / or a suffix with a m-th EG code. iii. In one example, the value of (Num PltIdx -K) may be signalled in a truncated binary binarization method. iv. In one example, the value of (Num PltIdx -K) may be signalled in a truncated unary binarization method. v. In one example, the value of (Num PltIdx -K) may be signalled in a m-th EG binarization method. c. In one example, the compliant bitstream satisfies that Num PltIdx is greater than or equal to K. d. In one example, the compatible bitstream is Num PltIdx The condition is met that is less than or equal to K'. i. In one example, K' is set to (block width * block height).
[0120] 32. Whether and / or how the above methods are applied may be based on the following: a. Video content (for example, screen content or natural content) b. Messages transmitted in DPS / SPS / VPS / PPS / APS / Picture Header / Slice Header / Tile Group Header / Largest Coding Unit (LCU) / Coding Unit (CU) / LCU Row / LCU Group / TU / PU Block / Video Coding Unit c. Location of CU / PU / TU / block / video encoding unit d. Block dimensions of the current block and / or neighboring blocks e. Block shape of the current block and / or neighboring blocks f. Color format instructions (e.g., 4:2:0, 4:4:4, RGB, or YUV) g. Encoded tree structure (dual tree or single tree, etc.) h. Slice / tile group type and / or picture type i. Color components (for example, applicable only to the lumens and / or chroma components) j. Temporal Layer ID k. Profile / level / tier of the standard l. Whether the current block has one escape sample or not i. In one example, the above methods may only be applicable if the current block has at least one escape sample. m. Whether the current block is encoded in lossless mode (e.g., cu_transquant_bypass_flag) i. In one example, the above methods assume that the current block is encoded in lossless mode. do not have This may only apply in certain cases. n. Whether lossless coding is enabled or not (e.g., transquant_bypass_enabled, cu_transquant_bypass_flag) i. In one example, the above methods are only applicable when lossless coding is disabled.
[0121] BDPCM related 34. If a single block is encoded in BDPCM and is divided into multiple transform blocks or subblocks, residual prediction may be performed at the block level, and residual signaling may be performed at the subblock / transform block level. a. Alternatively, the reconfiguration of one subblock is not permitted during the reconfiguration process of another subblock. b. Alternatively, residual prediction and residual signal transmission may be performed at the subblock / transform block level. i. In this way, the reconfiguration of one subblock can be utilized in the reconfiguration process of another subblock.
[0122] ChromaQP Table related 44. For a given index, the chroma QP table value for joint_cb_cr mode can be constrained by both the chroma QP table value for Cb and the chroma QP table value for Cr. c. For example, the value of the chroma QP table for the joint_cb_cr mode may be constrained to the range between the value of the chroma QP table for Cb and the value of the chroma QP table for Cr (including both ends).
[0123] Unblocking related 45. The comparison of MV in deblocking may depend on whether an alternative half-pixel interpolation filter is used (for example, as indicated by hpelIfIdx in the JVET-O2001-vE specification). d. In one example, blocks using different interpolation filters may be treated as having different MVs. e. In one example, if an alternative half-pixel interpolation filter is involved, a certain offset may be added to the MV difference for deblocking comparison.
[0124] 5. Embodiments The embodiments are based on JVET-O2001-vE. Newly added text is enclosed in double brackets. For example, {{a}} indicates that "a" has been added. Deleted text is enclosed in double brackets. For example, [[b]] indicates that "b" has been deleted.
[0125] 5.1 Embodiment #1 Decoding process for palette mode The inputs to this process are as follows: • The position (xCb, yCb) that specifies the top-left rumor sample of the current block relative to the top-left rumor sample of the current picture. • Variable startComp specifies the first color component in the palette table. • Variable cIdx specifies the color component of the current block. • Two variables, nCbW and nCbH, specify the width and height of the current block, respectively. The output of this process is an array recSamples[x][y] specifying the reconstructed sample values for the block, where x=0..nCbW-1 and y=0..nCbH-1.
[0126] Depending on the value of cIdx, the variables nSubWidth and nSubHeight are derived as follows: If cIdx is equal to 0, nSubWidth is set to 1 and nSubHeight is set to 1. Otherwise, nSubWidth is set to SubWidthC and nSubHeight is set to SubHeightC.
[0127] The (nCbW × nCbH) block of the reconstructed sample sequence recSamples at position (xCb, yCb) is represented by recSamples[x][y](x=0..nCTbW-1 and y=0..nCbH-1), and the values of recSamples[x][y] for each x in the range from 0 to nCTbW-1 (including both ends) and each y in the range from 0 to nCbH-1 (including both ends) are derived as follows: The variables xL and yL are derived as follows: xL=palette_transpose_flag ? x*nSubHeight : x*nSubWidth (8-268) yL=palette_transpose_flag ? y*nSubWidth : y*nSubHeight (8-269) The variable bIsEscapeSample is derived as follows: If PaletteIndexMap[xCb+xL][yCb+yL] is equal to MaxPaletteIndex and palette_escape_val_present_flag is equal to 1, then bIsEscapeSample is set to equal to 1. Otherwise, bIsEscapeSample is set to equal to 0. If bIsEscapeSample is equal to 0, the following applies: recSamples[x][y]=CurrentPaletteEntries[cIdx][PaletteIndexMap[xCb+xL][yCb+yL]] (8-270) Otherwise, if cu_transquant_bypass_flag is equal to 1, the following applies: recSamples[x][y]=PaletteEscapeVal[cIdx][xCb+xL][yCb+yL] (8-271) Otherwise (bIsEscapeSample is equal to 1 and cu_transquant_bypass_flag is equal to 0), the following ordered steps apply:
[0128] 1. The quantization parameter qP is derived as follows: If cIdx is equal to 0, qP = Max(0, Qp'Y) (8-272) Otherwise, if cIdx is equal to 1, qP = Max(0, Qp'Cb) (8-273) Otherwise (if cIdx is equal to 2), qP = Max(0, Qp'Cr) (8-274) 2. The variable bitDepth is derived as follows: bitDepth=(cIdx==0) ? BitDepth Y : BitDepth C (8-275) 3.[[The list levelScale[] is specified as levelScale[k]={40,45,51,57,64,72}, where k=0..5]] 4. The following applies: [[tmpVal=(PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]*levelScale[qP%6])<<(qP / 6)+32)>>6 (8-276)]] {{T is set to be equal to (internal_bit_depth - input_bit_depth) for component cIdx. Nbits = max(T, (qP-4) / 6) If Nbits is equal to T, recSamples[x][y]=PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]< <Nbits • Otherwise, recSamples[x][y]=(PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]< <Nbits)+(1<<(Nbits-1)}} [[recSamples[x][y]=Clip3(0,(1< <bitDepth)-1,tmpVal) (8-277)]]
[0129] The following conditions: cIdx is equal to 0 and numComps is equal to 1; cIdx is equal to 2 If either of the following is true, The variable PredictorPaletteSize[startComp] and the array PredictorPaletteEntries are derived or modified as follows: [Table 6] The value of PredictorPaletteSize[startComp] must be within the range of 0 to PaletteMaxPredictorSize (including both ends) as a requirement for bitstream compatibility.
[0130] 5.2 Embodiment #2 This embodiment describes the derivation of the pallet index. Palette coding semantics The variable [[adjustedRefPaletteIndex] is derived as follows: [Table 7] If CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows: if(CurrPaletteIndex>=adjustedRefPaletteIndex) CurrPaletteIndex++]]
[0131] Binary conversion process for palette_idx_idc The inputs to this process are the request for binarization of the syntax element palette_idx_idc and the variable MaxPaletteIndex. The output of this process is the binarization of the said syntax element. The variable cMax is derived as follows: · If this process is called for the first time for the current block, cMax is set equal to MaxPaletteIndex. · Otherwise (if this process has not been called for the first time for the current block), cMax is set equal to MaxPaletteIndex minus 1. The binarization of palette_idx_idc is derived by calling the TB binarization process defined in section 9.3.3.4 with cMax.
[0132] 5.3 Embodiment #3 [Table 8]
[0133] 8.4.5.3 Decoding process for palette mode The inputs to this process are as follows: · The position (xCb, yCb) specifying the top-left luma sample of the current block with respect to the top-left luma sample of the current picture · The variable startComp specifying the first color component in the palette table · The variable cIdx specifying the color component of the current block, · Two variables nCbW and nCbH respectively specifying the width and height of the current block The output of this process is an array recSamples[x][y] specifying the reconstructed sample values for the block, where x = 0..nCbW-1, y = 0..nCbH-1
[0134] Depending on the value of cIdx, the variables nSubWidth and nSubHeight are derived as follows: ... Otherwise (bIsEscapeSample is equal to 1 and cu_transquant_bypass_flag is equal to 0), the following ordered steps apply:
[0135] 5. The quantization parameter qP is derived as follows: If cIdx is equal to 0, qP = Max(0, Qp'Y) (8-272) Otherwise, if cIdx is equal to 1, qP = Max(0, Qp'Cb) (8-273) Otherwise (if cIdx is equal to 2), qP = Max(0, Qp'Cr) (8-274) 6. The variable bitDepth is derived as follows: bitDepth=(cIdx==0) ? BitDepth Y : BitDepth C (8-275) 3.[[The list levelScale[] is specified as levelScale[k]={40,45,51,57,64,72}, where k=0..5]] 4. The following applies: [[tmpVal=(PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]*levelScale[qP%6])<<(qP / 6)+32)>>6 (8-276)]] {{shift=(max(QpPrimeTsMin,qP)-4) / 6 tmpVal=(PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]< <shift)}} recSamples[x][y]=Clip3(0,(1< <bitDepth)-1,tmpVal) (8-277)
[0136] 5.4 Embodiment #4 A copy_above_palette_indices_flag equal to 1 specifies that the palette index is equal to the palette index at the same position in the row above if horizontal cross-sectional scanning is used, or at the same position in the column to the left if vertical cross-sectional scanning is used. A copy_above_palette_indices_flag equal to 0 specifies that the palette index indication for the sample is encoded or inferred in the bitstream. ... The variable adjustedRefPaletteIndex is derived as follows: [Table 9] If CopyAboveIndicesFlag[xC][yC] is equal to 0, the variable CurrPaletteIndex is derived as follows: if(CurrPaletteIndex>=adjustedRefPaletteIndex) CurrPaletteIndex++ (7-158)
[0137] 5.5 Embodiment #5 [Table 10] 8.4.5.3 Decoding process for palette mode The inputs to this process are as follows: • The position (xCb, yCb) that specifies the top-left rumor sample of the current block relative to the top-left rumor sample of the current picture. • Variable startComp specifies the first color component in the palette table. • Variable cIdx specifies the color component of the current block. • Two variables, nCbW and nCbH, specify the width and height of the current block, respectively. The output of this process is an array recSamples[x][y] that specifies the reconstructed sample values for the blocks, where x = 0..nCbW-1, y = 0..nCbH-1
[0138] Depending on the value of cIdx, the variables nSubWidth and nSubHeight are derived as follows: …… · Otherwise (bIsEscapeSample equals 1 and cu_transquant_bypass_flag equals 0), the following ordered steps apply:
[0139] 9. The quantization parameter qP is derived as follows: · If cIdx equals 0, qP = Max(0, Qp'Y) (8-272) · Otherwise, if cIdx equals 1, qP = Max(0, Qp'Cb) (8-273) · Otherwise (cIdx equals 2), qP = Max(0, Qp'Cr) (8-274) 10. The variable bitDepth is derived as follows: bitDepth = (cIdx == 0)? BitDepth Y : BitDepth C (8-275) 11. [The list levelScale[] is specified as levelScale[k] = {40, 45, 51, 57, 64, 72}. Here, k = 0..5] 12. The following applies: [[tmpVal=(PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]*levelScale[qP%6])<<(qP / 6)+32)>>6 (8-276)]] {{shift=min(bitDepth-1,(max(QpPrimeTsMin,qP)-4) / 6) tmpVal=(PaletteEscapeVal[cIdx][xCb+xL][yCb+yL]< <shift)}} recSamples[x][y]=Clip3(0,(1< <bitDepth)-1,tmpVal) (8-277)
[0140] 5.6 Embodiment #6 This embodiment demonstrates a design for skipping conversion shifts for conversion skipping and is based on JVET-O2001-vE.
[0141] 8.7.2 Scaling and Transformation Process The inputs to this process are as follows: • The rumor position (xTbY, yTbY) specifies the top-left sample of the current rumor transformation block relative to the top-left rumor sample of the current picture. • Variable cIdx specifies the color component of the current block. • Variable nTbW specifies the width of the conversion block. • Transformation block: Variable nTbH specifies the height. The output of this process is the residual sample (nTbW)×(nTbH) sequence resSamples[x][y], where x=0..nTbW-1 and y=0..nTbH-1. The variables bitDepth, bdShift, and tsShift are derived as follows: bitDepth=(cIdx==0) ? BitDepth Y : BitDepth C (8-942) bdShift=Max(20-bitDepth,0) (8-943) [[tsShift=5+((Log2(nTbW)+Log2(nTbH)) / 2) (8-944) The variable codedCIdx is derived as follows: If cIdx is equal to 0, or if TuCResMode[xTbY][yTbY] is equal to 0, then codedCIdx is set to be equal to cIdx. Otherwise, if TuCResMode[xTbY][yTbY] is equal to 1 or 2, codedCIdx is set to equal to 1. Otherwise, codedCIdx is set to equal 2. The variable cSign is set to equal to (1-2*slice_joint_cbcr_sign_flag). The residual sample (nTbW)×(nTbH) sequence resSample is derived as follows:
[0142] 1. The scaling process for the conversion coefficients specified in Section 8.7.3 is invoked with the conversion block position (xTbY, yTbY), conversion block width nTbW, conversion block height nTbH, color component variable cIdx set to equal codedCIdx, and the current color component bit depth bitDepth as input, and the output is an (nTbW) × (nTbH) array of the scaled conversion coefficients d. 2. The (nTbW)×(nTbH) sequence of the residual sample r is derived as follows: If [[transform_skip_flag[xTbY][yTbY] is equal to 1 and cIdx is equal to 0, then with x=0..nTbW-1 and y=0..nTbH-1, the residual sample array values r[x][y] are derived as follows: r[x][y]=d[x][y]< <tsShift (8-945)]] ·[[If transform_skip_flag[xTbY][yTbY] is equal to 0, or if cIdx is not equal to 0]]The transformation process for scaled transformation coefficients as defined in Section 8.7.4.1 is invoked with the transformation block position (xTbY, yTbY), transformation block width nTbW and transformation block height nTbH, color component variable cIdx and a (nTbW) × (nTbH) array of scaled transformation coefficients d as input, and the output is a (nTbW) × (nTbH) array of residual samples r. 3. Assuming x = 0..nTbW-1 and y = 0..nTbH-1, the intermediate residual sample res[x][y] is derived as follows: ·If {{transform_skip_flag[xTbY][yTbY] is equal to 1 and cIdx is equal to 0, then the following applies: res[x][y]=d[x][y]}} ·{{Otherwise (transform_skip_flag[xTbY][yTbY] is equal to 0, or cIdx is not equal to 0), the following applies:}} res[x][y]=(r[x][y]+(1<<(bdShift-1)))>>bdShift (8-946) 4. Assuming x = 0..nTbW-1 and y = 0..nTbH-1, the residual samples resSamples[x][y] are derived as follows: If cIdx is equal to codedCIdx, the following applies: resSamples[x][y]=res[x][y] (8-947) Otherwise, if TuCResMode[xTbY][yTbY] is equal to 2, the following applies: resSamples[x][y]=cSign*res[x][y] (8-948) Otherwise, the following applies: resSamples[x][y]=(cSign*res[x][y])>>1
[0143] 8.7.3 Scaling process for conversion coefficients ... The variable rectNonTsFlag is derived as follows: rect[[NonTs]]Flag=(((Log2(nTbW)+Log2(nTbH))&1)==1[[&&]] (8-955) [[transform_skip_flag[xTbY][yTbY]=]]=0) The variables bdShift, rectNorm, and bdOffset are derived as follows: ·If {{transform_skip_flag[xTbY][yTbY] is equal to 1 and cIdx is equal to 0, then the following applies: bdshift=10}} • {{Otherwise, the following applies:}} bdShift=bitDepth+((rect[[NonTs]]Flag ? 1 : 0)+ (Log2(nTbW)+Log2(nTbH)) / 2)-5+dep_quant_enabled_flag (8-956) dOffset=(1<<bdShift)> >1 (8-957) The list levelScale[][] is specified as levelScale[j][k]={{40,45,51,57,64,72},{57,64,72,80,90,102}, where j=0..1 and k=0..5. The (nTbW)×(nTbH) array dz is set to be equal to the (nTbW)×(nTbH) array TransCoeffLevel[xTbY][yTbY][cIdx]. Given x=0..nTbW-1 and y=0..nTbH-1, the following applies to the derivation of the scaled transformation coefficients d[x][y]: The intermediate scaling factor m[x][y] is derived as follows: If one or more of the following conditions are true, then m[x][y] is set to equal 16: • sps_scaling_list_enabled_flag is equal to 0. • transform_skip_flag[xTbY][yTbY] is equal to 1. Otherwise, the following applies: m[x][y]=ScalingFactor[Log2(nTbW)][Log2(nTbH)][matrixId][x][y] Here, matrixId is specified in Table 7 (8-958). The scaling factor ls[x][y] is derived as follows: If dep_quant_enabled_flag is equal to 1, the following applies: ls[x][y]=(m[x][y]*levelScale[rect[[NonTs]]Flag][(qP+1)%6])<<((qP+1) / 6) (8-960) Otherwise (dep_quant_enabled_flag is equal to 0), the following applies: ls[x][y]=(m[x][y]*levelScale[rect[[NonTs]]Flag][qP%6])<<(qP / 6) (8-961) If BdpcmFlag[xTbY][yYbY] is equal to 1, then dz[x][y] is modified as follows: If BdpcmDir[xTbY][yYbY] is equal to 0 and x is greater than 0, then the following applies: dz[x][y]=Clip3(CoeffMin,CoeffMax,dz[x-1][y]+dz[x][y]) (8-961) Otherwise, if BdpcmDir[xTbY][yYbY] is equal to 1 and y is greater than 0, then the following applies: dz[x][y]=Clip3(CoeffMin,CoeffMax,dz[x][y-1]+dz[x][y]) (8-962) The value dnc[x][y] is derived as follows: dnc[x][y]=(dz[x][y]*ls[x][y]+bdOffset)>>bdShift (8-963) The scaled transformation coefficients d[x][y] are derived as follows: d[x][y]=Clip3(CoeffMin,CoeffMax,dnc[x][y]) (8-964)
[0144] Figure 8 is a block diagram of a video processing device 800. The device 800 may be used to implement one or more of the methods described herein. The device 800 may be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 800 may include one or more processors 802, one or more memories 804, and video processing hardware 806. The processor 802 may be configured to implement one or more of the methods described herein. The memory(s) 804 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 806 may be used in hardware circuitry to implement some of the techniques described herein. In some implementations, the hardware 806 may be part or all of the processor 802, for example, a graphics processor.
[0145] Some embodiments of the disclosed technology include making a judgment or decision to enable a video processing tool or mode. In one example, if a video processing tool or mode is enabled, the encoder uses or implements that tool or mode in processing blocks of video, but does not necessarily have to modify the resulting bitstream based on the use of that tool or mode. That is, the conversion from blocks of video to a bitstream representation of video uses that video processing tool or mode when enabled based on the judgment or decision. In another example, if a video processing tool or mode is enabled, the decoder processes the bitstream using the knowledge that the bitstream has been modified based on that video processing tool or mode. That is, the conversion from a bitstream representation of video to blocks of video is performed using the video processing tool or mode enabled based on the judgment or decision.
[0146] Some embodiments of the disclosed technology include making judgments or decisions to disable video processing tools or modes. In one example, if a video processing tool or mode is disabled, the encoder does not use that tool or mode when converting blocks of video into a bitstream representation of the video. In another example, if a video processing tool or mode is disabled, the decoder processes the bitstream using the knowledge that the bitstream has not been modified using a video processing tool or mode that was enabled based on the judgment or decision.
[0147] Figure 9 is a block diagram showing an exemplary video processing system 900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 900. System 900 may include an input 902 for receiving video content. The video content may be received in raw or uncompressed format, for example, in the form of 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet and passive optical networks (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0148] System 900 may include an encoding component 904 that can implement the various encoding or encoding methods described in this paper. The encoding component 904 may reduce the average bitrate of the video from input 902 to the output of the encoding component 904 to produce an encoded representation of the video. Thus, the encoding technique is sometimes called video compression or video transcoding technique. The output of the encoding component 904 may be stored or transmitted via connected communication, as represented by component 906. The stored or transmitted bitstream (or encoded) representation of the video received at input 902 may be used by component 908 to generate pixel values or a displayable video to be sent to the display interface 910. The process of generating a video that can be seen by the user from the bitstream representation is sometimes called video decompression. Furthermore, certain video processing operations are referred to as “encoding” operations or tools, but it will be understood that encoding tools or operations are used in the encoder, and corresponding decoding tools or operations that invert the result of encoding are performed in the decoder.
[0149] Examples of peripheral bus interfaces or display interfaces include Universal Serial Bus (USB), High-Definition Multimedia Interface (MDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, and IDE interfaces. The techniques described in this paper can be implemented in a variety of electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0150] Figure 10 is a block diagram showing an exemplary video coding system 100 that can utilize the techniques described herein.
[0151] As shown in Figure 10, the video encoding system 100 may include a source device 110 and a destination device 120. The source device 110 may be called a video encoding device and generates encoded video data. The destination device 120 may be called a video decoding device and may decode the encoded video data generated by the source device 110.
[0152] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0153] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to produce a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include the encoded picture and associated data. The encoded picture is the encoded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntactic structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or transmitter. The encoded video data may be transmitted directly to the destination device 120 via the I / O interface 116 through the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0154] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0155] The I / O interface 126 may include a receiver and / or modem. The I / O interface 126 may acquire encoded video data from the source device 110 or storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be outside the destination device 120 configured to interface with an external display device.
[0156] The video encoder 114 and video decoder 124 may operate in accordance with video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Multipurpose Video Coding (VVC) standard, and other current and / or further standards.
[0157] Figure 11 is a block diagram showing an example of a video encoder 200, which may be the video encoder 114 in the system 100 shown in Figure 10.
[0158] The video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example in Figure 11, the video encoder 200 includes several functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0159] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206, a residual generation unit 207, a conversion unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse conversion unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0160] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intrablock copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0161] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but for illustrative purposes, they are represented separately in the example in Figure 11.
[0162] The partitioning unit 201 can partition the picture into one or more video blocks. The video encoder 200 and video decoder 300 can support various video block sizes.
[0163] The mode selection unit 203 can, for example, select one of the encoding modes, intra or inter, based on the error result, and provide the resulting intra or inter-encoded blocks to the residual generation unit 207, which generates residual block data, and to the reconstruction unit 212, which reconstructs the encoded blocks for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter-prediction (CIIP) mode, where the prediction is based on an inter-prediction signal and an intra-prediction signal. The mode selection unit 203 may also select the resolution of the motion vectors for the blocks (e.g., sub-pixel or integer pixel precision) in the case of inter-prediction.
[0164] To perform interpretation for the current video block, the motion estimation unit 204 may generate motion information about the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0165] The motion estimation unit 204 and the motion compensation unit 205 may perform different actions for the current video block depending, for example, whether the current video block is in an I slice, a P slice, or a B slice.
[0166] In some examples, the motion estimation unit 204 may perform a one-way prediction for the current video block, and the motion estimation unit 204 may search for a reference video block for the current video block by looking up reference pictures in List 0 or List 1. The motion estimation unit 204 can then generate a reference index indicating the reference picture in List 0 or List 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0167] In other examples, the motion estimation unit 204 may perform bidirectional prediction for the current video block, and may search for a reference video block for the current video block in reference pictures in List 0, and may also search for another reference video block for the current video block in reference pictures in List 1. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in Lists 0 and 1 containing the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0168] In some examples, the motion estimation unit 204 may output a full set of motion information for the decoder's decoding process.
[0169] In some cases, the motion estimation unit 204 may not output a full set of motion information for the current video. Rather, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0170] In one example, the motion estimation unit 204 may indicate to the video decoder 300 that the current video block has the same motion information as another video block in the syntactic structure associated with the current video block.
[0171] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntactic structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0172] As described above, the video encoder 200 may predictively transmit motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merged mode signaling.
[0173] The intra-prediction unit 206 can perform intra-prediction on the current video block. When the intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntactic elements.
[0174] The residual generation unit 207 can generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block (for example, indicated by a minus sign). The residual data for the current video block may include residual video blocks corresponding to different sample components of the sample in the current video block.
[0175] In other examples, for instance in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform the subtraction operation.
[0176] The conversion processing unit 208 can generate one or more conversion factor video blocks for the current video block by applying one or more conversions to the residual video block associated with the current video block.
[0177] After the conversion processing unit 208 generates a conversion coefficient video block associated with the current video block, the quantization unit 209 can quantize the conversion coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0178] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the conversion coefficient video block in order to reconstruct the residual video block from the conversion coefficient video block. The reconstruction unit 212 adds the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0179] After the reconstruction unit 212 has reconstructed the video block, loop filtering operations may be performed to reduce video blocking artifacts within the video block.
[0180] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, it can perform one or more entropy coding operations to generate entropy coded data and output a bitstream containing the entropy coded data.
[0181] Figure 12 is a block diagram showing an example of a video decoder 300, which may be the video decoder 114 in the system 100 shown in Figure 10.
[0182] The video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example in Figure 12, the video decoder 300 includes several functional components. The techniques described in this disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0183] In the example shown in Figure 12, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding pass that is roughly the reverse of the encoding pass described for the video encoder 200 (Figure 11).
[0184] The entropy decoding unit 301 can extract the encoded bitstream. The encoded bitstream may contain entropy-encoded video data (for example, encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.
[0185] The motion compensation unit 302 can potentially perform interpolation based on an interpolation filter to generate motion-compensated blocks. An identifier for the interpolation filter used with sub-pixel precision may be included in the syntax element.
[0186] The motion compensation unit 302 can calculate interpolated values for integer-less pixels of a reference block using an interpolation filter used by the video encoder 20 during video block encoding. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to the received syntax information and generate a predicted block using the interpolation filter.
[0187] The motion compensation unit 302 can use some of the syntactic information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of the picture in the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and a list of reference frames) for each inter-encoded block, and other information for decoding the encoded video sequence.
[0188] The intra-prediction unit 303 may, for example, use the intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies the inverse transform.
[0189] The reconstruction unit 306 can add the residual blocks with the corresponding predicted blocks generated by the motion compensation unit 202 or the intra-prediction unit 303 to form the decoded blocks. If desired, a deblocking filter may also be applied to filter the decoded blocks to remove blocking artifacts. The decoded video blocks are then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra-prediction and also generates the decoded video for presentation on a display device.
[0190] In some embodiments, the following methods are based on the list of examples and embodiments enumerated above. In one example, these methods can be carried out using the implementations shown in Figures 8 to 12, but are not limited thereto.
[0191] Figure 13 is a flowchart of an exemplary method for video processing. As shown therein, method 1300 includes determining in operation 1310 that, for a conversion between the current block of video and a bitstream representation of the video, the current block is encoded using palette mode and escape symbol values are signaled in the bitstream representation, thereby using clipped quantization parameters for the current block in the conversion, the clipped quantization parameters used for the chroma components of the video are derived based on quantization parameters after a mapping operation of a quantization or dequantization process. Method 1300 further includes, in operation 1320, performing the conversion based on the decision.
[0192] Figure 14 is a flowchart of an exemplary method for video processing. As shown therein, method 1400 includes determining in operation 1410 that, for a conversion between the current block of video and a bitstream representation of the video, the current block is encoded using palette mode and escape symbol values are signaled in the bitstream representation, thereby using clipped quantization parameters for the current block in the conversion, the clipped quantization parameters used for the chroma components of the video are derived based on the quantization parameters prior to the mapping operation of the quantization or dequantization process. Method 1400 further includes, in operation 1420, performing the conversion based on the decision.
[0193] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (for example, Item 1) as preferred features of some embodiments.
[0194] 1. A video processing method comprising the steps of determining that, for a conversion between a current block of video and a bitstream representation of the video, the current block is encoded using palette mode and escape symbol values are signaled in the bitstream representation; and, based on the determination, performing the conversion, wherein, in the conversion, clipped quantization parameters for the current block are used; the palette mode encoding tool represents the current video block using a palette of representative color values; the escape values are used for a sample of the current video block that is encoded without using the representative color values; and the clipped quantization parameters used for the chroma components of the video are derived based on quantization parameters after a mapping operation of a quantization or dequantization process.
[0195] 2. A video processing method comprising the steps of determining that, for a conversion between a current block of video and a bitstream representation of the video, the current block is encoded using palette mode and escape symbol values are signaled in the bitstream representation; and, based on the determination, performing the conversion, wherein, in the conversion, clipped quantization parameters for the current block are used; the palette mode encoding tool represents the current video block using a palette of representative color values; the escape values are used for a sample of the current video block that is encoded without using the representative color values; and the clipped quantization parameters used for the chroma components of the video are derived based on quantization parameters prior to a mapping operation of a quantization or dequantization process.
[0196] 3. The method of Solution 1 or 2, wherein the minimum value of the clipped quantization parameter is based on the minimum allowable quantization parameter for the transformation skip mode.
[0197] 4. The method of Solution 3, where the minimum value of the clipped quantization parameter is based on sps_min_qp_prime_ts.
[0198] 5. Solution 4 method, where the minimum value of the clipped quantization parameter is equal to QpPrimeTsMin.
[0199] 6. The method of Solution 1 or 2, wherein the indication of the minimum value of the clipped quantization parameter for each color component of the video is signaled in the sequence parameter set (SPS), picture parameter set (PPS), video parameter set (VPS), decoding parameter set (DPS), tile, or slice header in the bitstream representation.
[0200] 7. The method of Solution 1 or 2, wherein the minimum value of the clipped quantization parameter is (bd-ibd)×6+4, where bd is the internal bit depth and ibd is the input bit depth for the color components of the video.
[0201] 8. A method according to any one of Solutions 1 to 7, wherein the palette mode encoding tool is applied to a color component of the video.
[0202] 9. A method according to any one of solutions 1 to 8, wherein performing the conversion includes generating the bitstream representation from the current block.
[0203] 10. A method according to any one of solutions 1 to 8, wherein performing the conversion includes generating the current block from the bitstream representation.
[0204] 11. A device in a video system comprising a processor and non-temporary memory having instructions, wherein, when executed by the processor, the instructions cause the processor to implement one of the methods of Solutions 1 to 10.
[0205] 12. A computer program product stored on a non-temporary computer-readable medium, which includes program code for performing the method of any one of solutions 1 through 10.
[0206] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, or one or more combinations thereof, including the structures disclosed herein and their structural equivalents. The disclosed and other embodiments may be implemented as one or more modules of computer program instructions encoded on a computer-readable medium for execution by one or more computer program products, i.e., for control of the operation of a data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage board, a memory device, a composition of a material that realizes a machine-readable propagating signal, or one or more combinations thereof. The term “data processing device” encompasses all devices and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, a device may include code that generates an execution environment for the computer program in question, such as processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagating signal is an artificially generated signal, such as a mechanically generated electrical, optical, or electromagnetic signal produced to encode information for transmission to a suitable receiving device.
[0207] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including languages that are compiled or interpreted, and can be deployed in any form, including as a standalone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file containing one or more modules, subprograms, or code sections). A computer program can be deployed to run on one computer, or on multiple computers located in one site, or distributed across multiple sites and interconnected by a communication network.
[0208] The processes and logic flows described in this paper can be executed by one or more programmable processors running one or more computer programs, performing their functions by acting on input data and producing outputs. The processes and logic flows can also be executed by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the devices can be implemented in such manner.
[0209] Processors suitable for executing computer programs include, for example, both general-purpose and dedicated microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random-access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operationally coupled to receive data from them, or transfer data to them, or both. However, a computer is not required to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory can be supplemented by or incorporated into special-purpose logic circuits.
[0210] While this patent document contains many specific details, these should not be interpreted as limitations on the scope of any subject matter or the scope of claims, but rather as descriptions of features that may be specific to particular embodiments of a particular technique. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any preferred subcombination. Furthermore, features are described above as acting in certain combinations, and may even be initially described in claims as such, but one or more features from a claimed combination may, in some cases, be removed from the combination, and the claimed combination may be directed towards subcombinations or variations of subcombinations.
[0211] Similarly, although the drawings show operations in a specific order, this should not be interpreted as requiring that such operations be performed in a specific order or sequentially, or that all shown operations be performed, in order to achieve the desired result. Furthermore, the isolation of various system components in the embodiments described in this patent document should not be understood as requiring such isolation in all embodiments.
[0212] Only a few implementations and examples are described, and other implementations, improvements, and modifications may be made based on what is described and explained in this patent document.
Claims
1. A method of video processing: A step in which it is determined that a prediction mode in the current video block is applied for the conversion between the current video block and the bitstream of the video, wherein the reconfigured samples of the current video block are represented by escaped samples; For the current video block, the steps include determining the quantization parameters used to derive the escaped samples; The steps include at least the step of performing the transformation based on the quantization parameters, The current video block is a chroma block, The quantization parameters are modified using Max(QpPrimeTsMin,Qp) before being used to derive the escaped samples. QpPrimeTsMin represents the minimum allowable quantization parameter for the transformation skip mode. Qp represents the quantization parameter, The value of QpPrimeTsMin is equal to (4 + 6 × n), n represents the value of a syntactic element included in the sequence parameter set in the bitstream. method.
2. The method according to claim 1, wherein the conversion includes encoding the current video block into the bitstream.
3. The method according to claim 1, wherein the conversion includes decoding the current video block from the bitstream.
4. A video data processing device having a processor and non-temporary memory having instructions, wherein when an instruction is executed by the processor, the processor: A step in which it is determined that a prediction mode in the current video block is applied for the conversion between the current video block and the bitstream of the video, wherein the reconfigured samples of the current video block are represented by escaped samples; For the current video block, the steps include determining the quantization parameters used to derive the escaped samples; This involves performing at least the step of performing the transformation based on the quantization parameters, The current video block is a chroma block, The quantization parameters are modified using Max(QpPrimeTsMin,Qp) before being used to derive the escaped samples. QpPrimeTsMin represents the minimum allowable quantization parameter for the transformation skip mode. Qp represents the quantization parameter, The value of QpPrimeTsMin is equal to (4 + 6 × n), n represents the value of a syntactic element included in the sequence parameter set in the bitstream. Device.
5. A non-temporary, computer-readable storage medium storing instructions, wherein the instructions are transmitted to a processor: A step in which it is determined that a prediction mode in the current video block is applied for the conversion between the current video block and the bitstream of the video, wherein the reconfigured samples of the current video block are represented by escaped samples; For the current video block, the steps include determining the quantization parameters used to derive the escaped samples; This involves performing at least the step of performing the transformation based on the quantization parameters, The current video block is a chroma block, The quantization parameters are modified using Max(QpPrimeTsMin,Qp) before being used to derive the escaped samples. QpPrimeTsMin represents the minimum allowable quantization parameter for the transformation skip mode. Qp represents the quantization parameter, The value of QpPrimeTsMin is equal to (4 + 6 × n), n represents the value of a syntactic element included in the sequence parameter set in the bitstream. storage medium.
6. A method for storing a video bitstream, the method being: A step in which it is determined that a prediction mode in the current video block is applied for the conversion between the current video block and the bitstream of the video, wherein the reconfigured samples of the current video block are represented by escaped samples; For the current video block, the steps include determining the quantization parameters used to derive the escaped samples; The steps include generating the bitstream based on at least the quantization parameters; The step includes storing the bitstream on a non-temporary computer-readable recording medium, The current video block is a chroma block, The quantization parameters are modified using Max(QpPrimeTsMin,Qp) before being used to derive the escaped samples. QpPrimeTsMin represents the minimum allowable quantization parameter for the transformation skip mode. Qp represents the quantization parameter, The value of QpPrimeTsMin is equal to (4 + 6 × n), n represents the value of a syntactic element included in the sequence parameter set in the bitstream. method.
Citation Information
Patent Citations
Data encoding and decoding
JP2016519903A
Cited By
Quantization parameter derivation for palette mode
JP2024096202A