Block size restriction for cross-component video coding
By optimizing video coding through cross-component linear model prediction and transform skip mode, the problem of low efficiency in chroma block coding is solved, resulting in more efficient video coding and decoding effects and improved video quality under different sampling formats.
Patent Information
- Application Number
- CN202080076944.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-01
- Filing Date
- 2020-11-02
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2040-11-02
AI Technical Summary
Existing video encoding and decoding technologies suffer from low encoding efficiency and poor decoding quality when processing chroma blocks, especially under different resolution chroma and luminance sampling formats, making it difficult to effectively utilize the characteristics of the human visual system for efficient encoding.
The cross-component linear model prediction technique is adopted to generate predicted values of chroma blocks by using downsampled luminance samples. Combined with transform skip mode and multiple transform sets, the encoding and decoding process of video blocks is optimized, allowing different maximum allowable block sizes for different parts to adapt to the characteristics of different video regions.
It improves the efficiency and quality of video encoding and decoding, especially the encoding efficiency under different chroma and luminance sampling formats, reduces bandwidth requirements and improves video quality.
Smart Images

Figure CN114667730B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority and benefit from International Patent Application No. PCT / CN2019 / 115034, filed November 1, 2019, in accordance with the patent law and / or rules applicable under the Paris Convention. For all purposes required by law, the entire disclosure of the aforementioned application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This document covers video and image encoding and decoding technologies. Background Technology
[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] The disclosed techniques can be used by video or image decoder or encoder embodiments to perform encoding or decoding using cross-component linear model prediction.
[0006] In one example aspect, a method for processing video is disclosed. The method includes: a conversion between a chroma block of the video and a bitstream representation of the video, deriving parameters of a cross-component linear model using downsampled luminance samples generated from N upper adjacent lines of a collocated luminance block of the chroma block using a downsampling filter, where N is a positive integer; and performing the conversion with a predicted chroma block generated using the cross-component linear model.
[0007] In another example, a method for processing video is disclosed. The method includes: a conversion between video regions of video components and a bitstream representation of the video; determining the maximum allowed block size of video blocks encoded using a transform skip mode; and performing the conversion based on that determination.
[0008] In another example, a method for processing video is disclosed. The method includes performing a conversion between a video comprising video blocks and a bitstream representation of the video according to a first rule and a second rule, wherein a transform skip codec tool is used to encode and decode a first portion of the video blocks, wherein a transform codec tool is used to encode and decode a second portion of the video blocks, wherein the first rule specifies the maximum allowed block size for the first portion of the video blocks, and the second rule specifies the maximum allowed block size for the second portion of the video blocks, and wherein the maximum allowed block size for the first portion of the video blocks is different from the maximum allowed block size for the second portion of the video blocks.
[0009] In another example, a method for processing video is disclosed. This method includes performing a conversion between a video comprising one or more chroma blocks and a bitstream representation of the video, wherein the bitstream representation conforms to a format rule specifying whether syntax elements indicating the use of a transform skip tool are included in the bitstream representation depending on the maximum allowed size of the chroma blocks encoded and decoded using the transform skip tool.
[0010] In another example, a method for processing video is disclosed. The method includes performing a conversion between a video comprising one or more first video blocks of a first chroma component and one or more second video blocks of a second chroma component, and a bitstream representation of the video, wherein the bitstream representation conforms to a format rule specifying the use of syntax elements that jointly indicate the availability of transform skipping tools for encoding and decoding one or more first chroma blocks and one or more second chroma blocks.
[0011] In another example, the above method can be implemented by a video encoder device that includes a processor.
[0012] In yet another example, these methods can be embodied in processor-executable instructions and stored on a computer-readable program medium.
[0013] These and other aspects are further described in this document. Attached Figure Description
[0014] Figure 1A The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown.
[0015] Figure 1B An example of a video encoder is shown.
[0016] Figure 2 Examples of 67 intra-frame prediction modes are shown.
[0017] Figure 3 Examples of horizontal and vertical traversal scans are shown.
[0018] Figure 4 An example of the sample point locations used for the derivation of α and β is shown.
[0019] Figure 5 An example of dividing a 4×8 sample block into two independent decodeable regions is shown.
[0020] Figure 6 An example sequence of processing for rows of pixels that maximizes throughput for a 4×N block using vertical predictors is shown.
[0021] Figure 7 An example of the Low Frequency Inseparable Transform (LFNST) process is shown.
[0022] Figure 8 An example of adjacent chroma samples and downsampled co-bit adjacent luminance samples used in the derivation of CCLM parameters for 4:2:2 video is shown.
[0023] Figure 9 An example of a video processing device is shown.
[0024] Figure 10 A block diagram of a video encoder is shown.
[0025] Figure 11 This is a flowchart illustrating an example of a video processing method based on some implementations of the disclosed technology.
[0026] Figure 12 This is a block diagram of an example video processing system.
[0027] Figure 13 This is a block diagram illustrating an example video codec system.
[0028] Figure 14 This is a block diagram illustrating an encoder according to some embodiments of the disclosed technology.
[0029] Figure 15 This is a block diagram illustrating a decoder according to some embodiments of the disclosed technology.
[0030] Figure 16A and Figure 16B This is a flowchart illustrating an example of a video processing method based on some implementations of the disclosed technology. Detailed Implementation
[0031] This document provides various techniques that decoders of image or video bitstreams can use to improve the quality of decompressed or decoded digital video or images. For simplicity, the term "video" used herein includes both sequences of images (traditionally referred to as video) and individual images. Furthermore, video encoders may also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0032] Chapter headings are used in this document for ease of understanding and do not limit the embodiments and techniques to the corresponding chapters. Thus, embodiments from one part can be combined with embodiments from other parts.
[0033] 1. Brief Overview
[0034] This invention relates to video codec technology. Specifically, it relates to cross-component linear model prediction and other codec tools in image / video codecs. It can be applied to existing video codec standards like HEVC, or to standards that are yet to be finalized (Multi-Functional Video Codec). It can also be applied to future video codec standards or video codecs.
[0035] 2. Introduction to Video Encoding and Decoding
[0036] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) established the Joint Video Experts Group (JVET) to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.
[0037] 2.1. Color Space and Chroma Subsampling
[0038] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes the range of colors as tuples of numbers, typically 3 or 4 values or color components (e.g., RGB). Essentially, a color space is a detailed description of a coordinate system and its subspaces.
[0039] For video compression, the most commonly used color spaces are YCbCr and RGB.
[0040] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a color space family used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, while CB and CR are the blue and red chromaticity components. Y′ (with an apostrophe) is distinct from Y, where Y is luminance, meaning that the light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.
[0041] Chromatic subsampling is a practice of encoding images by taking advantage of the fact that the human visual system is less sensitive to color differences than to brightness, thereby achieving a lower resolution for chromatic information than for luminance information. 2.1.1. 4:4:4
[0043] Each of the three Y'CbCr components has the same sampling rate, therefore there is no chromaticity subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2
[0045] The two chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third, but with almost no visual difference. (In the VVC working draft...) Figure 1A Examples of the nominal vertical and horizontal positions for the 4:2:2 color format are described. 2.1.3. 4:2:0
[0047] Compared to 4:1:1, the horizontal sampling in 4:2:0 is doubled, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating line in this scheme. The data rate is therefore the same. The Cb and Cr channels are subsampled in both the horizontal and vertical directions by a factor of 2. There are three variants of the 4:2:0 scheme with different horizontal and vertical addressing.
[0048] In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are located between pixels in the vertical direction (in the gap).
[0049] • In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in the gap, in the middle between alternating luminance samples.
[0050] In a 4:2:0 DV configuration, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternating lines.
[0051] Table 2-1. SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag
[0052]
[0053] 2.2. Encoder / decoder streams of typical video codecs
[0054] Figure 1BAn example of a VVC encoder block diagram is shown, comprising three in-loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the raw samples of the current image. With the signaling from the encoding / decoding process informing the offset and filter coefficients, they reduce the mean square error between the raw and reconstructed samples by adding an offset and applying a Finite Impulse Response (FIR) filter, respectively. ALF is located in the final processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.
[0055] 2.3. Intra-mode encoding and decoding with 67 intra-prediction modes
[0056] To capture arbitrary edge directions presented in natural video, the number of intra-frame directional modes has increased from 33 used by HEVC to 65. Additional directional modes are added in... Figure 2 The plane and DC modes remain unchanged and are depicted as red dashed arrows. These denser directional intra-prediction modes are applicable to all block sizes and to both luma and chroma intra-prediction.
[0057] Traditional intra-frame prediction direction is defined as clockwise from 45 degrees to -135 degrees, such as... Figure 2 As shown. In VTM, for non-square blocks, several traditional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are signaled using the original method and remapped to the wide-angle mode index after parsing. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding and decoding remain unchanged.
[0058] In HEVC, each intra-codec block has a square shape, and the length of each side is a power of 2. Therefore, generating intra-predictors using DC mode does not require division. In VVC, blocks can have rectangular shapes, making division for each block necessary in general. To avoid division in DC prediction, only the longer sides are used to calculate the average of non-square blocks.
[0059] Figure 2 Examples of 67 intra-frame prediction modes are shown.
[0060] 2.4. Inter-frame prediction
[0061] For each inter-frame predicted CU, motion parameters include motion vectors, reference picture indexes, reference picture list usage indexes, and additional information required for the new encoding / decoding features of the VVC used to generate inter-frame predicted samples. Motion parameters can be signaled explicitly or implicitly. When encoding / decoding a CU in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded / decoded motion vector increments, or reference picture indexes. A merge mode is specified, thereby obtaining the motion parameters of the current CU from neighboring CUs (including spatial and temporal candidates) and the additional scheduling introduced in the VVC. The merge mode can be applied to any inter-frame predicted CU, not just skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where motion vectors, corresponding reference picture indexes for each reference picture list, reference picture list usage flags, and other necessary information are explicitly signaled for each CU.
[0062] 2.5. Intra-Block Copying (IBC)
[0063] Intra-Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well-known for significantly improving the encoding and decoding efficiency of screen content footage. Since IBC mode is implemented as a block-level encoding / decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector indicates the displacement from the current block to a reference block that has already been reconstructed within the current frame. The luma block vector of the CU encoded with IBC is integer precision. The chroma block vector is also rounded to integer precision. When used in conjunction with AMVR, IBC mode can switch between 1-pixel and 4-pixel motion vector precision. CUs encoded with IBC are considered a third prediction mode, distinct from intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.
[0064] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height of no more than 16 luminance samples. For non-merge modes, a block vector search is first performed using a hash-based search. If the hash search does not return any valid candidates, a local search based on block matching is performed.
[0065] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4×4 sub-blocks. For larger current blocks, a hash key is determined to be the matching reference block's hash key when all hash keys in all 4×4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the current block's hash key, the block vector cost of each matching reference is calculated, and the one with the lowest cost is selected.
[0066] In block matching search, the search scope is set to cover both the previous and current CTUs.
[0067] At the CU level, IBC mode is signaled using a flag and can be signaled as either IBC AMVP mode or IBC skip / merge mode, as shown below:
[0068] –IBC Skip / Merge Mode: The merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codecs is used to predict the current block. The merge list includes spatial candidates, HMVP candidates, and paired candidates.
[0069] –IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (if IBC encoding and decoding). When either neighbor is unavailable, the default block vector will be used as the predictor. Signaling notification flags indicate the block vector predictor index.
[0070] 2.6. Palette Mode
[0071] For palette mode signaling, the palette mode is encoded and decoded into the prediction mode of the codec unit; that is, the prediction mode of the codec unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. If palette mode is used, the pixel values in the CU are represented by a small set of representative color values. This set is called the palette. For pixels with values close to the palette colors, the palette index is signaled. For pixels with values outside the palette, the pixel is represented by an escape symbol, and the quantized pixel value is signaled directly.
[0072] To decode a palette-encoded block, the decoder needs to decode the palette colors and indices. Palette colors are described by a palette table and encoded by a palette table encoder / decoder. Escape symbols are signaled to each CU to indicate whether an escape symbol exists in the current CU. If an escape symbol exists, the palette table is incremented by 1, and the last index is assigned to the escape mode. The palette indices of all pixels in the CU form a palette index map and are encoded by a palette index map encoder / decoder.
[0073] For the encoding and decoding of the palette table, the palette predictors are maintained. Predictors are initialized at the beginning of each stripe, where they are reset to 0. For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. The reuse flag is sent using zero-run-length encoding and decoding. Subsequently, the number of new palette entries is signaled using zero-order exponent Golomb codes. Finally, the component values of the new palette entries are signaled. After encoding the current CU, the palette predictors are updated using the current palette, and entries from previous palette predictors that are not reused in the current palette are added to the end of the new palette predictor until the maximum allowed size (palette fill) is reached.
[0074] To encode and decode the palette index map, use methods such as... Figure 3 The horizontal and vertical traversal scans shown encode and decode the index. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag.
[0075] Figure 3 Examples of horizontal and vertical traversal scans are shown.
[0076] The palette index is encoded and decoded using two main palette sample modes: 'INDEX' and 'COPY_ABOVE'. A flag is used to signal the mode except when the top row is used in horizontal scan or the first column or preceding column is in 'COPY_ABOVE' mode in vertical scan. In 'COPY_ABOVE' mode, the palette index of the samples in the upper row is copied. In 'INDEX' mode, the palette index is explicitly signaled. For both 'INDEX' and 'COPY_ABOVE' modes, a running value specifying the number of pixels encoded and decoded using the same mode is signaled.
[0077] The encoding order of the index mapping is as follows: First, the signaling informs the number of index values for the CU. Next, truncated binary encoding / decoding is used to signal the actual index values for the entire CU. Both the number of indices and the index values are encoded / decoded in bypass mode. This groups the bypass binary bits associated with the index together. Then, the palette mode (INDEX or COPY_ABOVE) and run length are signaled in an interleaved manner. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and encoded / decoded in bypass mode. After signaling the index values, the signaling informs the appended syntax element `last_run_type_flag`. This syntax element, combined with the number of indices, eliminates the need to signal the run length value corresponding to the last run in the block.
[0078] In VTM, a dual-tree is enabled for I-stripes, which separates the codec units for luma and chroma. Therefore, in this proposal, the color palette is applied separately to the luma (Y component) and chroma (Cb and Cr components). If the dual-tree is disabled, the color palette will be applied jointly to the Y, Cb, and Cr components, the same as in HEVC.
[0079] 2.7. Cross-component linear model prediction
[0080] In VVC, the Cross-Component Linear Model (CCLM) is used to predict the mode. For this mode, chromaticity samples are predicted based on reconstructed luminance samples from the same CU using the following linear model:
[0081] pred C (i,j)=α·rec L ′(i,j)+β (2-1)
[0082] Among them, pred C (i,j) represents the predicted chromaticity sample points in the CU, and rec L (i,j) represents the downsampled reconstructed luminance sample points of the same CU.
[0083] Figure 4 This shows examples of the positions of the left and top samples involved in the LM mode, as well as the samples of the current block.
[0084] Figure 4 An example of the location of the sample points used to derive α and β is shown.
[0085] Besides the top and left templates, which can be used together to calculate linear model coefficients in LM mode, they can also be used alternately in two other LM modes, called LM_A and LM_L modes. In LM_A mode, only the top template is used to calculate linear model coefficients. To obtain more samples, the top template is extended to (W+H). In LM_L mode, only the left template is used to calculate linear model coefficients. To obtain more samples, the left template is extended to (H+W). For non-square blocks, the top template is extended to W+W, and the left template is extended to H+H.
[0086] The CCLM parameters (α and β) are derived from at most four adjacent chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block dimension is W×H, then W' and H' are set to...
[0087] – When applying the LM pattern, W' = W, H' = H
[0088] – When applying the LM-A mode, W' = W + H;
[0089] – When applying the LM-L pattern, H' = H + W;
[0090] The adjacent positions above are denoted as S[0, -1]…S[W'-1, -1], and the adjacent positions to the left are denoted as S[-1, 0]…S[-1, H'-1]. These four sample points are then selected as…
[0091] –When LM mode is applied and the upper and left adjacent samples are available, S[W' / 4, -1], S[3W' / 4, -1], S[-1, H' / 4], S[-1, 3H' / 4];
[0092] – When applying LM-A mode or when only the upper adjacent sample points are available, S[W' / 8, -1], S[3W' / 8, -1], S[5W' / 8, -1], S[7W' / 8, -1];
[0093] – When applying LM-L mode or when only the left adjacent sample point is available, S[-1, H' / 8], S[-1, 3H' / 8], S[-1, 5H' / 8], S[-1, 7H' / 8];
[0094] Four adjacent brightness samples at the selected location are downsampled and compared four times to find the two smaller values: x 0 A and x 1 A and two larger values: x 0 B and x 1 BTheir corresponding chromaticity sample values are represented as y 0 A y 1 A, y 0 B and y 1 B Then x A x B y A and y B It can be deduced as:
[0095] X a =(x 0 A +x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;
[0096] Y b =(y 0 B +y 1 B +1)>>1 (2-1)
[0097] Finally, the linear model parameters α and β are obtained according to the following equation.
[0098]
[0099] β=Y b -α·X b (2-4)
[0100] The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented using exponent notation. For example, diff is approximated using a 4-bit significant part and an exponent. Therefore, for the 16 significant values, the table for 1 / diff is simplified to 16 elements, as follows:
[0101] DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}
[0102] (2-3)
[0103] This will help reduce the complexity of the calculations and the memory size required to store the necessary tables.
[0104] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filters is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type-0" and "Type-2" content, respectively.
[0105]
[0106]
[0107] Please note that when the upper reference line is at the CTU boundary, only one luminance line (the universal line buffer in intra-frame prediction) is used to form the downsampled luminance sample.
[0108] This parameter calculation is performed as part of the decoding process, not just as part of the encoder's search operation. Therefore, the α and β values are not passed to the decoder using syntax.
[0109] For chroma intra-mode encoding and decoding, a total of eight intra-modes are allowed. These modes include five traditional intra-modes and three cross-component linear model modes (LM, LM_A, and LM_L). Table 2-2 shows the chroma mode signaling and derivation process. Chroma mode encoding and decoding directly depend on the intra-prediction mode of the corresponding luma block. Since separate block partitioning structures for luma and chroma components are enabled in I-strips, one chroma block can correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra-prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.
[0110] Table 2-1. Derivation of chromaticity prediction mode from luminance mode when CCLM is enabled.
[0111]
[0112] 2.8. Block Differential Pulse Encoding / Decoding Modulation / Decoding
[0113] BDPCM was proposed in JVET-M0057. Since the shape of the horizontal (correspondingly, vertical) predictor is predicted using the left (A) (correspondingly, the top (B)) pixels of the current pixel, the most throughput-efficient way to process this block is to process all pixels in a column (correspondingly, a line) in parallel and then process these columns (correspondingly, the lines) sequentially. To increase throughput, we introduce the following process: when the predictor selected on the block is vertical, a block with a width of 4 is divided into two halves with horizontal boundaries, and when the predictor selected on the block is horizontal, a block with a height of 4 is divided into two halves with vertical boundaries.
[0114] When blocks are divided, samples from one region are not allowed to use pixels from another region to compute predictions: if this happens, the predicted pixel is replaced by a reference pixel in the prediction direction. This is in... Figure 5 The different positions of the current pixel X in the 4×8 block for vertical prediction are shown in the figure.
[0115] Figure 5 An example of dividing a 4×8 sample block into two independent decodeable regions is shown.
[0116] Because of this characteristic, it is now possible to process 4×4 blocks in 2 cycles, and 4×8 or 8×4 blocks in 4 cycles, and so on. Figure 6 As shown.
[0117] Figure 6 An example sequence of processing for rows of pixels that maximizes throughput for a 4×N block using vertical predictors is shown.
[0118] Table 2-2 outlines the number of cycles required to process a block based on its block size. It is easy to see that blocks with any two dimensions greater than or equal to 8 can be processed at 8 pixels or more per cycle.
[0119] Table 2-2. Worst-case throughput of blocks with sizes of 4xN and Nx4
[0120]
[0121] 2.9. Quantization Residual Domain BDPCM
[0122] In JVET-N0413, Quantization Residual Domain BDPCM (hereinafter referred to as RBDPCM) was proposed. Intra-prediction of the entire block is performed by copying samples in the prediction direction (horizontal or vertical prediction) similar to intra-prediction. The residual is quantized, and the increment between the quantized residual and the quantized value of its predictor (horizontal or vertical) is encoded and decoded.
[0123] For a block of size M (rows) × N (columns), let ri,j ,0≤i≤M-1,0≤j≤N-1 is the prediction residual after performing horizontal (copying the left adjacent pixel value line by line across the prediction block) or vertical (copying the top adjacent line to each line in the prediction block) intra-frame prediction using unfiltered samples from the upper or left block boundary samples. Let Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1 represent residuals r i,j The quantized version, where the residual r i,j This is the difference between the original block value and the predicted block value. Then, the block DPCM is applied to the quantized residual samples to produce elements... Modified M×N array When signaling is sent to the vertical BDPCM:
[0124]
[0125] For horizontal prediction, similar rules apply, and the residual quantization samples are obtained by the following formula.
[0126] Residual Quantization Samples It is sent to the decoder.
[0127] On the decoder side, the above calculation is reversed to produce Q(r). i,j ), 0≤i≤M-1, 0≤j≤N-1. For the vertical prediction case,
[0128]
[0129] Regarding the horizontal situation
[0130]
[0131] Inverse quantization residual Q -1 (Q(r i,j The values are added to the intra-block prediction values to produce reconstructed sample values.
[0132] The main advantage of this approach is that the inverse DPCM transform can be dynamically completed by simply adding predictors as the coefficients are resolved during the coefficient resolution process, or this can be performed after resolution.
[0133] Transform skipping is always used in the quantized residual domain BDPCM.
[0134] 2.10. VVC Multiple Transform Set (MTS)
[0135] VTM supports large block size transforms up to 64×64, which is primarily useful for higher resolution video such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both) equal to 64, high-frequency transform coefficients are zeroed out, thus retaining only low-frequency coefficients. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the top 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without zeroing out any values. VTM also supports the configurable maximum transform size in SPS, allowing the encoder to flexibly choose transform sizes up to 16, 32, or 64 in length, depending on the specific implementation requirements.
[0136] In addition to DCT-II already used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding and decoding of inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. The table shows the basis functions of the selected DST / DCT.
[0137] Table 2-4. Transform basis functions for DCT-II / VIII and DSTVII for N-point inputs
[0138]
[0139] To maintain the orthogonality of the transformation matrix, the transformation matrix is quantized more precisely than the HEVC transformation matrix. To keep the intermediate values of the transformation coefficients within the 16-bit range, all coefficients will have 10 bits after the horizontal and vertical transformations.
[0140] To control the MTS scheme, separate enable flags are specified at the SPS level for intra-frame and inter-frame operations. When MTS is enabled at the SPS, signaling notifies the CU level flag to indicate whether MTS should be applied. Here, MTS applies only to luminance. Signaling notifies the MTS CU level flag when the following conditions are met.
[0141] - Both width and height are less than or equal to 32
[0142] -CBF logo equals one
[0143] If the MTS CU flag is zero, DCT2 is applied in both directions. However, if the MTS CU flag is 1, additional signaling informs two other flags to indicate the transform type used for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 2-3. Transform selection using a unified ISP and implicit MTS is achieved by removing intra-frame mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra-frame block and both intra-frame and inter-frame explicit MTS are enabled, only DST7 is used for both the horizontal and vertical transform cores. For transform matrix precision, an 8-bit primary transform core is used. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. Furthermore, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8, use an 8-bit primary transform core.
[0144] Table 2-3. Transformation and Signalling Mapping Table
[0145]
[0146] To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are zeroed out for DST-7 and DCT-8 blocks with dimensions (width or height, or both) equal to 32. Only the coefficients in the 16×16 low-frequency region are retained.
[0147] In HEVC, block residuals can be encoded and decoded using transform skip mode. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to zero. The block size limit for transform skip is the same as the block size limit for MTS in JEM4, indicating that transform skip applies to the CU when both the block width and height are equal to or less than 32. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Furthermore, implicit MTS can still be enabled when MTS is enabled for inter-frame encoding / decoding blocks.
[0148] 2.11. Low-Frequency Non-Separable Transform (LFNST)
[0149] In VVC, LFNST (Low Frequency Inseparable Transform), known as the simplified secondary transform, is applied between the forward primary transform and quantization (on the codec side) and between dequantization and the inverse primary transform (on the decoder side), such as... Figure 7As shown. In the LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. For example, a 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and an 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4).
[0150] Figure 7 An example of the low-frequency non-separable transform (LFNST) process is shown.
[0151] The application of the non-separable transform being used in the LFNST is described below by taking an input as an example. To apply a 4×4 LFNST, the 〖4×4〗 input block X
[0152]
[0153] is first represented as a vector
[0154]
[0155] The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16×16 transform matrix. The 16×1 coefficient vector is then reorganized into a 4×4 block using the scan order (horizontal, vertical, or diagonal) of the block. Coefficients with smaller indices will be placed in the 4×4 coefficient block together with smaller scan indices.
[0156] 2.11.1. Simplified Non-Separable Transform
[0157] The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method so that it is implemented in a single pass without multiple iterations. However, it is necessary to reduce the dimension of the non-separable transform matrix to minimize the computational complexity and the memory space for storing the transform coefficients. Therefore, the simplified non-separable transform (or RST) method is used in the LFNST. The main idea of the simplified non-separable transform is to map an N-dimensional vector (for an 8×8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix, as follows:
[0158]
[0159] The R rows of the transform form an R basis in N-dimensional space. The inverse transform matrix of RT is the transpose of its forward transform. For the 8×8 LFNST, a simplification factor of 4 is applied, and the 64×64 direct matrix, which is the size of a regular 8×8 inseparable transform matrix, is simplified to a 16×48 direct matrix. Therefore, a 48×16 inverse RST matrix is used on the decoder side to generate the core (primary) transform coefficients in the upper left region of the 8×8. When the 16×48 matrix is applied instead of the 16×64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4×4 blocks in the upper left 8×8 block, excluding the lower right 4×4 block. With the help of the reduced dimensionality, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, a reasonable performance decrease. To reduce complexity, the LFNST is restricted to only being applied when all coefficients outside the first coefficient subgroup are invalid. Therefore, when the LFNST is applied, all primary-only transform coefficients must be zero. This allows adjustment of the LFNST index signaling at the last valid position, thus avoiding the extra coefficient scan in the current LFNST design, which is only needed to check valid coefficients at specific positions. The worst-case handling of LFNST (in terms of per-pixel multiplication) restricts the non-separable transformation of 4×4 and 8×8 blocks to 8×16 and 8×48 transformations, respectively. In these cases, for other sizes less than 16, the last valid scan position must be less than 8 when LFNST is applied. For blocks with shapes of 4×N and N×4 where N>8, the proposed restriction means that LFNST is now applied only once, and only to the top-left 4×4 region. Since all primary-only coefficients are zero when LFNST is applied, the number of operations required for the primary transformation is reduced in this case. From the codec's perspective, coefficient quantization is significantly simplified when testing the LFNST transformation. For the first 16 coefficients (in scan order), rate-distortion optimized quantization must be performed to the maximum extent, and the remaining coefficients are forced to zero.
[0160] 2.11.2. LFNST Transform Selection
[0161] There are a total of 4 transform sets in LFNST, and each transform set uses 2 inseparable transform matrices (kernels). The mapping from intra-prediction modes to transform sets is predefined, as shown in Table 2-4. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by an explicitly signaled LFNST index. After the transform coefficients, each intra-CU signals the index once in the bitstream.
[0162] Table 2-4. Transformation Selection Table
[0163] Intra-prediction mode Transform set index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0
[0164] 2.11.3. LFNST Index Signaling and Interaction with Other Tools
[0165] Because LFNST is restricted to application only when all coefficients outside the first coefficient subgroup are invalid, LFNST index encoding / decoding depends on the position of the last valid coefficient. Furthermore, LFNST indexing is context-coded but not dependent on the intra-prediction mode, and only the first binary bit (bin) is context-coded. Additionally, LFNST applies to intra-CUs in both intra-frame and inter-frame stripes, and to both luma and chroma. If dual-tree is enabled, the LFNST indexes for luma and chroma are signaled separately. For inter-frame stripes (where dual-tree is disabled), a single LFNST index is signaled and used for both luma and chroma.
[0166] When ISP mode is selected, LFNST is disabled, and the RST index is not signaled because the performance improvement is negligible even if RST is applied to every feasible partition block. Furthermore, disabling RST for the residuals of ISP predictions reduces coding complexity. When MIP mode is selected, LFNST is also disabled, and the index is not signaled.
[0167] Considering that due to the existing maximum transform size limit (64×64), large CUs larger than 64×64 are implicitly partitioned (TU tiling), LFNST index search could increase the data buffer size by a factor of four for a given number of decoding pipeline stages. Therefore, the maximum size allowed by LFNST is limited to 64×64. Note that LFNST is only enabled in the DCT case.
[0168] 2.12. Skip chroma transformations
[0169] Chroma Transform Skip (TS) was introduced into VVC. The motivation was to unify the TS and MTS signaling between luma and chroma by relocating `transform_skip_flag` and `mts_idx` to the residual coding section. A context model was added for chroma TS. For `mts_idx`, no context model and binary representation were changed. Furthermore, TS residual coding / decoding was also applied when using chroma TS.
[0170] Semantics
[0171] `transform_skip_flag[x0][y0][cIdx]` specifies whether to apply a transformation to the associated luminance transform block. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the considered transform block relative to the top-left luminance sample of the slice. `transform_skip_flag[x0][y0][cIdx]` equal to 1 indicates that no transformation is applied to the current luminance transform block. The transform block applies any transform. The array index cIdx specifies an indicator for the color components; it is 0 for luminance, 1 for Cb, and 2 for Cr. The fact that transform_skip_flag[x0][y0][cIdx] is equal to 0 indicates that the decision to apply a transform to the current transform block depends on other syntax elements. When transform_skip_flag[x0][y0][cIdx] does not exist, it is inferred to be equal to 0.
[0172] 2.13. BDPCM for Chromaticity
[0173] In addition to chroma TS support, BDPCM is added for the chroma components. If sps_bdpcm_enable_flag is 1, another syntax element, sps_bdpcm_chroma_enable_flag, is added to SPS. These flags have the following behavior as indicated in the table.
[0174] Table 2-7. SPS Marker for Luminosity and Chromaticity BDPCM
[0175]
[0176] When BDPCM is available only for luma, the current behavior remains unchanged. When BDPCM is also available for chroma, a bdpcm_chroma_flag is sent for each chroma block. This indicates whether BDPCM is used on the chroma blocks. When enabled, BDPCM is used for both chroma components, and an additional bdpcm_dir_chroma flag is encoded and decoded to indicate the prediction direction for both chroma components.
[0177] The deblocking filter is deactivated at the boundary between two Block-DPCM blocks because neither block uses the transform stage typically responsible for blocking artifacts. This deactivation occurs independently for the luma and chroma components.
[0178] 3. Examples of technical problems solved by publicly available solutions
[0179] The current design for linear parameter derivation in CCLM and TS has the following problems:
[0180] 1. For non-4:4:4 color formats, the derivation of the linear parameter in CCLM involves adjacent chroma samples and downsampled co-located adjacent luminance samples. For example... Figure 8 As shown, in the current VVC, when the nearest line is not at the CTU boundary, for 4:2:2 video, the second line above the current block is used to derive the downsampled co-located adjacent top luma sample. However, for 4:2:2 video, the vertical resolution remains unchanged. Therefore, there is a phase shift between the downsampled co-located adjacent top luma sample and the adjacent chroma sample.
[0181] Figure 8 An example is shown showing the adjacent chroma samples and downsampled co-located adjacent luminance samples used in the derivation of CCLM parameters in a 4:2:2 video.
[0182] 2. In the current VVC, the same maximum block size is used for the conditional checks of both the luminance transition skip flag and chrominance transition skip flag signaling. This design does not take color format into account, which is undesirable.
[0183] a. Similar issues exist for the signaling of the luminance BDPCM flag and the signaling of the chrominance BDPCM flag, where the same maximum block size is used in the condition check.
[0184] 4. List of embodiments and technologies
[0185] The list below should be considered as examples for explaining general concepts. These items should not be interpreted in a narrow way. Furthermore, these items can be combined in any way.
[0186] In this document, the term "CCLM" refers to an encoding / decoding tool that uses cross-color component information to predict the samples / residuals of the current color component, or to derive the reconstruction of the samples in the current color component. It is not limited to the CCLM technique described in VVC.
[0187] Derivation of linear parameters in CCLM
[0188] 1. When deriving the CCLM parameters of a chroma block, one or more of its co-occurring luminance blocks above adjacent lines can be used to derive its downsampled co-occurring adjacent top luminance samples.
[0189] a. In one example, when the current chroma block is not at the top CTU boundary, the nearest line above the co-occurring luminance block can be used instead of the second line above to deduce the downsampled co-occurring adjacent top luminance sample.
[0190] i. In one example, the same downsampling filter can be used to derive the downsampled co-occurring top luminance sample and the downsampled co-occurring left luminance sample.
[0191] 1) For example, a [1 2 1] filter can be used. More specifically, pDsY[x] = (pY[2*x–1][-1] + 2*pY[2*x][-1] + pY[2*x+1][-1] + 2) >> 2, where pY[2*x][-1], pY[2*x–1][-1], and pY[2*x+1][-1] are luminance samples from the nearest upper adjacent line, and pDstY[x] represents the downsampled co-located adjacent top luminance sample.
[0192] ii. In one example, different downsampling filters (e.g., different filter taps / different filter coefficients) can be used to derive the co-occurring top luminance sample and the co-occurring left luminance sample of the downsampling.
[0193] iii. In one example, the same downsampling filter can be used to derive the co-occurring top luminance samples of the downsampled chroma blocks, regardless of the location of the chroma blocks (e.g., the chroma blocks may or may not be located at the top CTU boundary).
[0194] iv. In one example, the above method can be applied only to images / videos in 4:2:2 format.
[0195] b. In one example, when the current chroma block is not at the top CTU boundary, the upper adjacent luminance sample, which includes the nearest upper line of the co-located luminance block but does not include the second upper line, can be used to derive the co-located adjacent top luminance sample of the downsampled chroma block.
[0196] c. In one example, the derivation of the co-occurring top brightness samples of the downsampled data can depend on the samples located on multiple lines.
[0197] i. In one example, it could depend on both the second nearest line and the nearest line above the co-bit luminance block.
[0198] ii. In one example, for different color formats (e.g., 4:2:0 and 4:2:2), the same downsampling filter can be used to derive the co-occurring top luminance samples of the downsampled color.
[0199] 1) In one example, a 6-tap filter can be used (e.g., [1 2 1; 1 2 1]).
[0200] a) In one example, the downsampled co-located adjacent top luminance samples can be derived as: pDsY[x]=(pY[2*x–1][-2]+2*pY[2*x][-2]+pY[2*x+1][-2]+pY[2*x–1][-1]+2*pY[2*x][-1]+pY[2*x+1][-1]+4)>>3, where pY is the corresponding luminance sample and pDstY[x] represents the downsampled co-located adjacent top luminance samples.
[0201] b) Alternatively, the above method can be applied when sps_cclm_colocated_chroma_flag equals 0.
[0202] 2) In one example, a 5-tap filter can be used (e.g., [0 1 0; 1 4 1; 0 1 0]).
[0203] a) In one example, the downsampled co-located adjacent top luminance sample can be derived as: pDsY[x]=(pY[2*x][-2]+pY[2*x-1][-1]+4*pY[2*x][-1]+pY[2*x+1][-1]+pY[2*x][0]+4)>>3, where pY is the corresponding luminance sample and pDstY[x] represents the downsampled co-located adjacent top luminance sample.
[0204] b) Alternatively, the above method can be applied when sps_cclm_colocated_chroma_flag equals 1.
[0205] iii. In one example, the above method may only be applied to images / videos in 4:2:2 format.
[0206] Transform the maximum block size of skipped codec blocks (For example, where transform_skip_flag equals 1, or BDPCM or bypass transformation process / other modes using identity transform)
[0207] 2. The maximum block size for transform skip codec blocks can depend on the color components. Let MaxTsSizeY and MaxTsSizeC represent the maximum block sizes for luminance and chrominance transform skip codec blocks, respectively.
[0208] a. In one example, the maximum block size for the luminance and chrominance components can be different.
[0209] b. In one example, the maximum block size of the two chromaticity components can be different.
[0210] c. In one example, the maximum block size of the luminance and chrominance components, or the maximum block size of each color component, can be signaled separately.
[0211] i. In one example, MaxTsSizeC / MaxTsSizeY can be signaled at the sequence level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header.
[0212] ii. In one example, MaxTsSizeY can be conditionally signaled, such as whether transformation skip is enabled or whether BDPCM is enabled.
[0213] iii. In one example, MaxTsSizeC can conditionally signal whether skipping based on color format / transformation is enabled or whether BDPCM is enabled.
[0214] iv. Alternatively, predictive coding and decoding can be used between the maximum block sizes of the luminance and chrominance components.
[0215] d. In one example, MaxTsSizeC can depend on MaxTsSizeY.
[0216] i. In one example, MaxTsSizeC can be set to be equal to MaxTsSizeY.
[0217] ii. In one example, MaxTsSizeC can be set to equal MaxTsSizeY / N (where N is an integer). For example, N = 2.
[0218] e. In one example, MaxTsSizeC can be set based on the chroma subsampling rate.
[0219] i. In one example, MaxTsSizeC is set to be equal to MaxTsSizeY >> SubWidthC, where SubWidthC is defined in the table.
[0220] ii. In one example, MaxTsSizeC is set to be equal to MaxTsSizeY >> SubHeightC, where SubHeightC is defined in the table.
[0221] iii. In one example, MaxTsSizeC is set to be equal to MaxTsSizeY >> max(SubWidthC, SubHeightC).
[0222] iv. In one example, MaxTsSizeC is set to be equal to MaxTsSizeY >> min(SubWidthC, SubHeightC).
[0223] 3. The maximum allowed block size width and height of the transform codec block can be defined differently.
[0224] a. In one example, the maximum allowed block size width and height can be notified separately by signaling.
[0225] b. In one example, the maximum allowed block size width and height of the chroma transform codec block can be represented as MaxTsSizeWC and MaxTsSizeHC, respectively. MaxTsSizeWC can be set to be equal to MaxTsSizeY >> SubWidthC, and MaxTsSizeHC can be set to be equal to MaxTsSizeY >> SubHeightC.
[0226] i. In one example, MaxTsSizeY is defined in bullet point 2.
[0227] 4. Whether signaling is used to notify the chroma block of the transform skip flag (e.g., transform_skip_flag[x0][y0][1] and / or transform_skip_flag[x0][y0][2]) may depend on the maximum allowed size of the chroma transform skip codec block.
[0228] a. In one example, the chroma transformation skip flag can be conditionally signaled based on the following conditions:
[0229] i. In one example, the conditions are: tbW is less than or equal to MaxTsSizeC, tbH is less than or equal to MaxTsSizeC, where tbW and tbH are the width and height of the current chroma block.
[0230] 1) In one example, MaxTsSizeC is defined as shown in bullet points 2-3.
[0231] ii. In one example, the conditions are: tbW is less than or equal to MaxTsSizeWC, tbH is less than or equal to MaxTsSizeHC, where tbW and tbH are the width and height of the current chroma block, and MaxTsSizeWC and MaxTsSizeHC represent the maximum allowed block size width and height for skipping encoding / decoding blocks in the chroma transformation, respectively.
[0232] 1) In one example, MaxTsSizeWC and / or MaxTsSizeHC can be defined as shown in bullet point 3.
[0233] b. In one example, by replacing "transform skip" with "BDPCM", the above method can be applied to the encoding and decoding of chroma BDPCM flags (e.g., intra_bdpcm_chroma_flag).
[0234] 5. Instead of encoding and decoding two TS tags for two chroma color components, it is proposed to use a single syntax to indicate the use of TS for two chroma color components.
[0235] a. In one example, instead of encoding / decoding transform_skip_flag[x0][y0][1] and / or transform_skip_flag[x0][y0][2], a single syntax element (e.g., TS_chroma_flag) can be encoded / decoded.
[0236] i. In one example, the value of a single syntax element is a binary value.
[0237] 1) Alternatively, in addition, the two chroma component blocks share the same on / off control of the TS mode based on a single syntax element.
[0238] a) In one example, a value of 0 for a single syntax element indicates that TS is disabled for both.
[0239] b) In one example, a value of 0 for a single syntax element indicates that TS is enabled for both.
[0240] 2) Alternatively, the second syntax element can be further signaled based on whether the value of a single syntax element is equal to K (e.g., K=1).
[0241] a) In one example, a single syntax element value of 0 indicates that TS is disabled for both; a single syntax element value of 0 indicates that TS is enabled for at least one of the two chroma components.
[0242] b) The second syntax element can be used to indicate which of the two chromaticity components TS is applied to and / or whether TS is applied to both of the two chromaticity components.
[0243] ii. In one example, the value of a single syntax element is a non-binary value.
[0244] 1) In one example, a single syntax element value equal to K0 indicates that TS is disabled for both.
[0245] 2) In one example, the value of a single syntax element equals K1, indicating that TS is enabled for the first chromaticity color component and disabled for the second color component.
[0246] 3) In one example, the value of a single syntax element equal to K2 indicates that TS is disabled for the first chromaticity color component and enabled for the second color component.
[0247] 4) In one example, the value of a single syntax element equals K3, indicating that TS is enabled for both.
[0248] 5) In one example, a single syntax element can be encoded or decoded using fixed-length, unary, truncated unary, and k-order EG binary methods.
[0249] iii. In one example, a single syntax element and / or a second syntax element may be either context-coded or bypass-coded.
[0250] General Declaration
[0251] 6. Whether and / or how the methods disclosed above can be signaled at the sequence level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header.
[0252] 7. Whether and / or how the methods disclosed above are applied may depend on encoding / decoding information, such as color format and single / dual tree splitting.
[0253] 5. Examples
[0254] This section illustrates example embodiments and how the current VVC standard has been modified to describe these embodiments. Changes to the VVC specification are highlighted in bold and italics. Deleted text is marked with double brackets (e.g., [[a]] indicates the deletion of the character "a").
[0255] 5.1. Example 1
[0256] The working draft specified in JVET-P2001-v9 can be modified as follows.
[0257] 8.4.5.2.13 Specifications for INTRA_LT_CCLM, INTRA_L_CCLM, and INTRA_T_CCLM Intra-Prediction Modes
[0258] …
[0259] 3. The derivation of the downsampled co-position brightness sample points pDsY[x][y], where x=0..nTbW-1 and y=0..nTbH-1, is as follows:
[0260] –If both SubWidthC and SubHeightC are equal to 1, then the following applies:
[0261] –pDsY[x][y], where x=1..nTbW-1, y=1..nTbH-1, is derived as follows:
[0262] pDstY[x][y]=pY[x][y] (8-159)
[0263] –Otherwise, the following applies:
[0264] – The one-dimensional filter coefficient arrays F1 and F2, and the two-dimensional filter coefficient arrays F3 and F4 are specified as follows.
[0265] F1[i]=1, where i=0..1 (8-160)
[0266] F2[0]=1,F2[1]=2,F2[2]=1 (8-161)
[0267] F3[i][j]=F4[i][j]=0, where i=0..2, j=0..2 (8-162)
[0268] –If both SubWidthC and SubHeightC are equal to 2, then the following applies:
[0269] F1[0]=1,F1[1]=1 (8-163)
[0270] F3[0][1]=1,F3[1][1]=4,F3[2][1]=1,F3[1][0]=1,F3[1][2]=1 (8-164)
[0271] F4[0][1]=1, F4[1][1]=2, F4[2][1]=1 (8-165)
[0272] F4[0][2]=1, F4[1][2]=2, F4[2][2]=1 (8-166)
[0273] –Otherwise, the following applies:
[0274] F1[0] = 2, F1[1] = 0 (8-167)
[0275] F3[1][1]=8 (8-168)
[0276] F4[0][1]=2, F4[1][1]=4, F4[2][1]=2, (8-169)
[0277] …
[0278] 5. When numSampT is greater than 0, the selected adjacent top chromaticity sample pSelC[idx] is set to equal to p[pickPosT[idx cntL]][-1], where idx = cntL..cntL+cntT-1, and the downsampled adjacent top luminance sample pSelDsY[idx], where idx = 0..cntL+cntT-1, is specified as follows:
[0279] …
[0280] – Otherwise (sps_cclm_colocated_chroma_flag equals 0), the following applies:
[0281] – If x is greater than 0, the following applies:
[0282] –If bCTUboundary equals FALSE, then the following applies:
[0283]
[0284]
[0285] – Otherwise (bCTUboundary equals TRUE), the following applies:
[0286]
[0287] – Otherwise (x equals 0), the following applies:
[0288] – If availTL equals TRUE and bCTUboundary equals FALSE, then the following applies:
[0289]
[0290] – Otherwise, if availTL equals TRUE and bCTUboundary equals TRUE, then the following applies:
[0291] pSelDsY[idx]=(F2[0]*pY[-1][-1]+F2[1]*pY[0][-1]+F2[2]*pY[1][-1]+2)>>2(8-196)
[0292] – Otherwise, if availTL equals FALSE and bCTUboundary equals FALSE, then the following applies:
[0293] pSelDsY[idx]=(F1[1]*pY[0][-2]+F1[0]*pY[0][-1]+1)>>1 (8-197)
[0294] – Otherwise (availTL equals FALSE and bCTUboundary)
[0295] If TRUE is true, then the following applies:
[0296] pSelDsY[idx]=pY[0][-1] (8-198)
[0297] …
[0298] 5.2. Example 2
[0299] This embodiment illustrates an example of chroma transform skip flag encoding / decoding based on the maximum permissible transform skip codec block size. The working draft specified in JVET-P2001-v9 can be modified as follows.
[0300] 7.3.9.10 Transformation Unit Syntax
[0301] …
[0302]
[0303]
[0304] …
[0305] 5.3. Example 3
[0306] This embodiment illustrates an example of chroma BDPCM flag encoding / decoding that skips the codec block size based on the maximum permissible chroma transformation. The working draft specified in JVET-P2001-v9 can be modified as follows.
[0307] 7.3.9.5 Encoding / Decoding Unit Syntax
[0308] …
[0309]
[0310] …
[0311] Figure 9This is a block diagram of a video processing apparatus 900. Apparatus 900 can be used to implement one or more of the methods described herein. Apparatus 900 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 900 may include one or more processors 902, one or more memories 904, and video processing hardware 906. The processors 902(s) may be configured to implement one or more methods described herein. The memories 904(s) may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 906 may be used to implement some of the techniques described herein (e.g., those listed in the previous section) in hardware circuitry. In some embodiments, hardware 906 may be partially or wholly within processor 902, such as a graphics processor.
[0312] Figure 10 A block diagram of a video encoder is shown.
[0313] Figure 11 This is a flowchart of method 1100 for processing video. Method 1100 includes a conversion between a chroma block for the video and a codec representation of the video, deriving parameters of a cross-component linear model by using downsampled co-occurring top luminance samples generated from N upper adjacent lines of the co-occurring luminance block using a downsampling filter, where N is a positive integer, and performing (1104) the conversion with a predicted chroma block generated using the cross-component linear model.
[0314] Figure 12 This is a block diagram of an example video processing system that can implement the disclosed technology.
[0315] Figure 12 This is a block diagram illustrating an example video processing system 1200, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 1200. System 1200 may include an input 1202 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1202 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0316] System 1200 may include codec component 1204, which can implement the various codec or encoding methods described in this document. Codec component 1204 can reduce the average bit rate of the video from input 1202 to the output of codec component 1204 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1206, the output of codec component 1204 can be stored or transmitted via connected communication. Component 1208 can use the stored or transmitted bitstream (or codec) representation of the video received at input 1202 to generate pixel values or displayable video sent to display interface 1210. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that encoding tools or operations are used at the encoder, and corresponding decoding tools or operations, the opposite of the encoding result, are performed by the decoder.
[0317] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0318] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from video blocks to a bitstream representation of video will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of video to video blocks will be performed using the video processing tool or mode enabled based on that decision or determination.
[0319] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion of video blocks to a bitstream representation of video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that it has not been modified using a video processing tool or mode enabled based on the decision or determination.
[0320] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed herein and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a combination of a machine-readable storage device, a machine-readable storage substrate, a memory device, a substance that implements a machine-readable propagating signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagating signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiving device.
[0321] A computer program (also called a program, software, software application, script, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer, or on multiple computers located in one location or distributed across multiple locations and interconnected through a communication network.
[0322] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0323] For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented or incorporated into dedicated logic circuitry.
[0324] Figure 13 This is a block diagram illustrating an example video encoding / decoding system 100 that can utilize the technology of the present invention.
[0325] like Figure 13 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 decodes the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0326] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0327] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage media / server 130b for access by destination device 120.
[0328] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0329] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to connect to an external display device.
[0330] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or further standards.
[0331] Figure 14 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 13 The video encoder 114 in the system 100 shown.
[0332] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 14 In the example, the video encoder 200 includes multiple functional components. The techniques described in this invention can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0333] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0334] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference image is the picture containing the current video block.
[0335] Furthermore, some components (such as motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but for interpretative purposes... Figure 14 The example is shown separately.
[0336] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0337] The mode selection unit 203 can select one of the encoding / decoding modes (intra-frame or inter-frame) based, for example, on the error result, and provide the resulting intra-frame or inter-frame encoded / decoded blocks to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded blocks for use as reference images. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vectors for the blocks (e.g., sub-pixel or integer pixel precision).
[0338] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.
[0339] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0340] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating a reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the motion vector of the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0341] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference images in lists 0 and 1 containing the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0342] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.
[0343] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block to another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.
[0344] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0345] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0346] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0347] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0348] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0349] In other examples, there may be no residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform the subtraction operation.
[0350] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0351] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0352] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is then stored in buffer 213.
[0353] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0354] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.
[0355] Figure 15 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 13 The video decoder 124 in the system 100 shown.
[0356] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 15 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0357] exist Figure 15 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform functions typically associated with the video encoder 200 (e.g., Figure 14 The encoding process described is the opposite of the decoding process.
[0358] Entropy decoding unit 301 can retrieve encoded bitstreams. The encoded bitstreams may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine this information, for example, by executing AMVP and merge modes.
[0359] Motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The syntax elements may include an identifier for the interpolation filter to be used with sub-pixel precision.
[0360] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[0361] The motion compensation unit 302 may use some syntax information to determine the size of the blocks of (multiple) frames and / or (multiple) stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0362] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. Inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies the inverse transform.
[0363] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation.
[0364] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the codec will use or implement the tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream representation of the video will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to a video block will be performed using the video processing tool or mode enabled based on that decision or determination.
[0365] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-located or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the error residual values after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream.
[0366] The following is a list of preferred terms for some embodiments.
[0367] The first set of clauses illustrates example embodiments of the techniques discussed in the preceding chapters.
[0368] 1. A video processing method comprising: a conversion between a chroma block of a video and a codec representation of the video; deriving parameters of a cross-component linear model by using downsampled co-occurring top luminance samples generated from N upper adjacent lines of the co-occurring luminance block using a downsampling filter, wherein N is a positive integer; and performing the conversion using a predicted chroma block generated using the cross-component linear model.
[0369] 2. The method according to Clause 1, wherein, since the chroma block is not at the boundary of the top codec tree unit, the N upper adjacent lines correspond to the nearest upper line of the co-bit luma block.
[0370] 3. The method according to any one of Clauses 1-2, wherein a downsampling filter is further applied to generate downsampled co-occurring left-side luminance samples.
[0371] 4. The method according to any one of Clauses 1-2, wherein the downsampling filter is different from another downsampling filter used to generate co-occurring adjacent left-hand luminance samples for downsampling.
[0372] 5. The method according to any one of Clauses 1, wherein the downsampling filter is independent of the position of the chroma block relative to the top boundary of the codec tree unit.
[0373] 6. The method according to any one of Clause 1, wherein the method is selectively applied because the video has a 4:2:2 format.
[0374] 7. The method described in Clause 1, wherein N is greater than one.
[0375] 8. The method according to Clause 7, wherein the N adjacent upper lines include the nearest upper line and the second nearest upper line.
[0376] 9. The method according to Clause 1, wherein the downsampling filter depends on the color format of the video.
[0377] 10. The method according to any one of clauses 1-9, wherein the downsampling filter is a 6-tap filter.
[0378] 11. The method according to any one of clauses 1-9, wherein the downsampling filter is a 5-tap filter.
[0379] 12. The method according to any one of clauses 1 to 11, wherein the conversion includes encoding the video into a codec representation.
[0380] 13. The method according to any one of clauses 1 to 11, wherein the conversion includes decoding the encoding / decoding representation to generate pixel values of the video.
[0381] 14. A video decoding apparatus, comprising a processor configured to implement one or more of the methods described in clauses 1 to 13.
[0382] 15. A video encoding apparatus, comprising a processor configured to implement one or more of the methods described in clauses 1 to 13.
[0383] 16. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of clauses 1 to 13.
[0384] 17. A method, apparatus or system as described in this document.
[0385] The second set of clauses describes certain features and aspects of the technology disclosed in the previous chapters (e.g., Item 1).
[0386] 1. A video processing method comprising: a conversion between a chroma block of a video and a bitstream representation of the video, deriving parameters of a 1602 cross-component linear model by using downsampled luminance samples generated from N upper adjacent lines of a co-occurring luminance block of the chroma block using a downsampling filter, wherein N is a positive integer; and performing the conversion using a predicted chroma block generated using the cross-component linear model.
[0387] 2. The method according to Article 1, wherein, since the chroma block is not at the boundary of the top codec tree unit, the N upper adjacent lines correspond to the nearest upper line of the co-bit luminance block.
[0388] 3. The method according to any one of Clauses 1-2, wherein a downsampling filter is further applied to generate additional downsampled luminance samples, which are generated from the left adjacent line of the co-occurring luminance block.
[0389] 4. The method according to any one of Clauses 1-2, wherein another downsampling filter is applied to generate additional downsampled luminance samples, which are generated from the left adjacent line of the co-occurring luminance block.
[0390] 5. The method according to any one of clauses 1-4, wherein the downsampling filter has filter coefficients [1, 2, 1].
[0391] 6. The method according to any one of clauses 1-5, wherein the downsampled luminance sample pDsY[x] satisfies the equation pDsY[x]=(pY[2*x–1][-1]+2*pY[2*x][-1]+pY[2*x+1][-1]+2)>>2, where pY[2*x][-1], pY[2*x–1][-1] and pY[2*x+1][-1] are luminance samples from the nearest upper adjacent line, and x is an integer.
[0392] 7. The method according to any one of clauses 1-6, wherein the downsampling filter is independent of the position of the chroma block relative to the top boundary of the codec tree unit.
[0393] 8. The method according to any one of clauses 1-6, wherein the method is selectively applied due to the 4:2:2 color format of the video.
[0394] 9. The method according to Clause 1, wherein, since the chroma block is not at the top codec tree unit boundary, the N upper adjacent lines include the nearest upper line of the co-bit luma block, but do not include the second nearest upper line.
[0395] 10. The method described in Clause 1, wherein N is greater than one.
[0396] 11. The method according to Clause 10, wherein the N adjacent upper lines include the nearest upper line and the second nearest upper line.
[0397] 12. The method according to Clause 1, wherein the downsampling filter depends on the color format of the video.
[0398] 13. The method according to any one of clauses 1-12, wherein the downsampling filter is a 6-tap filter.
[0399] 14. The method according to any one of clauses 1-12, wherein the downsampling filter is a 5-tap filter.
[0400] 15. The method according to any one of clauses 1 to 14, wherein the conversion includes encoding the video into a bitstream representation.
[0401] 16. The method according to any one of clauses 1 to 14, wherein the conversion includes decoding video from a bitstream representation.
[0402] 17. A video processing apparatus, comprising a processor configured to implement the method described in any one or more of the provisions 1 to 16.
[0403] 18. A computer-readable medium storing program code that, when executed, causes a processor to perform any one or more of the methods described in clauses 1 to 16.
[0404] 19. A computer-readable medium for storing a bitstream representation generated according to any of the above methods.
[0405] The third set of clauses describes certain features and aspects of the technology disclosed in the previous chapters (e.g., items 2 to 7).
[0406] 1. A video processing method (e.g., such as...) Figure 16A The method shown (1610) includes: a conversion between a video region of a video component and a bitstream representation of the video; determining 1612 the maximum allowed block size of a video block encoded using a transform skip mode; and performing a conversion 1614 based on the determination.
[0407] 2. The method according to Clause 1, wherein the transform skipping mode includes encoding and decoding the residual of a video block during encoding without applying a non-identity transform, or determining the decoded video block during decoding without applying a non-identity inverse transform to the residual encoded and decoded in the bitstream representation.
[0408] 3. The method according to Clause 1, wherein the transform skip mode includes BDPCM (Block Differential Pulse Modulation), which corresponds to an intra-frame codec tool that uses Differential Pulse Modulation (DPCM) at the block level.
[0409] 4. The method according to Clause 1, wherein the maximum allowed block size depends on whether the block skipped by the transformation is a chroma block or a luma block.
[0410] 5. The method according to Clause 1, wherein the maximum allowed block size depends on the chromaticity components of the blocks skipped by the transformation.
[0411] 6. The method according to Clause 1, wherein the maximum allowed block size (MaxTsSizeY) of the luminance block and the maximum allowed block size (MaxTsSizeC) of the chrominance block are separately signaled in the bitstream representation.
[0412] 7. The method described in Clause 6, wherein MaxTsSizeC and / or MaxTsSizeY are notified at the sequence level, picture level, strip level, or slice level signaling.
[0413] 8. The method according to Clause 6, wherein MaxTsSizeY is conditionally signaled based on the enabled state of the transformation skip mode.
[0414] 9. The method according to Clause 6, wherein MaxTsSizeY is conditionally signaled based on the enabled status of the color format and / or transformation skip mode.
[0415] 10. The method according to Clause 1, wherein the conversion is performed by using predictive encoding / decoding between the maximum block sizes of the luminance and chrominance components.
[0416] 11. The method according to Clause 1, wherein the video block is a chroma video block, and wherein the maximum permissible block size (MaxTsSizeC) of the video block depends on the maximum permissible block size (MaxTsSizeY) of another video block that is a luminance component.
[0417] 12. The method according to Clause 11, wherein MaxTsSizeC is set to be equal to MaxTsSizeY.
[0418] 13. The method according to Clause 11, wherein MaxTsSizeC is set to be equal to MaxTsSizeY / N, where N is an integer.
[0419] 14. The method according to Clause 1, wherein the video block is a chroma video block, and wherein the maximum permissible block size (MaxTsSizeC) of the video block is set according to the chroma subsampling rate.
[0420] 15. The method according to Clause 14, wherein MaxTsSizeC is set to equal i) MaxTsSizeY >> SubWidthC, ii) MaxTsSizeY >> SubHeightC, iii) MaxTsSizeY >> max(SubWidthC, SubHeightC), iv) MaxTsSizeY >> min(SubWidthC, SubHeightC), wherein MaxTsSizeY indicates the maximum block size of the luminance video block, and SubWidthC and SubHeightC are predefined.
[0421] 16. A video processing method (e.g., such as...) Figure 16A The method 1610 shown includes: performing a conversion between a video block and a bitstream representation of the video, according to a first rule and a second rule. A transform skip codec tool is used to encode and decode a first portion of the video block, and a transform codec tool is used to encode and decode a second portion of the video block. The first rule specifies the maximum allowed block size for the first portion of the video block, the second rule specifies the maximum allowed block size for the second portion of the video block, and the maximum allowed block size for the first portion of the video block differs from the maximum allowed block size for the second portion of the video block.
[0422] 17. The method according to Clause 16, wherein the maximum permissible block size corresponds to the width and height of the corresponding block.
[0423] 18. The method according to Clause 17, wherein the width and height of the maximum permissible block size are signaled separately.
[0424] 19. The method according to Clause 17, wherein, for the second part of a video block that is a chroma block, the width (MaxTsSizeWC) is set to be equal to MaxTsSizeY >> SubWidthC, and the height (MaxTsSizeHC) is set to be equal to MaxTsSizeY >> SubHeightC, wherein MaxTsSizeY indicates the maximum permissible block size of the luminance block.
[0425] 20. A video processing method (e.g., such as...) Figure 16AThe method shown (1610) includes: performing a conversion between a video comprising one or more chroma blocks and a bitstream representation of the video. The bitstream representation conforms to a format rule that specifies whether a syntax element indicating the use of a transform skip tool is included in the bitstream representation depends on the maximum allowed size of the chroma blocks encoded and decoded using the transform skip tool.
[0426] 21. The method according to Clause 20, wherein the transformation skipping tool includes bypass transformation or application of identity transformation.
[0427] 22. The method according to Clause 20, wherein signaling notifies the syntax element when tbW is less than or equal to MaxTsSizeC and tbH is less than or equal to MaxTsSizeC, wherein tbW and tbH are the width and height of the chroma block, respectively, and MaxTsSizeC is the maximum permissible size of the chroma block.
[0428] 23. The method according to Clause 20, wherein, when tbW is less than or equal to MaxTsSizeWC and tbH is less than or equal to MaxTsSizeHC, the signaling notifies the syntax element, wherein tbW and tbH are the width and height of the chroma block, respectively, and MaxTsSizeWC and MaxTsSizeHC represent the width and height of the maximum permissible size of the chroma block, respectively.
[0429] 24. The method according to Clause 20, wherein the transform skipping tool includes a BDPCM (Block Differential Pulse Modulation) mode, the BDPCM mode corresponding to an intra-frame codec tool that uses differential pulse modulation (DPCM) at the block level.
[0430] 25. A video processing method (e.g., such as...) Figure 16A The method shown (1610) includes: performing a conversion between a video and a bitstream representation of a video comprising one or more first video blocks including a first chroma component and one or more second video blocks including a second chroma component. The bitstream representation conforms to a format rule specifying a syntax element to be used, which jointly indicates the availability of a transform skip tool for encoding and decoding one or more first chroma blocks and one or more second chroma blocks.
[0431] 26. The method described in Clause 25, wherein the syntax element has a binary value.
[0432] 27. The method according to Clause 25, wherein the transform skip tool is enabled or disabled in one or more first video blocks and one or more second video blocks according to the syntax element.
[0433] 28. The method according to Clause 25, wherein the formatting rule further specifies that additional syntax elements are included in the bitstream representation based on whether the value of a syntax element is equal to K, where K is an integer.
[0434] 29. The method according to Clause 28, wherein the second syntax element is used to indicate which of the one or more first video blocks and one or more second video blocks the transform skip tool is applied to.
[0435] 30. The method described in Clause 25, wherein the syntax element has a non-binary value.
[0436] 31. The method described in Article 30, wherein the syntax elements are encoded or decoded using a fixed-length, unary, truncated unary, or k-order Exp-Golomb (EG) binary method.
[0437] 32. The method described in accordance with Clause 25, wherein the syntax element is context-coded or bypass-coded.
[0438] 33. The method according to any one of the preceding clauses, wherein whether and / or how the method is applied is signaled at the sequence level, picture level, strip level, or slice level.
[0439] 34. The method according to any one of the preceding clauses, wherein the method is further based on encoded / decoded information.
[0440] 35. The method according to any one of clauses 1 to 34, wherein the conversion includes encoding the video into a bitstream representation.
[0441] 36. The method according to any one of clauses 1 to 34, wherein the conversion includes decoding video from a bitstream representation.
[0442] 37. A video processing apparatus, comprising a processor configured to implement any one or more of the methods described in clauses 1 to 36.
[0443] 38. A computer-readable medium storing program code that, when executed, causes a processor to perform any one or more of the methods described in clauses 1 to 36.
[0444] 39. A computer-readable medium storing a codec representation or bitstream representation generated according to any of the above methods.
[0445] While this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features characteristic of specific embodiments of a particular technology. Certain features described in the context of individual embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0446] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0447] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
Claims
1. A video processing method, comprising: For the conversion between video regions of video components and the bitstream of the video, determine the maximum allowed block size of video blocks encoded and decoded using transform skip mode; and The conversion is performed based on the determination; in, The maximum allowed block size depends on the chroma components of the blocks skipped by the transform. Specifically, based on the enabled state of the transformation skip mode, the maximum allowed block size MaxTsSizeY of the luminance block is conditionally signaled.
2. The method according to claim 1, wherein, The transform skipping mode includes: during encoding, encoding and decoding the residuals of the video block without applying a non-identity transform, or during decoding, determining the video block to be decoded without applying a non-identity inverse transform to the residuals encoded and decoded in the bitstream.
3. The method according to claim 1, wherein, The transform skipping modes include BDPCM (Block Differential Pulse Modulation), which corresponds to intra-frame coding tools that use differential pulse modulation (DPCM) at the block level.
4. The method according to any one of claims 1-3, wherein, The MaxTsSizeY and the maximum allowed block size MaxTsSizeC for chroma blocks are separately signaled in the bitstream.
5. The method according to claim 4, wherein, The MaxTsSizeC and / or MaxTsSizeY are notified at the sequence level, picture level, strip level, or slice level signaling.
6. The method according to claim 4, wherein, Based on the color format and / or the enabled status of the transformation skip mode, the MaxTsSizeC is conditionally signaled.
7. The method according to any one of claims 1-3, wherein, The conversion is performed by using predictive encoding and decoding between the maximum block sizes of the luminance and chrominance components.
8. The method according to any one of claims 1-3, wherein, The video block is a chroma video block, and the maximum allowed block size MaxTsSizeC of the video block depends on the maximum allowed block size MaxTsSizeY of another video block of the luminance component.
9. The method according to claim 8, wherein, The MaxTsSizeC is set to be equal to the MaxTsSizeY.
10. The method according to claim 8, wherein, The MaxTsSizeC is set to be equal to the MaxTsSizeY / N, where N is an integer.
11. The method according to any one of claims 1-3, wherein, The video block is a chroma video block, and the maximum allowed block size MaxTsSizeC of the video block is set according to the chroma subsampling rate.
12. The method according to claim 11, wherein, The MaxTsSizeC is set to equal i) MaxTsSizeY >> SubWidthC, ii) MaxTsSizeY >> SubHeightC, iii) MaxTsSizeY >> max (SubWidthC, SubHeightC), iv) MaxTsSizeY >> min (SubWidthC, SubHeightC), where MaxTsSizeY indicates the maximum block size of the luminance video block, and SubWidthC and SubHeightC are predefined.
13. The method according to claim 1, further comprising: Perform a conversion between a video comprising one or more first video blocks with a first chroma component and one or more second video blocks with a second chroma component, and the bitstream of said video. The bitstream therein conforms to a format rule specifying the syntax elements to be used, which together indicate the availability of a transform skip tool for encoding and decoding the one or more first video blocks and the one or more second video blocks.
14. The method according to claim 13, wherein, The syntax elements have binary values.
15. The method according to claim 13 or 14, wherein, The transform skip tool can be enabled or disabled in one or more first video blocks and one or more second video blocks according to the syntax elements described therein.
16. The method according to claim 13 or 14, wherein, The formatting rules also specify that additional syntax elements are included in the bitstream based on whether the value of the syntax element is equal to K, where K is an integer.
17. The method according to claim 16, wherein, The second syntax element is used to indicate which of the one or more first video blocks and one or more second video blocks the transform skip tool is applied to.
18. The method according to claim 13, wherein, The syntax element has a non-binary value.
19. The method according to claim 18, wherein, The syntax elements are encoded and decoded using fixed-length, unary, truncated unary, or k-order Exp-Golomb (EG) binary methods.
20. The method according to any one of claims 13-14, 18, and 19, wherein, The syntax elements are either context-coded or bypass-coded.
21. A video processing method, comprising: The conversion between the video, including video blocks, and the bitstream of the video is performed according to the first and second rules. The transform skip encoding / decoding tool is used to encode and decode the first part of the video block. The transform codec tool is used to encode and decode the second part of the video block. The first rule specifies the maximum allowed block size of the first portion of the video block, and the second rule specifies the maximum allowed block size of the second portion of the video block. The maximum allowed block size of the first portion of the video block is different from the maximum allowed block size of the second portion of the video block; The maximum allowable block size of the first part depends on the chromaticity components of the blocks skipped by the transform; Specifically, based on the enabled state of the transformation skip codec tool, the maximum allowed block size MaxTsSizeY of the luminance block in the first part is conditionally signaled.
22. The method according to claim 21, wherein, The maximum allowed block size corresponds to the width and height of the corresponding block.
23. The method according to claim 22, wherein, The width and height of the maximum allowed block size are separately signaled.
24. The method according to claim 22 or 23, wherein, For the second portion of the video block that is a chroma block, the width MaxTsSizeWC is set to be equal to MaxTsSizeY >> SubWidthC, and the height MaxTsSizeHC is set to be equal to MaxTsSizeY >> SubHeightC.
25. A video processing method, comprising: Perform a conversion between a video comprising one or more chroma blocks and the bitstream of that video. The bitstream conforms to a format rule that specifies whether a syntax element indicating the use of a transform skip tool is included in the bitstream depends on the maximum allowed size of the chroma block encoded or decoded using the transform skip tool. The maximum allowable size depends on the chroma components of the blocks skipped by the transform. Specifically, based on the enabled state of the transformation skip encoding / decoding tool, the maximum allowed block size MaxTsSizeY of the luma block is conditionally signaled.
26. The method of claim 25, wherein, The transformation skipping tool includes bypass transformation or application of identity transformation.
27. The method according to claim 25 or 26, wherein, When tbW is less than or equal to MaxTsSizeC and tbH is less than or equal to MaxTsSizeC, the signaling notifies the syntax element, where tbW and tbH are the width and height of the chroma block, respectively, and MaxTsSizeC is the maximum allowed size of the chroma block.
28. The method according to claim 25 or 26, wherein, When tbW is less than or equal to MaxTsSizeWC and tbH is less than or equal to MaxTsSizeHC, the signaling notifies the syntax element, where tbW and tbH are the width and height of the chroma block, respectively, and MaxTsSizeWC and MaxTsSizeHC represent the width and height of the maximum allowed size of the chroma block, respectively.
29. The method according to claim 25 or 26, wherein, The transform skipping tool includes the BDPCM (Block Differential Pulse Modulation) mode, which corresponds to the intra-frame codec tool that uses differential pulse modulation (DPCM) at the block level.
30. The method according to any one of claims 1-3, 13-14, 18-19, 21-23, and 25-26, wherein, Whether and / or how the method is applied is signaled at the sequence level, picture level, strip level, or slice level.
31. The method according to any one of claims 1-3, 13-14, 18-19, 21-23, and 25-26, wherein, The method is also based on encoding and decoding information.
32. The method according to any one of claims 1-3, 13-14, 18-19, 21-23, and 25-26, wherein, The conversion includes encoding the video into the bitstream.
33. The method according to any one of claims 1-3, 13-14, 18-19, 21-23, and 25-26, wherein, The conversion includes decoding the video from the bitstream.
34. A video processing apparatus comprising a processor configured to implement the method of any one of claims 1 to 33.
35. A computer-readable medium storing program code, which, when executed, causes a processor to perform the method of any one of claims 1 to 33.
Citation Information
Patent Citations
Transform coding of video data
WO2018171751A1