Method and apparatus for encoding and decoding pictures

By applying constraints to luma/chroma coding tree interdependencies and enabling/disabling tools like luma-dependent chroma residual scaling within authorized rectangular areas, the challenges of processing delays in decoder pipelines are addressed, improving coding efficiency and reducing structural delays in video decoding.

JP7720367B2Active Publication Date: 2025-08-07INTERDIGITAL VC HOLDINGS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023147275
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-02
Filing Date
2023-09-12
Publication Date
2025-08-07
Estimated Expiration
2040-02-25

AI Technical Summary

Technical Problem

The separation of luma and chroma coding trees in video coding systems poses challenges for hardware implementation, particularly in decoder pipelines, due to interdependencies between chroma and co-located luma samples, leading to large structural delays in processing chroma blocks.

Method used

Implement constraints on the use of separate luma/chroma coding trees and coding tools like luma-dependent chroma residual scaling, enabling or disabling these tools based on the size of luma and chroma blocks to ensure chroma blocks are processed efficiently within authorized rectangular areas (ARAs), thereby reducing pipeline delays.

Benefits of technology

This approach enhances coding efficiency for chroma components by optimizing processing delays and maintaining decoder pipeline efficiency, ensuring seamless decoding without significant coding loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007720367000010
    Figure 0007720367000010
  • Figure 0007720367000011
    Figure 0007720367000011
  • Figure 0007720367000012
    Figure 0007720367000012
Patent Text Reader

Abstract

To provide an encoding and decoding method that improves efficiency by enabling an inter-component dependency tool.SOLUTION: A decoding method includes activating an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block colocated with the chroma block, and decoding the chroma block on the basis of the determined luma value in response to the activation of the inter-component dependency tool.SELECTED DRAWING: Figure 25
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] 1.Technical Field At least one embodiment of the present invention relates generally to methods and apparatus for encoding and decoding pictures, and more particularly to methods and apparatus for encoding and decoding pictures with independent luma and chroma splitting. [Background technology]

[0002] 2. Background technology To achieve high compression efficiency, image and video coding schemes typically use prediction and transform to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit correlation within or between frames, and then the difference between the original and predicted image blocks, often denoted as prediction error or prediction residual, is transformed, quantized, and entropy coded. During encoding, the original image blocks are typically partitioned / divided into sub-blocks, possibly using quadtree partitioning. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to prediction, transformation, quantization, and entropy coding. Summary of the Invention

[0003] 3. A concise overview According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block co-located with the chroma block; decoding said chroma blocks in response to said enabling of said inter-component dependency tool; A method for decoding video data is presented, including:

[0004] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block that colocates with the chroma block; and decoding said chroma blocks in response to said enabling of said inter-component dependency tool; An apparatus for decoding video data is presented, the apparatus including one or more processors configured to:

[0005] According to another general aspect of at least one embodiment, a bitstream is formatted to include signals generated according to the encoding method described above.

[0006] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block co-located with the chroma block; encoding the chroma blocks in response to the enabling of the inter-component dependency tool. A method for encoding video data is presented, including:

[0007] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block that colocates with the chroma block; and encoding the chroma blocks in response to the enabling of the inter-component dependency tool. An apparatus for encoding video data is presented, the apparatus including one or more processors configured to:

[0008] According to another general aspect of at least one embodiment, a bitstream is formatted to include signals generated according to the encoding method described above.

[0009] According to another general aspect, the method for decoding or encoding video, or the apparatus for decoding or encoding video, further includes determining a location in a picture of a given sample location in a chroma block, determining a co-located luma block, which is a luma block including a luma sample that co-locates with the location in the chroma block, determining neighboring luma samples of the co-located luma block, determining luma values from the determined neighboring luma samples, determining a scale factor based on the determined luma values, and applying scaling of a residual of the chroma block according to the scale factor.

[0010] One or more embodiments of the present invention also provide a computer-readable storage medium storing instructions for encoding or decoding video data according to at least a portion of any of the above methods. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the above encoding methods. One or more embodiments also provide methods and apparatus for transmitting or receiving a bitstream generated according to the above encoding methods. One or more embodiments also provide a computer program product including instructions for performing at least a portion of any of the above methods. [Brief explanation of the drawings]

[0011] 4. A brief summary of the drawing [Figure 1] Denotes a coding tree unit (CTU). [Figure 2] 1 shows the coding tree units split into coding units, prediction units, and transform units. [Figure 3] The CTUs are divided into coding units using Quad-Tree Plus Binary-Tree (QTBT). [Figure 4] The CTUs are divided into coding units using Quad-Tree Plus Binary-Tree (QTBT). [Figure 5] Indicates the division mode of the coding unit defined by the Quad-Tree plus Binary-Tree coding tool. [Figure 6] Further partitioning modes of the coding unit are shown, such as asymmetric binary partitioning mode and ternary tree partitioning mode. [Figure 7] 1 illustrates a flow diagram of a method for enabling or disabling a cross-component dependency coding tool according to one embodiment. [Figure 8A] 8 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 7. [Figure 8B] 8 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 7. [Figure 8C] 8 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 7. [Figure 8D] 8 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 7. [Figure 9] 1 illustrates a flow diagram of a method for enabling or disabling a cross-component dependency coding tool according to one embodiment. [Figure 10] 1 illustrates an example chroma block and its co-located luma block. [Figure 11] 10 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 9. [Figure 11A] 1 shows a simplified block diagram of the activation of the CCLM process. [Figure 11B] 1 shows a simplified block diagram of an example of an index selection process from a history list. [Figure 11C] The process is illustrated by considering a luma ARA of size 32x32 and a chroma ARA of size 16x16. [Figure 11D]For example, consider a luma ARA of size 32x32 and a chroma ARA of size 16x16, and if a square luma block of size 64x64 is not divided, the corresponding chroma block (of size 32x32 in 4:2:0 format) can be divided by QT division only into four 16x16 blocks that can possibly be further divided (dotted lines in the figure) or cannot be divided at all. Examples that follow these constraints are shown below. [Figure 11E] For example, consider a luma ARA of size 32x32 and a chroma ARA of size 16x16, and if a square luma block of size 64x64 is divided into four 32x32 blocks, the corresponding chroma block (of size 32x32 in 4:2:0 format) can be divided by QT decomposition only into four 16x16 blocks that can possibly be further divided (dotted lines in the figure) or cannot be divided at all. Examples that comply with these constraints are shown below. [Figure 11F] For example, consider a luma ARA of size 32x32 and a chroma ARA of size 16x16, and if a square luma block of size 64x64 is divided into two blocks of 32 lines and 64 columns, the corresponding chroma block (of size 32x32 in 4:2:0 format) can only be divided into two blocks of 16 lines and 32 columns by horizontal BT division, or into four 16x16 blocks by QT division. [Figure 11G] For example, consider a luma ARA of size 32x32 and a chroma ARA of size 16x16, and if a square luma block of size 64x64 is divided into two blocks of 64 lines and 32 columns, the corresponding chroma block can only be divided into two blocks of 32 lines and 16 columns. [Figure 12] 1 illustrates a flow diagram of a method for enabling or disabling a cross-component dependency coding tool according to one embodiment. [Figure 13] 13 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 12. [Figure 14]13 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 12. [Figure 15] 1 illustrates a flow diagram of a method for enabling or disabling a cross-component dependency coding tool according to one embodiment. [Figure 16] 16 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to the embodiment of FIG. 15. [Figure 17] 1 illustrates a flow diagram of a method for obtaining a scale factor or scale factor index for a chroma block, according to one embodiment. [Figure 18] Shows the current chroma block and some neighboring chroma blocks. [Figure 19] 1 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to one embodiment. [Figure 20] 1 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to one embodiment. [Figure 21] 1 illustrates the principle of enabling / disabling the inter-component dependency coding tool according to one embodiment. [Figure 22] 1 illustrates a flow diagram of a method for determining scale factors used in chroma residual scaling, according to one embodiment. [Figure 23A] Chroma blocks and co-located luma blocks are shown. [Figure 23B] Chroma blocks and co-located luma blocks are shown. [Figure 23C] Chroma blocks and co-located luma blocks are shown. [Figure 23D] Chroma blocks and co-located luma blocks are shown. [Figure 24A] 1 shows an example of the distance between the top left neighboring luma sample and the neighboring chroma sample. [Figure 24B] 1 shows an example of the distance between the top left neighboring luma sample and the neighboring chroma sample. [Figure 25]1 shows a flow diagram of a method for checking the availability of luma samples based on its own co-located chroma samples and based on the chroma samples of a current block. [Figure 25A] The second square block is shown. [Figure 26] 1 shows a block diagram of a video encoder according to one embodiment; [Figure 27] 1 shows a block diagram of a video decoder according to one embodiment; [Figure 28] 1 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0012] 5. Detailed Description In HEVC coding, a picture is divided into square CTUs with a configurable size, typically 64x64, 128x128, or 256x256. As shown in Figure 1, a CTU is the root of a quadtree partition into four square coded units (CUs) of equal size, i.e., half the width and height of the parent block. A quadtree is a tree in which a parent node can be divided into four child nodes, and each child node can be the parent node of another division into four child nodes. In HEVC, coded blocks (CBs) are divided into one or more predictive blocks (PBs), which form the root of the quadtree partition into transform blocks (TBs). As shown in Figure 2, corresponding to coded blocks, predictive blocks, and transform blocks, coded units (CUs) include prediction units (PUs) and a set of tree-structured transform units (TUs), where a PU contains prediction information for all color components (e.g., intra- or inter-prediction parameters), and a TU contains a residual coding syntax structure for each color component. Intra or inter coding mode is specified for the CU. The sizes of the luma components CB, PB, and TB apply to the corresponding CU, PU, and TU.

[0013] In more recent coding systems, a CTU is the root of a coding tree division into coded units (CUs). A coding tree is a tree in which a parent node (usually corresponding to a block) can be divided into child nodes (e.g., two, three, or four child nodes), and each child node can be the parent node of another division into child nodes. In addition to the quadtree partitioning mode, new partitioning modes (binary tree symmetric partitioning mode, binary tree asymmetric partitioning mode, and ternary tree partitioning mode) are also defined to increase the total number of possible partitioning modes. A coding tree has a unique root node, e.g., a CTU. The leaves of the coding tree are the terminal nodes of the tree. Each node of the coding tree represents a block that can be further divided into smaller blocks, also named subblocks. Once the division of a CTU into CUs has been determined, the CUs corresponding to the leaves of the coding tree are coded. The division of a CTU into CUs and the coding parameters used to code each CU (corresponding to the leaves of the coding tree) can be determined at the encoder side by a rate-distortion optimization procedure.

[0014] Figure 3 illustrates the division of a CTU into CUs, which can be split according to both the quadtree and symmetric binary tree partitioning modes. This type of partitioning is known as QTBT (Quad-Tree plus Binary-Tree). The symmetric binary tree partitioning mode is defined to allow a CU to be split horizontally or vertically into two equally sized coded units. In Figure 3, solid lines indicate quadtree partitioning, and dotted lines indicate binary tree partitioning of a CU into symmetric CUs. Figure 4 illustrates the associated coding tree. In Figure 4, solid lines indicate quadtree partitioning, and dotted lines indicate binary partitioning spatially embedded within the leaves of the quadtree. Figure 5 illustrates the four partitioning modes used in Figure 3. NO_SPLIT mode indicates no further division of the CU. QT_SPLIT mode indicates that the CU is divided into four quadrants according to the quadtree, with the quadrants separated by two partition lines. HOR mode indicates that the CU is divided horizontally into two equally sized CUs separated by a single partition line. VER indicates that a CU is vertically divided into two CUs of equal size separated by a dividing line, which is represented by a dashed line in Figure 5.

[0015] In QTBT decomposition, CUs have square or rectangular shapes. The size of a coded unit is a power of two, typically ranging from 4 to 128 in both directions (horizontal and vertical). QTBT has several differences from HEVC decomposition. First, QTBT decomposition of a CTU consists of two stages. The CTU is first decomposed into a quadtree. Each leaf of the quadtree can then be further decomposed into a binary tree. This is illustrated on the right side of Figure 3, where the solid lines indicate the stages of the quadtree decomposition and the dotted lines represent the binary decomposition that is spatially embedded within the leaves of the quadtree.

[0016] Second, the partitioning structure of luma and chroma blocks, especially within a slice, can be separated and determined independently.

[0017] Additional types of partitioning can also be used. As shown in Figure 6, an asymmetric binary tree (ABT) partitioning mode is defined to allow a CU to be divided horizontally into two coded units with respective rectangular sizes (w, h / 4) and (w, 3h / 4), or to allow a CU to be divided vertically into two coded units with respective rectangular sizes (w / 4, h) and (3w / 4, h). The two coded units are separated by a single partition line, represented by the dashed line in Figure 6.

[0018] 6 also shows a ternary tree partitioning mode that splits one coding unit into three coding units in both the vertical and horizontal directions. In the horizontal direction, the CU is split into three coding units of respective sizes (w, h / 4), (w, h / 2), and (w, h / 4). In the vertical direction, the CU is split into three coding units of respective sizes (w / 4, h), (w / 2, h), and (w / 4, h).

[0019] The partitioning of the coded units is determined at the encoder side by rate-distortion optimization, which involves determining a representation of the CTU with the minimum rate-distortion cost.

[0020] In this application, the term "block" or "picture block" may be used to refer to any one of CTU, CU, PU, TU, CB, PB, and TB. In addition, the term "block" or "picture block" may be used to refer to macroblocks, partitions, and sub-blocks defined within H.264 / AVC or other video coding standards, and more broadly to refer to arrays of samples of various sizes.

[0021] In this application, the terms "reconstruct" and "decode" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "slice," "tile," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstruct" is used on the encoder side, while "decode" is used on the decoder side.

[0022] The use of these new topologies allows for significant improvements in coding efficiency. Specifically, significant gains are obtained for the chroma components. This gain arises from the separation / independence of the luma and chroma coding trees (or splits).

[0023] However, this separation of luma and chroma coding trees at the CTU level poses some problems for hardware implementation. The size of a CTU is typically 128x128 or 256x256. Completely separating the coding trees for the luma and chroma components implies that the luma and chroma components are also completely separate in the compressed domain and therefore appear differently in the coded bitstream. This causes some problems for decoder implementations that want to ensure that they can realize a decoding pipeline based on a maximum decoding unit size that may be smaller than the size of the CTU. Typically, a 64x64-based decoder pipeline is desirable in some decoder implementations. To do so, a maximum transform block size equal to 64x64 is chosen in the Versatile Video Coding Test Model.

[0024] Combining the use of separate luma / chroma coding trees with coding tools that involve interdependencies between chroma and its co-located luma samples (e.g., luma-dependent chroma residual scaling) can be problematic. Indeed, to process a chroma block, the luma samples of the luma block that is co-located with the considered chroma block need to be processed before the chroma block. If the luma block is of a large size, this can result in a large structural pipeline delay before the chroma block can be processed.

[0025] At least one embodiment applies constraints on the combined use of separate luma / chroma coding trees and coding tools that involve interdependencies between chroma and its co-located luma samples (e.g., chroma residual scaling, a component-to-component linear model known as CCLM) depending on the size of at least one luma coding block that is collocated with a chroma coding block.

[0026] At least one embodiment enables or disables a component interdependent coding tool (e.g., luma-dependent chroma residual scaling for a chroma-coded block) depending on the size (or horizontal / vertical dimension) of the chroma-coded block and depending on the size (or horizontal / vertical dimension) of the luma-coded block that co-locates with the sample of the chroma-coded block under consideration.

[0027] At least one embodiment provides a solution for chroma blocks when chroma residual scaling is disabled due to the above-mentioned constraints on the combined use of separate luma / chroma coding trees and chroma residual scaling.

[0028] Luma-dependent chroma residual scaling is an example of a coding tool that involves interdependence between chroma and its co-located luma samples. Luma-dependent chroma residual scaling is disclosed in JVET-M0427 by Lu et al., entitled "CE12: Mapping functions (test CE12-1 and CE12-2)." Luma-dependent chroma residual scaling involves using a scaling or inverse scaling table indexed by the luma value. This table is either explicitly signaled in the stream or inferred from a table coded in the stream.

[0029] On the encoder side, the process works as follows: When encoding a chroma block, a luma value representing the co-located luma block is calculated. This is typically the average of the luma samples in the luma prediction (or reconstructed) block that is co-located with the chroma block under consideration. From the calculated luma value, a scale value is obtained from a scale table. The scale value is applied as a multiplicative factor to the chroma prediction residual before applying a transform and then quantization to the chroma residual signal.

[0030] On the decoder side, the process works as follows: When decoding a chroma block, calculate a luma value representing the luma block that is co-located with the chroma block under consideration. This is typically the average of the luma samples in the luma prediction (or reconstructed) block that is co-located with the chroma block under consideration. From the calculated luma value, obtain an inverse scale value from an inverse scaling table that is signaled in the stream or inferred from data signaled in the stream. The inverse scale value is applied to the residual of the chroma prediction after applying inverse quantization and then an inverse transform.

[0031] JVET-M0427 proposes disabling chroma residual scaling when using separate luma / chroma coding trees, a solution that introduces coding loss for the chroma components.

[0032] In various embodiments, a separate luma / chroma tree splitting is enabled, i.e., the splitting of luma and chroma can be realized independently. The proposed embodiments improve coding efficiency, especially with respect to the chroma components. In the following, the chroma format is considered to be 4:2:0, i.e., the dimensions of the chroma components are half the dimensions of the luma components. However, it will be understood that embodiments of the present invention are not limited to this particular chroma format. Other chroma formats can be used by modifying the ratio between the luma dimensions and the chroma dimensions.

[0033] For simplicity and readability, the following figures show the same resolution for luma and chroma components even if the actual resolution is different, e.g., in the case of a 4:2:0 chroma format, in which case a simple scaling should be applied to the chroma picture.

[0034] When separate luma / chroma tree partitioning is enabled, maximum horizontal / vertical dimensions Wmax / Hmax are specified that constrain the use of luma-dependent chroma residual scaling for a given coded chroma block. The picture is divided into non-overlapping rectangular regions of size Wmax / Hmax for luma and size WmaxC / HmaxC for chroma (typically Wmax=Hmax=32 for luma, or WmaxC=HmaxC=16 for chroma), hereafter referred to as "authorized rectangular areas" (ARA). In several variant embodiments, activation of luma-dependent chroma residual scaling (hereafter referred to as LDCRS) is conditioned on the location and size of chroma blocks and co-located luma blocks relative to the ARA grid.

[0035] Various embodiments refer to luma blocks that are co-located with the chroma blocks under consideration. The luma blocks that are co-located with the chroma blocks can be defined as follows: -The luma block contains pixels that collocate with a given position in the chroma block, such that: The center of a chroma block, e.g., defined as its relative position within the chroma block ((x0+Wc) / 2, (y0+Hc) / 2), where (x0, y0) corresponds to the relative position within the chroma picture of the top-left sample of the chroma block, and (Wc, Hc) are the horizontal / vertical dimensions of the chroma block. The top left position within the chroma block, defined as a relative position (x0, y0) within the chroma picture, The bottom right position within the chroma block, defined as the relative position (x0+Wc-1, y0+Hc-1) within the chroma picture, The top right position within the chroma block, defined as the relative position (x0+Wc-1, y0) within the chroma picture, The bottom left position within the chroma block, defined as the relative position (x0, y0+Hc-1) within the chroma picture, Consider a luma block that is co-located with several given positions within a chroma block as described above, for example, four chroma block positions: top left, top right, bottom left, bottom right (see Figure 10), or Luma blocks that colocate with all chroma sample positions of the chroma block under consideration.

[0036] Embodiment 1 - LDCRS disabled when chroma blocks cross ARA boundaries In at least one embodiment, the LDCRS is disabled if the chroma block under consideration is not entirely contained within a single chroma ARA, which corresponds to the following equation: If ((x0 / WmaxC) !=((x0+Wc-1) / WmaxC))||((y0 / HmaxC) !=((y0+Hc-1) / HmaxC)) is true, disable LDCRS. where x||y is the Boolean "or" of x and y, and "!=" means "not equal to."

[0037] A simplified block diagram of this process is shown in Figure 7. Step 300 checks whether the chroma block is within a single chroma ARA. If this condition is true, then enable the LDCRS (step 303). If this condition is false, then disable the LDCRS (step 302).

[0038] Figures 8A, 8B, 8C, and 8D show some examples of chroma splitting, where gray blocks correspond to chroma blocks with LDCRS enabled and white blocks correspond to chroma blocks with LDCRS disabled. The outlines of the chroma blocks are shown by thick black lines. The grid defined by WmaxC and HmaxC is shown by dashed lines.

[0039] In FIG. 8A, two chroma blocks cross several ARAs and according to an embodiment of the present invention LDCRS is disabled for both chroma blocks. In FIG. 8B, two chroma blocks are in one chroma ARA and according to an embodiment of the present invention LDCRS is enabled for both chroma blocks. In Figure 8C, one rectangular vertical chroma block crosses two chroma ARAs, and according to an embodiment of the present invention, LDCRS is disabled for this chroma block. All other chroma blocks are within one chroma ARA, and therefore LDCRS can be enabled for those chroma blocks. In Figure 8D, one horizontal chroma block crosses two chroma ARAs, and according to an embodiment of the present invention, LDCRS is disabled for this chroma block. All other chroma blocks are within one chroma ARA, and therefore LDCRS can be enabled for those chroma blocks.

[0040] Embodiment 2 - LDCRS disabled when a chroma block crosses a chroma ARA boundary or when at least one of the co-located luma blocks crosses a co-located luma ARA boundary In one embodiment, the LDCRS is disabled if at least one of the following conditions is true: The chroma block under consideration is not completely contained within a single chroma ARA At least one of the luma blocks co-located with the chroma block under consideration is not completely contained within the luma ARA co-located with the chroma ARA, i.e., crosses the boundary of the co-located luma ARA.

[0041] A simplified block diagram of this process is shown in Figure 9. Step 400 checks whether the chroma blocks are within a single chroma ARA. If this condition is false, disable LDCRS (step 401). If this condition is true, identify the luma blocks that are co-located with the chroma block under consideration (step 402). Step 403 checks whether all co-located luma blocks are within the luma ARA that co-locates with the chroma ARA. If this condition is false, disable LDCRS (step 404). If this condition is true, enable LDCRS (step 405).

[0042] 11 shows two examples of embodiments of the present invention. At the top of the drawing, a chroma block is contained within a single chroma ARA (top left), with three co-located luma blocks (coloc luma blocks 1, 2, and 3, top right). Two of them (coloc luma blocks 1 and 2 in the dashed block) are within the co-located luma ARA, while the third (coloc luma block 3) is outside the co-located luma ARA. Because this third luma block is outside the co-located luma ARA, LDCRS is disabled according to embodiments of the present invention.

[0043] At the bottom of the figure, a chroma block is contained within a single chroma ARA, with three co-located luma blocks (coloc luma blocks 1, 2, and 3), three of which are within the co-located luma ARA. According to an embodiment of the present invention, LDCRS is enabled.

[0044] As shown in Figures 5 and 6, the current VVC specification supports three types of partitioning: quad-tree (QT), binary tree (BT, horizontal or vertical), and ternary tree (TT, horizontal or vertical). In a modified form, BT and TT partitioning are enabled from a given ARA dimension (e.g., 32x32 for luma and 16x16 for chroma when considering a 4:2:0 chroma format). Beyond this dimension, only QT partitioning or no partitioning is enabled. Thus, the conditions presented in embodiments 1 and 2 are systematically met, except for the "no partitioning" case for blocks larger than the luma ARA / chroma ARA dimension (e.g., 64x64 is not partitioned into four 32x32 blocks for luma blocks).

[0045] In at least one embodiment, BT and TT partitioning is only enabled for luma block sizes of 32x32 or smaller, or for chroma block sizes of 16x16 or smaller (when using 4:2:0 chroma format). Above this size, only QT partitioning (or no partitioning) is enabled. The restriction is for mode LDCRS or mode CCLM. This process is illustrated in FIG. 11C, which considers a luma ARA of size 32x32 and a chroma ARA of size 16x16. Step 1300 checks whether the block being considered is a luma block of size 64x64 or larger, or a chroma block of size 32x32 or larger. If this test is true, step 1301 disables BT and TT partitioning while enabling no partitioning and QT partitioning. If this test is false, step 1302 enables no partitioning, QT partitioning, BT partitioning, and TT partitioning.

[0046] In at least one embodiment, to reduce the delay of processing chroma in the case of modes LDCRS or CCLM, if a square luma block larger than the luma ARA is not divided, the corresponding chroma block can be divided by QT division only into four blocks that can possibly be further divided or not divided at all. For example, consider a luma ARA of size 32x32 and a chroma ARA of size 16x16, and if a square luma block of size 64x64 is not divided, the corresponding chroma block (of size 32x32 in 4:2:0 format) can be divided by QT division only into four 16x16 blocks that can possibly be further divided (dotted lines in the figure) or not divided at all. An example that follows these constraints is shown in Figure 11D.

[0047] In at least one embodiment, to reduce the delay of processing chroma in the case of modes LDCRS or CCLM, if a square luma block larger than the luma ARA is divided into four blocks by QT division, the corresponding chroma block can be divided by QT division only into four blocks that can possibly be further divided or not divided at all. For example, considering a luma ARA of size 32x32 and a chroma ARA of size 16x16, if a square luma block of size 64x64 is divided into four 32x32 blocks, the corresponding chroma block (of size 32x32 in 4:2:0 format) can be divided by QT division only into four 16x16 blocks that can possibly be further divided (dotted lines in the figure) or not divided at all. An example that follows these constraints is shown in Figure 11E.

[0048] In at least one embodiment, to reduce the delay of processing chroma in the case of modes LDCRS or CCLM, if a square luma block larger than the luma ARA is divided into two blocks by horizontal BT division, the corresponding chroma block can only be divided into two blocks by horizontal BT division or into four blocks by QT division. For example, considering a luma ARA of size 32x32 and a chroma ARA of size 16x16, if a square luma block of size 64x64 is divided into two blocks of 32 lines and 64 columns, the corresponding chroma block (of size 32x32 in 4:2:0 format) can only be divided into two blocks of 16 lines and 32 columns by horizontal BT division or into four 16x16 blocks by QT division. An example that follows these constraints is shown in Figure 11F.

[0049] In at least one embodiment, to reduce the delay in processing chroma in the case of modes LDCRS or CCLM, if a square luma block larger than luma ARA is not split, or is split into four blocks by QT split, or is split into two blocks by horizontal QT split, the corresponding chroma block cannot be split into two blocks by vertical binary split, but can be split into four blocks by QT split, or can be split into two blocks by horizontal BT split, or may not be split.

[0050] In at least one embodiment, to reduce the delay of processing chroma in the case of mode LDCRS or CCLM, if a square luma block larger than the luma ARA is divided into two blocks by vertical BT partitioning, the corresponding chroma block can only be divided into two blocks by vertical BT partitioning. For example, consider a luma ARA of size 32x32 and a chroma ARA of size 16x16, and if a square luma block of size 64x64 is divided into two blocks of 64 lines and 32 columns, the corresponding chroma block can only be divided into two blocks of 32 lines and 16 columns. An example that follows these constraints is shown in Figure 11G.

[0051] Example syntax based on VTM5.0 syntax described in document JVET-N1001 (version 9 - date 2019-06-25 13:45:21) The text in small font below corresponds to an example of syntax corresponding to an example implementation of the above embodiment, based on the syntax described in document JVET-N1001 version 9. The section numbering corresponds to the numbering used in JVET-N1001 version 9.

[0052] Version in which BT / TT splitting is prohibited for blocks larger than VDPU The syntax description below corresponds to the embodiment that restricts splitting as described above, with changes to the VTM5 v9 specification highlighted in grey.

[0053] 6.4.2 Allowed Binary Split Processes The inputs to this process are: -Binary split mode btSplit, - cbWidth, the width of the coded block in luma samples, - cbHeight, the height of the coded block in luma samples, the position (x0, y0) of the top left luma sample of the considered coded block relative to the top left luma sample of the picture, - multitype tree depth mttDepth, - maximum multitype tree depth maxMttDepth with offset, - Maximum binary tree size maxBtSize, - partition index partIdx, - Variable treeType that specifies whether a single tree (SINGLE_TREE) or a dual tree is used to split the CTU, and if a dual tree is used, whether the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) is currently being processed. The output of this process is the variable allowBtSplit.

[0054] [Table 1]

[0055] The variables parallelTtSplit and cbSize are derived as specified in Table 6-2.

[0056] The variable allowBtSplit is derived as follows: - Set allowBtSplit equal to false if one or more of the following conditions are true: -cbSize is less than or equal to MinBtSizeY -cbWidth is greater than maxBtSize -cbHeight is greater than maxBtSize -treeType is not equal to SINGLE_TREE and cbWidth is greater than 32 -treeType is not equal to SINGLE_TREE and cbHeight is greater than 32 -mttDepth is greater than or equal to maxMttDepth -treeType is equal to DUAL_TREE_CHROMA and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 16. [Ed.(SL): Is "less than or equal to" necessary here?] Otherwise, set allowBtSplit equal to false if all of the following conditions are true: ... Otherwise, set allowBtSplit equal to true.

[0057] 6.4.3 Allowed Three-Way Split Processes The inputs to this process are: - ternary split mode ttSplit, - cbWidth, the width of the coded block in luma samples, - cbHeight, the height of the coded block in luma samples, the position (x0, y0) of the top left luma sample of the considered coded block relative to the top left luma sample of the picture, - multitype tree depth mttDepth, - maximum multitype tree depth maxMttDepth with offset, - maximum ternary tree size maxTtSize, - Variable treeType that specifies whether a single tree (SINGLE_TREE) or a dual tree is used to split the CTU, and if a dual tree is used, whether the luma (DUAL_TREE_LUMA) or chroma component (DUAL_TREE_CHROMA) is currently being processed. The output of this process is the variable allowTtSplit.

[0058] [Table 2]

[0059] The variable cbSize is derived as specified in Table 6-3.

[0060] The variable allowTtSplit is derived as follows: - Set allowTtSplit equal to false if one or more of the following conditions are true: -cbSize is less than or equal to 2*MinTtSizeY -cbWidth is greater than Min(MaxTbSizeY, maxTtSize) -cbHeight is greater than Min(MaxTbSizeY, maxTtSize) -treeType is not equal to SINGLE_TREE and cbWidth is greater than 32 -treeType is not equal to SINGLE_TREE and cbHeight is greater than 32 -mttDepth is greater than or equal to maxMttDepth -x0+cbWidth is greater than pic_width_in_luma_samples -y0+cbHeight is greater than pic_height_in_luma_samples -treeType is equal to DUAL_TREE_CHROMA and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 32 Otherwise, set allowTtSplit equal to true.

[0061] Versions where CCLM is prohibited when VDPU is divided into BT or TT The following syntax description corresponds to an embodiment where CCLM (or CRS) is not enabled, and partitions do not respect the above splitting restrictions. For example, CCLM is disabled if a chroma block is of size 32x16, or if its co-located luma block is of size 64x32 or 32x64. CCLM (or CRS) is enabled in other cases. Changes to the VTM5 v9 specification are highlighted in gray.

[0062] 1.1.1 Chrominance Intra Prediction Mode Derivation Process The inputs to this process are: - a luma position (xCb, yCb) that specifies the top-left sample of the current chroma-coded block relative to the top-left luma sample of the current picture; - the variable cbWidth, which specifies the width of the current coded block in luma samples, - The variable cbHeight that specifies the height of the current coded block in luma samples. - variable treeType that specifies whether a single tree (SINGLE_TREE) or a dual tree is used to split the CTU.

[0063] In this process, the chrominance intra-prediction mode IntraPredModeC[xCb][yCb] is derived.

[0064] The corresponding luma intra prediction mode lumaIntraPredMode is derived as follows: - If intra_mip_flag[xCb][yCb] is equal to 1, lumaIntraPredMode is derived as specified in Table 8-4 using IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2] and sizeId set equal to MipSizeId[xCb][yCb]. Otherwise, lumaIntraPredMode is set equal to IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2]. The variable cclmEnabled is derived by invoking the inter-component chroma intra prediction mode checking process of the subordinate clause xxx [Ed.(SL): subordinate clause number is undetermined].

[0065] The chrominance intra prediction mode IntraPredModeC[xCb][yCb] is derived using intra_chroma_pred_mode[xCb][yCb] and lumaIntraPredMode as specified in Tables 8-5 and 8-6.

[0066] [Table 3]

[0067] [Table 4]

[0068] If chroma_format_idc is equal to 2, chroma intra prediction mode Y is derived using chroma intra prediction mode X in Tables 8-5 and 8-6 as specified in Table 8-7, and chroma intra prediction mode X is later set equal to chroma intra prediction mode Y.

[0069] [Table 5]

[0070] xxx Inter-component Chroma Intra Prediction Mode Inspection Process The inputs to this process are: - Luma position (xCb, yCb) that specifies the top-left sample of the current chroma-coded block relative to the top-left luma sample of the current picture - variable cbWidth specifying the width of the current coded block in luma samples - variable cbHeight specifying the height of the current coding block in luma samples The output to this process is: - A flag lmEnabled that specifies whether the inter-component chroma intra prediction mode is enabled for the current chroma coded block. In this process, cclmEnabled is derived as follows: - Set wColoc and hColoc equal to the width and height of the collocated luma coded block covering the position given by (xCB<<1, yCb<<1)). Set -cclmEnabled equal to 1. -If sps_cclm_enabled_flag is equal to 0, set cclmEnabled equal to 0. - Otherwise, if treeType is equal to SINGLE_TREE, set cclmEnabled equal to 1. - Otherwise, set cclmEnabled equal to 0 if one of the following conditions is false: -cbWidth or cbHeight is greater than or equal to 32 and cbWidth is not equal to cbHeight. -wColoc or hColoc is greater than or equal to 64 and wColoc is not equal to hColoc.

[0071] Embodiment 2a - Variation using inferred enabling flag for Luma ARA; generalization to CCLM The above concepts can also be applied to other luma-to-chroma modes such as CCLM.

[0072] In a variation of embodiment 2, the CCLM (or LDCRS) is disabled if at least one of the following conditions is true: The chroma block being considered is not completely contained within a single chroma ARA. The luma ARA that is co-located with the chroma ARA contains at least one luma block that is not completely contained within the luma ARA, i.e., one luma block that crosses the boundary of the co-located luma ARA.

[0073] This can be implemented by using an inference flag specified for each luma ARA, hereinafter referred to as luma_blocks_inside_flag. For each luma ARA, a flag is inferred from the luma partition tree (selected in the encoder or parsed in the decoder). Once the luma partition tree is generated for a CTU (or VDPU), a flag for each luma ARA of the CTU (or VDPU) is inferred as follows: If all luma blocks of a given luma ARA are strictly contained, the luma ARA, the flag is set to true. Otherwise, if at least one luma block crosses one or more of the boundaries of the luma ARA, the flag is set to false.

[0074] The flags can then be used as follows: Condition (as above): The luma ARA that is co-located with the chroma ARA contains at least one luma block that is not completely contained within the luma ARA, i.e., one luma block that crosses the boundary of the co-located luma ARA. can be equivalently formulated as follows: The flag luma_blocks_inside_flag of the luma ARA that is colocated with the chroma ARA is false.

[0075] Embodiment 2b - LDCRS or CCLM enabled when a chroma block is within its own co-located luma block In another variation, if a chroma block is not entirely within its co-located luma block, the luma-dependent mode (LDCRS or CCLM) is disabled.

[0076] This can be expressed as follows: LDCRS or CCLM is disabled if one of the following conditions is false: The luma block that colocates with the top-left sample of the chroma block (at position (x0, y0) and of size (Wc, Hc)) is within the 32x32 chroma ARA xColocY<=2*x0 ·xColocY+wColocY-1>=2*(x0+Wc-1) yColocY<=2*y0 ·yColocY+hColocY-1>=2*(y0+Hc-1) (xColocY, yColocY) is the position of the top left sample of the luma block that is colocated with the top left sample of the chroma block, and (wColocY, hColocY) is the dimension of this luma block.

[0077] Embodiment 2c - Use of fallback mode when LDCRS or CCLM is disabled At least one embodiment relates to CCLM mode (and its variant, MDLM). This embodiment considers there to be a traditional CCLM mode and a fallback CCLM mode. The fallback CCLM can be used when the traditional CCLM mode is not enabled due to the constraints described herein.

[0078] The block diagram in Figure 11A shows a simplified block diagram of the activation of the CCLM process. In step 112, conditions for activating conventional CCLM mode for a given chroma block are checked (such conditions are described in various embodiments herein). If all enabling conditions are met for a given chroma block, conventional CCLM mode may be enabled and used by the encoder / decoder (113). If one of the enabling conditions is not met for a given chroma block, conventional CCLM mode is disabled and a fallback CCLM mode may be enabled and used by the encoder / decoder (114).

[0079] In at least one embodiment, the conventional CCLM mode corresponds to a mode in which CCLM parameters are derived from neighboring luma and chroma samples of the chroma block under consideration (e.g., chroma samples and their co-located luma samples belonging to one or more lines on top of the chroma block, and / or chroma samples and their co-located luma samples belonging to one or two columns to the left of the chroma block).

[0080] Possible fallback CCLM mode In at least one embodiment, the fallback CCLM mode uses CCLM parameters derived from chroma and luma samples adjacent to the current CTU, for example, one or two lines at the top of the CTU and / or one or two columns to the left of the CTU.

[0081] In at least one embodiment, the fallback CCLM mode uses CCLM parameters derived from chroma and luma samples adjacent to the current VDPU, for example, one or two lines on top of the VDPU and / or one or two columns to the left of the VDPU.

[0082] In at least one embodiment, the fallback CCLM mode uses CCLM parameters derived from chroma and luma samples adjacent to the current 16x16 chroma ARA and its corresponding 32x32 luma ARA, for example, the top one or two lines of the current 16x16 chroma ARA and its corresponding 32x32 luma ARA and / or the left one or two columns of the current 16x16 chroma ARA and its corresponding 32x32 luma ARA are used.

[0083] In at least one embodiment, the fallback CCLM mode uses CCLM parameters derived from luma samples adjacent to the luma block that are co-located with the top-left sample of the chroma block, and from chroma samples that are co-located with those luma samples.

[0084] In at least one embodiment, the fallback CCLM mode uses the last used CCLM parameters.

[0085] Using the CCLM parameter history list In at least one embodiment, the fallback CCLM mode uses CCLM parameters from a history list of CCLM parameters built in the encoder and decoder. A maximum size Nh is specified for the history list. The history list is populated with the last used CCLM parameters. If a chroma block uses a CCLM mode, its CCLM parameters are added to the history list. If the history list is already full, the oldest CCLM parameters are removed and new ones are inserted into the history list. In one variation, new CCLM parameters are added to the history list only if they are different from CCLM parameters already in the history list. In one variant, if a chroma block is coded using the fallback CCLM mode, the index of a CCLM parameter from the history list of CCLM parameters is coded for the chroma block. In one variant, the index of the CCLM parameter from the history list is inferred based on a similarity test of the luma and / or chroma samples within or surrounding the chroma block, which luma and / or chroma samples have been used to calculate the CCLM parameters of the list.

[0086] One possible implementation of the latter variant is as follows: When a CCLM parameter is inserted into the history list, the average value of the luma samples used to derive the parameter is calculated and stored in the list, denoted as avg_Ref_Y[idx], where idx is the index in the history list of the CCLM parameter (idx=0 to Nh-1, where Nh is the size of the history list).

[0087] If CCLM mode is used for a chroma block, calculate the average value of the luma samples (considering only available luma samples) adjacent to the chroma block, denoted as avg_Cur_Y.

[0088] Identify the CCLM parameter index idx0 in the list, used to perform CCLM prediction for the chroma block, as the index that minimizes the absolute value of the difference between avg_Ref_Y[i] and avg_Cur_Y: idx0=argmin(|avg_Cur_Y-avg_Ref_Y[idx]|)

[0089] The similarity metric can alternatively be based on chroma samples. When a CCLM parameter is inserted into the history list, the average values of the chroma samples used to derive the parameter are calculated and stored in the list, denoted as avg_Ref_Cb[idx], avg_Ref_Cr[idx].

[0090] If CCLM mode is used for a chroma block, calculate the average values of the chroma samples adjacent to the chroma block, denoted as avg_Cur_Cb and avg_Cur_Cr.

[0091] Identify the CCLM parameter index in the list, idx0, used to perform CCLM prediction for the chroma block, as the index that minimizes the absolute value of the difference between avg_Ref_Cb[idx], avg_Ref_Cr[idx] and avg_Cur_Cb, avg_Cur_Cr: idx0=argmin(|avg_Cur_Cb-avg_Ref_Cb[idx]|+|avg_Cur_Cr-avg_Ref_Cr[idx]|)

[0092] FIG. 11B shows a simplified block diagram of an example of an index selection process from a history list. The inputs to this process are luma and / or chroma samples neighboring the chroma block under consideration and the history list. The history list contains, for each index i, CCLM parameters and reference values related to the luma or chroma samples used to derive the CCLM parameters (e.g., avg_Ref_Y[i], avg_Ref_Cb[i], avg_Ref_Cr[i]). Step 1200 calculates local values from neighboring luma or chroma samples neighboring the chroma block (e.g., avg_Cur_Y, avg_Cur_Cb, avg_Cur_Cr). Step 1201 sets a parameter distMin to a large value. Step 1202 loops over indexes = 0 to Nh-1. Step 1203 calculates the distortion "dist" between the local value calculated in step 1200 and the reference value for index i. Step 1204 checks whether dist is less than minDist. If this is true, step 1205 updates the values of idx0 and minDist and the process continues to step 1206. If not (dist is not less than minDist), the process continues to step 1206. Step 1206 checks to see if the end of the loop over the history list indices is complete. If so, idx0 is the output index value. If not, the process returns to step 1202 to check the next index.

[0093] Embodiment 3 - LDCRS enabled when no chroma blocks cross ARA boundaries and when at least one of the co-located luma blocks does not cross a co-located luma ARA boundary In one embodiment, the LDCRS is enabled if all of the following conditions are true: The chroma block under consideration is completely contained within a single chroma ARA At least one of the luma blocks co-located with the chroma block under consideration is completely contained within the luma ARA co-located with the chroma ARA. Hereinafter, such a luma block is referred to as a "co-located enabled luma block."

[0094] A simplified block diagram of this process is shown in Figure 12. Step 500 checks whether the chroma blocks are within a single chroma ARA. If this condition is false, disable LDCRS (step 501). If this condition is true, identify luma blocks that are co-located with the chroma block under consideration (step 502). Step 503 checks whether at least one of the co-located luma blocks is within a luma ARA that is co-located with the chroma ARA. If this condition is false, disable LDCRS (step 504). If this condition is true, enable LDCRS (step 505).

[0095] Referring to Figure 13 (showing the same division as the top diagram of Figure 11), LDCRS is enabled because the chroma block is within a single chroma ARA and the co-located luma blocks 1 and 2 are within the co-located luma ARA, i.e., at least one of the luma blocks co-located with the chroma block under consideration is completely contained within the luma ARA co-located with the chroma ARA.

[0096] Referring to Figure 14, since a chroma block covers two chroma ARAs (first chroma ARA and second chroma ARA), LDCRS is disabled even if two of its co-located luma blocks (coloc luma blocks 1 and 2) are within the luma ARA that co-locates with the first chroma ARA.

[0097] Embodiment 4 - LDCRS enabled if at least one of the co-located luma blocks does not cross the co-located luma ARA boundary In one embodiment, even if the chroma block under consideration is not entirely contained within a single chroma ARA, LDCRS can be enabled if at least one of the luma blocks co-located with the chroma block under consideration is entirely contained within a luma ARA co-located with the first chroma ARA of the chroma block under consideration. In this embodiment, the chroma block under consideration may be larger than a single chroma ARA.

[0098] A simplified block diagram of this process is shown in Figure 15. Luma blocks that are co-located with the chroma block under consideration are identified (step 600). Step 601 checks whether at least one of the co-located luma blocks is in the luma ARA that is co-located with the first chroma ARA. If this condition is false, the LDCRS is disabled (step 602). If this condition is true, the LDCRS is enabled (step 603).

[0099] Figure 16 shows the same division as Figure 14. Even though a chroma block covers two chroma ARAs (first and second chroma ARAs), LDCRS is enabled because part of its co-located luma block (blocks 1 and 2) is within the luma ARA that co-locates with the first chroma ARA of the chroma block.

[0100] In a variation of this embodiment, when LDCRS is enabled and activated for a chroma block, the scale factors used to scale the chroma residual are derived only from luma samples that belong to the co-located enabled luma block: all chroma samples use the same scale factors that are derived from only a portion of the co-located luma samples.

[0101] For example, consider the case in Figure 16 and calculate the average of the luma samples from coloc luma blocks 1 and 2. This average luma value is used to identify the scale factor that is applied to the residual for the entire chroma block, so the luma samples from luma block 3 are not used.

[0102] Embodiment 5: Signaling scale factors for chroma blocks with LDCRS disabled In one embodiment, if a chroma block does not benefit from luma-based chroma residual scaling, the chroma scale factor is signaled in the bitstream. If a chroma block is made up of several transform units, the chroma scale factor is applied to the residual of each TU.

[0103] A simplified block diagram of this process is shown in Figure 17. Step 700 checks whether LDCRS is enabled for the chroma block under consideration. If this condition is false, a scale factor or scale factor index is coded / decoded (step 701). If this condition is true, a scale factor or scale factor index is derived from the co-located luma block (step 702).

[0104] In one embodiment, the chroma scale factor index is signaled. In one embodiment, the scale factor or scale factor index is predicted from the scale factor or scale factor index of one or several adjacent chroma blocks previously coded / decoded. For example, as shown in Figure 18, the scale factors or scale factor indexes from chroma blocks including block A, block B if block A is unavailable, block C if blocks A and B are unavailable, and block D if blocks A, B, and C are unavailable are used as predictors. Otherwise, the scale factor or scale factor index is not predicted.

[0105] In a variant, the scale factors are inferred by a prediction process using neighboring chroma blocks that have already been processed: no additional signaling is used, and only the inferred scale values are used for the chroma block under consideration.

[0106] Embodiment 6: Delta QP signaling for chroma blocks with LDCRS disabled In one embodiment, if a chroma block does not benefit from luma-based chroma residual scaling, the delta QP parameters signaled in the bitstream may be used for that chroma block. If the chroma block is made up of several transform units, the delta QP parameters are applied to the residual of each TU.

[0107] In one embodiment, delta QP parameters are signaled for chroma blocks.

[0108] In one embodiment, a delta QP parameter is signaled for a quantization group that includes a chroma block.

[0109] In one embodiment, delta QP parameters are signaled for CTUs that contain chroma blocks.

[0110] In one embodiment, chroma QP parameters are predicted based on QP values or scale factors from one or several previously coded / decoded neighboring chroma blocks. For example, as shown in Figure 18, QP values or scale factors from chroma blocks including block A, block B if block A is unavailable, block C if blocks A and B are unavailable, and block D if blocks A, B, and C are unavailable are used as predictors.

[0111] If the predictor is a scale factor sc, then this scale factor used to predict chroma QP is given by the equation: QPpred=-6*Log2(Sc) It is first converted to a value like QP denoted by QPpred using Log2, where Log2 is the base 2 logarithm function.

[0112] If the predictor is a QP value, then this QP value is used as the predictor QPpred.

[0113] In a variant, the QP value is inferred by a prediction process using neighboring chroma blocks that have already been processed: no additional signaling is used, and only the inferred QP value is used for the chroma block under consideration.

[0114] Embodiment 7: Extension to other luma-dependent chroma coding modes The disclosed embodiments with respect to luma-based chroma residual scaling can be used with other luma-dependent chroma coding modes, or more broadly with coding tools that involve inter-component dependencies, such as chroma-from-luma intra prediction (also known as CCLM or LM mode, and its variant Multiple Direction Linear Model, also known as MDLM).

[0115] In a further embodiment, for luma blocks with dimensions WL and HL larger than WmaxL and HmaxL, CCLM can be enabled only if the luma block is divided into transform units of maximum dimensions WmaxL and HmaxL.

[0116] Similarly, for luma blocks with dimensions WL and HL larger than WmaxL and HmaxL, chroma residual scaling can only be enabled if the luma block is divided into transform units of maximum dimensions WmaxL and HmaxL.

[0117] In a further embodiment, for luma blocks with dimensions WL and HL larger than WmaxL and HmaxL, the luma blocks are systematically partitioned into transform units of maximum size WmaxL and HmaxL. In this way, chroma residual scaling and / or CCLM modes can be used within the chroma components without suffering from the structural delay implied by the need for reconstructed collocated luma blocks to be available for processing the chroma blocks.

[0118] In a further modification, for chroma blocks with dimensions Wc and Hc larger than WmaxC and HmaxC, the chroma blocks are systematically divided into transform units of maximum size WmaxC and HmaxC. In this way, chroma residual scaling and / or CCLM modes can be used within the chroma components whatever the size of the chroma coding units.

[0119] The example in FIG. 19 shows the division of a luma block into two TUs (left side) or four TUs (right side).

[0120] In a variant, therefore, if a chroma block has dimensions Wc and Hc larger than WmaxC and HmaxC, CCLM can be enabled only if the chroma block is divided into transform units of maximum dimensions WmaxC and HmaxC.

[0121] Similarly, if a chroma block has dimensions Wc and Hc that are larger than WmaxC and HmaxC, chroma residual scaling can only be enabled if the chroma block is divided into transform units of maximum dimensions WmaxC and HmaxC.

[0122] Embodiment 8: Deactivation of separate luma / chroma trees for CTUs using availability of reference samples for CCLM The CCLM mode uses reference luma and chroma samples to determine parameters that are further used to predict a chroma block sample from its co-located luma sample. The reference luma and chroma samples are typically located at: Top line adjacent to chroma / colocated luma block, The left column adjacent to the chroma / colocated luma block, Possibly a top-left position adjacent to the chroma / co-located luma block. A combination of these three positions can be used depending on the CCLM mode considered.

[0123] In one embodiment, chroma reference samples used for CCLM mode are considered available if they belong to a chroma block contained in the neighboring chroma ARA above / left / top right of the chroma ARA containing the chroma block under consideration (FIG. 20). If the chroma reference samples are not contained in the neighboring chroma ARA above / left / top right, they are considered unavailable.

[0124] 21 shows a case where the top reference chroma samples (within the diagonally shaded rectangle above the chroma block) used for a chroma block belong to different luma blocks (blocks 1, 3, and 4), and some of those luma blocks (blocks 3 and 4) are not in the adjacent top chroma ARA. According to this embodiment, only the reference sample belonging to block 1 is considered available. Other reference samples from blocks 3 and 4 cannot be used for CCLM mode.

[0125] In a variant, CCLM mode is disabled if a given proportion (eg, 30%) of the chroma reference samples used for CCLM prediction are not contained within the adjacent top / left / top right chroma ARA.

[0126] In Figure 21, the number of available reference samples is equal to 25% of the total number of reference samples, therefore CCLM is disabled.

[0127] In a variant, CCLM mode is disabled as soon as at least one chroma reference sample used for CCLM prediction is not contained within the adjacent top / left / top right chroma ARA.

[0128] Referring to FIG. 21, in this illustration CCLM is disabled because at least one reference sample is unavailable.

[0129] Similarly, for luma reference samples used in CCLM, the following embodiment is proposed: Luma reference samples used in CCLM mode are considered available if they belong to a luma block contained in a luma ARA that is co-located with the neighboring chroma ARA above / left / top right of the chroma ARA containing the chroma block under consideration (FIG. 20). Luma reference samples are considered unavailable if they are not contained in the neighboring luma ARA above / left / top right that is co-located with the neighboring chroma ARA above / left / top right.

[0130] In a modified form, CCLM mode is disabled if a given percentage (e.g., 30%) of luma reference samples are not contained within the adjacent top / left / top right luma ARA that collocates with the top / left / top right adjacent chroma ARA.

[0131] In a modified form, CCLM mode is disabled if at least one luma reference sample is not contained within an adjacent top / left / top right luma ARA that is co-located with an adjacent top / left / top right chroma ARA.

[0132] Embodiment 9: Deactivation of Separate Luma / Chroma Trees for CTUs Using CCLM In one embodiment, if at least one chroma block in a CTU uses CCLM (or LDCRS), the separate luma / chroma tree is disabled and the chroma split for the CTU is inferred from the luma split.

[0133] In a variant, a flag is signaled at the CTU level to indicate whether CCLM (or LDCRS) is used within the CTU.

[0134] This concept can be generalized to different block types than CTUs (e.g., video decoding processing units, also known as VDPUs).

[0135] Embodiment 10: Chroma residual scaling based on neighboring samples of luma block co-located with the top-left corner of the chroma block In another embodiment, potential hardware latency issues are further reduced by determining the scale factor used for chroma residual scaling of a chroma block from already processed predicted or reconstructed luma samples that are neighbors of the luma block that colocates with a given position in the chroma block, hereafter referred to as "neighboring luma samples."

[0136] FIG. 22 shows a flow diagram of a method for determining the scale factor used for chroma residual scaling of a chroma block.

[0137] In step 800, a location in a picture of a given sample location in a chroma block is identified. In step 801, a luma block containing a luma sample that colocates with the location in the chroma block is identified. In step 802, already-processed luma samples that neighbor the luma block are identified. In step 803, a luma value is determined from the identified neighboring luma samples, e.g., an average value of the identified neighboring luma samples. In step 804, a scale factor for the chroma block is determined based on the luma value. In step 805, the residual of the chroma block is scaled using the scale factor.

[0138] In one embodiment, a given sample location within a chroma block is the top-left corner of the chroma block. This is shown in Figure 23A. The chroma block (rectangle) is shown in bold. The co-located luma block (square) relative to the top-left sample of the chroma block is in thin lines. Its neighboring samples are shown in gray. In a variant, a given sample location within a chroma block is the center of the chroma block.

[0139] In one embodiment, only one neighboring luma sample is used, corresponding to the sample above or to the left of the top-left sample of the co-located luma block, as shown in Figure 23B. A variant uses both the above sample and the left sample.

[0140] In another embodiment, adjacent luma samples are made up of adjacent top lines of size WL and adjacent left column lines of size HL, where WL and HL are the horizontal and vertical dimensions of the luma block, as shown in Figure 23C.

[0141] In another embodiment, the adjacent luma samples are made up of adjacent top lines of size minS and adjacent left column lines of size minS, where minS is the minimum value among WL and HL.

[0142] In another embodiment, assume that the adjacent luma samples are made up of adjacent top lines of size Wc*2 and adjacent left column lines of size Hc*2, where Wc and Hc are the horizontal and vertical dimensions of the chroma block, and the chroma format is 4:2:0.

[0143] In another embodiment, the adjacent luma samples are made up of adjacent top lines of size minSC*2 and adjacent left column lines of size minSC*2, where minSC is the minimum value among Wc and Hc.

[0144] In a variant, the adjacent top line of size minSC*2 starts at the same relative horizontal position as the top left corner of the chroma, and the adjacent left column of lines of size minSC*2 starts at the same relative vertical position as the top left corner of the chroma, as shown in Figure 23D. For example, if the top left corner in a chroma block is at position (xc, yc) in the chroma picture and the chroma format is 4:2:0, then the first sample in the top line of adjacent luma samples is at horizontal position 2*xc, and the first sample in the left column of adjacent luma samples is at vertical position 2*yc.

[0145] According to a further modification, the neighboring samples used to determine the scale factor to be applied to the current chroma block of the residual are made up of one or more reconstructed luma samples belonging to the line below and / or the column to the right of the luma samples belonging to the neighboring ARAs above and to the left of the ARA containing the current chroma block, respectively.

[0146] In another embodiment, no scaling of the residual of a chroma block is applied if a given ratio of neighboring luma samples is not available.

[0147] In another embodiment, a neighboring luma sample is considered unavailable if it is not in the same CTU as the chroma block.

[0148] In another embodiment, a neighboring luma sample is considered unavailable if it is not in the same VDPU as the chroma block.

[0149] In another embodiment, a neighboring luma sample is considered unavailable if it is not in the same luma ARA that colocates with the first chroma block ARA.

[0150] In another embodiment, residual scaling of a chroma block is not applied if a given proportion of chroma samples in the considered neighborhood is not available, which may occur if a neighboring chroma block is not yet available in its reconstructed state at the time the current chroma block is being processed.

[0151] In one embodiment, the neighboring luma samples used to derive the chroma residual scale factors are the samples used to predict the co-located luma block (often known as intra-prediction reference samples). In the binary tree case, for each given region (e.g., VDPU), it is common to process the luma blocks of the given region first, followed by the chroma blocks of the given region. When an embodiment of the present invention is applied, it means that for each luma block of the given region, its neighboring luma samples need to be stored, which requires additional memory storage and complicates the process because the number of reference samples varies depending on the size of the block. To reduce these negative effects, in one embodiment, only the luma values derived in step 803 from the neighboring luma samples are stored for each luma block. In another embodiment, only the scale factors derived in step 804 from the luma values derived in step 803 from the neighboring luma samples are stored for each luma block. This limits storage to a single value per luma block. This principle can be extended to modes other than LMCS, such as CCLM, where instead of storing the luma and chroma samples for each block, only the minimum and maximum values of the luma and chroma samples used to derive the CCLM parameters for the block are stored.

[0152] The concept of embodiment 10 can be generalized to CCLM mode. In this case, this embodiment consists of determining CCLM linear parameters used to predict a chroma block from already processed predicted or reconstructed luma samples (neighboring luma samples) and predicted or reconstructed chroma samples (neighboring chroma samples) in the neighborhood of the luma block co-located with a given position in the chroma block. Figures 23A to 23D illustrate this concept. The gray-filled rectangular area corresponds to the neighboring samples (luma and chroma) of the luma block co-located with the top-left sample of the current chroma block that may be used to derive the CCLM parameters.

[0153] According to a variant, the luma and / or chroma neighboring samples used to determine the linear model used to perform CCLM prediction of the current chroma block are made of at least one reconstructed or predicted luma and / or chroma sample belonging to the bottom line and / or right column of luma and / or chroma samples belonging to the upper and / or left neighboring luma and / or chroma ARAs, respectively, of the ARA containing the current chroma block. This concept can be extended to the bottom line of the upper right neighboring luma and / or chroma ARA of the ARA containing the current chroma block and to the right column of the lower left neighboring luma and / or chroma ARA of the ARA containing the current chroma block. This concept can be extended to the bottom line of the upper right neighboring luma and / or chroma ARA of the ARA containing the current chroma block and to the right column of the lower left neighboring luma and / or chroma ARA of the ARA containing the current chroma block.

[0154] Embodiment 11: Checking availability of luma samples based on position relative to chroma blocks In another embodiment, a neighboring luma sample at position (xL, yL) is considered unavailable if the distance between the neighboring luma sample at position (xC, yC) and the position of the top-left sample of the considered chroma block is greater than a given value. In other words, a neighboring luma sample is considered unavailable if the following condition is true: ((xC*2)-xL)>TH, or ((yC*2)-yL)>TV or equivalently (xC-(xL / 2))>TH / 2, or (yC-(yL / 2))>TV / 2 where TH and TV are predefined values or values signaled in the bitstream. Typically, TH=TC=16. In a variant, TH and TC depend on the picture resolution. For example, TH and TC are determined according to the following conditions: - If the picture resolution is 832x480 luma samples or less, TH=TC=8, - Otherwise, if the picture resolution is 1920x1080 luma samples or less, then TH=TC=16, -Otherwise TH=TC=32.

[0155] Figures 24A and 24B show examples of distances between the top-left neighboring luma sample and the neighboring chroma sample (scaled by 2 for 4:2:0 chroma format). Figure 24A shows the distance between the top-left chroma sample and the top-left neighboring chroma sample of a co-located luma block. Figure 24B shows the distance between the top-left chroma sample and the top line or left column.

[0156] In a modification, instead of strictly invalidating samples that are too far from the top-left chroma sample, a weighting depending on the distance between the neighboring chroma samples and the top-left chroma sample is applied to the reference samples considered.

[0157] In a modification, instead of strictly invalidating samples that are too far from the top-left chroma sample, a weighting depending on the distance between the neighboring chroma samples and the top-left chroma sample is applied to the reference samples considered.

[0158] Embodiment 12: Checking the availability of luma samples based on the values of their co-located chroma samples in relation to the values of the chroma samples of the current block In another embodiment, the availability of a neighboring luma sample neighborY is based on the values of its collocated Cb and Cr chroma samples neighborCb, neighborCr, and the values of the predicted Cb and Cr chroma samples of the current chroma block. For example, a neighboring luma sample neighbor is considered unavailable if the following condition is true: -Abs(topLeftCb-neighborCb)>Th_Ch, or Abs(topLeftCr-neighborCr)>Th_Ch where Th_Ch is a predefined value or a value signaled in the bitstream, and Abs is the absolute value function. Th_Ch may depend on the bit depth of the chroma samples. Typically, for 10-bit content, Th_Ch=64, and for 8-bit content, Th_Ch=32.

[0159] In a modified version, chroma residual scaling is applied only if the average Cb and Cr values avgNeighCb, avgNeighCr of the chroma samples co-located with the neighboring luma samples are not too far from the average Cb and Cr values avgCurrCb, avgCurrCr of the chroma samples of the current chroma block.

[0160] For example, we consider adjacent luma samples unavailable if the following conditions are true: -Abs(avgCurrCb-avgNeighCb)>Th_Ch, or Abs(avgCurrCr-avgNeighCr)>Th_Ch is established.

[0161] In a modification, instead of strictly invalidating samples that are too different from the top-left chroma sample, a weighting depending on the difference between the adjacent chroma sample and the top-left chroma sample is applied to the reference sample being considered.

[0162] In a modification, instead of strictly invalidating samples that are too different from the top-left chroma sample, a weighting depending on the difference between the adjacent chroma sample and the top-left chroma sample is applied to the reference sample being considered.

[0163] Figure 25 shows a flow diagram of a method for checking the availability of luma samples based on their co-located chroma samples and the chroma samples of a current block. Steps 800 to 802 are similar to the corresponding steps in Figure 22. Step 900 is inserted after step 802 to identify adjacent chroma samples that are co-located with the adjacent luma samples. Step 901 checks the similarity of the adjacent chroma samples with the chroma samples of the current block. If the chroma samples are similar, chroma residual scaling is enabled for the current chroma block (903). If the samples are deemed not similar, chroma residual scaling is disabled for the current chroma block (902).

[0164] Embodiment 12a: Availability of reference samples for MDLM The MDLM mode is a modification of the CCLM mode that uses the neighboring luma and chroma samples on the top (MDLM-top) or the neighboring luma and chroma samples on the left (MDLM-left) as reference samples to derive linear parameters for predicting chroma samples from luma samples.

[0165] In at least one embodiment, if the splitting process results in the splitting of a luma VDPU containing two upper square blocks with width / height half the width / height of the VDPU, or the splitting of a chroma VDPU containing two upper square blocks with width / height half the width / height of the VDPU, the reference samples for the MDLM left of the second square block (denoted by "2" in FIG. 25A ) are only the samples adjacent to the left boundary of block 2. The samples below those reference samples are considered unavailable. This advantageously limits the delay of processing block 2, as the blocks below block 2 are not needed to derive the MDLM parameters.

[0166] This embodiment can also be generalized to the case of further dividing luma or chroma block 1 of FIG. 25A into smaller partitions.

[0167] Embodiment 13: Low-level flag for activating or deactivating chroma residual scaling In another embodiment, a low-level flag is inserted in the bitstream syntax to enable or disable the chroma residual scaling tool at a level lower than the slice or tile group. This embodiment may only be applied in the case of separate luma / chroma split trees. This embodiment can also be extended to cases where separate luma / chroma split trees are not used.

[0168] The table below shows an example of syntax changes (highlighted in grey) compared to the VTM specification in document JVET-N0220, where the signaling is done at the CTU level and only in the case of separate luma / chroma split trees (identified by the syntax element qtbtt_dual_tree_intra_flag).

[0169] [Table 6]

[0170] The table below shows another example of syntax changes (highlighted in grey) compared to the VTM specification in document JVET-N0220, where the signaling is done at the CTU level, and the signaling is done whether or not any separate luma / chroma splitting tree is used.

[0171] [Table 7]

[0172] According to these modifications, the activation in the CTU of the chroma residual scaling is conditioned by the value of the flag ctu_chroma_residual_scale_flag.

[0173] Example of application of specification syntax according to embodiment 2 An example of syntax changes in the current VVC draft specification (see Benjamin Bross et al. “Versatile Video Coding (Draft 4)”, JVET 13th Meeting: Marrakech, MA, 9-18 Jan. 2019, JVET-M1001-v7) according to embodiment 2 is shown below. Changes compared to JVET-M1001 are highlighted in gray.

[0174] 8.4.3 Derivation Process for Chroma Intra Prediction Modes The inputs to this process are: - a luma position (xCb, yCb) that specifies the top-left sample of the current chroma-coded block relative to the top-left luma sample of the current picture; - the variable cbWidth, which specifies the width of the current coded block in luma samples, - The variable cbHeight that specifies the height of the current coded block in luma samples.

[0175] In this process, the chrominance intra-prediction mode IntraPredModeC[xCb][yCb] is derived. -Set the variable lmEnabled equal to sps_cclm_enabled_flag. If the following condition is true, lmEnabled is derived by calling the Inter-Component Chroma Intra Prediction Mode Check Process: -lmEnabled equals 1, -tile_group_type equals 2 (I tile group), -qtbtt_dual_tree_intra_flag equals 1.

[0176] The chrominance intra prediction mode IntraPredModeC[xCb][yCb] is derived using intra_chroma_pred_mode[xCb][yCb] and IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2] as specified in Tables 8-2 and 8-3.

[0177] [Table 8]

[0178] [Table 9]

[0179] Inter-component chroma intra prediction mode inspection process The inputs to this process are: - a luma position (xCb, yCb) that specifies the top-left sample of the current chroma-coded block relative to the top-left luma sample of the current picture; - the variable cbWidth, which specifies the width of the current coded block in luma samples, - The variable cbHeight that specifies the height of the current coded block in luma samples.

[0180] The output to this process is: - A flag lmEnabled that specifies whether the inter-component chroma intra prediction mode is enabled for the current chroma coded block. In this process, lmEnabled is derived as follows: If -((16+xCb) / 16) is not equal to ((16+xCb+cbWidth-1) / 16) or ((16+yCb) / 16) is not equal to ((16+yCb+cbHeight-1) / 16), set cclmEnabled equal to 0 - Otherwise the following applies: -For i=0, (cbHeight-1) and j=0..(cbWidth-1), the following applies: Set −(xTL, yTL) equal to the top-left sample position relative to the top-left luma sample of the current picture of the collocated luma coding block ColLumaBlock that covers the position given by ((xCb+j)<<1, (yCb+i)<<1)). -Set wColoc and hColoc equal to the width and height of the ColLumaBlock. - If one of the following conditions is false, set cclmEnabled equal to 0 and stop looping over i and j: -((32+xTL) / 32) is equal to ((32+xTL+wColoc-1) / 32) -((32+yTL) / 32) is equal to ((32+yTL+hColoc-1) / 32) When -cclmEnabled is equal to 1, for i=1..(cbHeight-2) and j=0, (cbWidth-1), the following applies: Set −(xTL, yTL) equal to the top-left sample position relative to the top-left luma sample of the current picture of the collocated luma coding block ColLumaBlock that covers the position given by ((xCb+j)<<1, (yCb+i)<<1)). -Set wColoc and hColoc equal to the width and height of the ColLumaBlock. - If one of the following conditions is false, set cclmEnabled equal to 0 and stop looping over i and j: -((32+xTL) / 32) is equal to ((32+xTL+wColoc-1) / 32) -((32+yTL) / 32) is equal to ((32+yTL+hColoc-1) / 32)

[0181] Example of application of specification syntax according to embodiment 10 An example of syntax for inclusion in the current VVC draft specification (eg, document JVET-M1001) according to embodiment 10 is shown below:

[0182] Picture reconstruction with a mapping process for chroma sample values - Patent Application 20070122997 The inputs to this process are: - predMapSamples, a (nCurrSwx2)x(nCurrShx2) array mapped, specifying the mapped luma prediction samples of the current block; - array recSamples specifying the reconstructed luma of the current picture if tile_group_type is equal to 2 (I tile group) and qtbtt_dual_tree_intra_flag is equal to 1, - a (nCurrSw) x (nCurrSh) array predSamples specifying the chroma prediction samples for the current block, - resamples, a (nCurrSw)x(nCurrSh) array specifying the chroma residual samples for the current block. - Arrays InputPivot[i] and ReshapePivot[i], where i is in the range from 0 to MaxBinIdx+1, derived in 7.4.4.1 - the arrays InvScaleCoeff[i] and ChromaScaleCoef[i], with i in the range 0 to MaxBinIdx, derived in 7.4.4.1

[0183] The output of this process is the reconstructed chroma sample array recSamples.

[0184] recSamples is derived as follows: -(!tile_group_reshaper_chroma_residual_scale_flag||((nCurrSw)x(nCurrSh)<=4)) is true recSamples[xCurr+i][yCurr+j]=Clip1c(predSamples[i][j]+resSamples[i][j]) i=0..nCurrSw-1, j=0..nCurrSh-1 -Otherwise, (tile_group_reshaper_chroma_residual_scale_flag&&((nCurrSw)x(nCurrSh)>4)) holds, and the following applies:

[0185] The variable varScale is derived as follows: -invAvgLuma is derived as follows: -If tile_group_type is equal to 2 (I tile group) and qtbtt_dual_tree_intra_flag is equal to 1, the following applies: -Identify the chroma position (xCh, yCh) of the top left sample of the current block of chroma relative to the top left chroma sample of the current picture. - Set the luma position (xTL, yTL) equal to the top-left sample position of the collocated luma coding block ColLumaBlock that covers the position given by (xCh<<1, yCh<<1) relative to the top-left luma sample of the current picture, and set wColoc and hColoc equal to the width and height of ColLumaBlock. If -((32+yCh) / 32) is equal to ((64+yTL-1) / 64), set invAvgLuma equal to recSamples[yTL-1][xTL] - Otherwise, if ((32+xCh) / 32) equals ((64+xTL-1) / 64), set invAvgLuma equal to recSamples[yTL][xTL-1] Otherwise, set invAvgLuma equal to -1. - Otherwise the following applies: invAvgLuma=Clip1 Y ((Σ i Σ j predMapSamples[i][j]+nCurrSw*nCurrSh*2) / (nCurrSw*nCurrSh*4)) If -invAvgLuma is not equal to -1, the following applies: - The variable idxYInv is derived by associating the identification of the segmentation function index as defined in Section 8.5.6.2 with the input of the sample value invAvgLuma. - Set varScale equal to ChromaScaleCoef[idxYInv] - Otherwise, set varScale equal to (1<<shiftC)

[0186] recSamples is derived as follows: - If tu_cbf_cIdx[xCurr][yCurr] is equal to 1, the following applies: shiftC = 11 recSamples[xCurr + i][yCurr + j] = ClipCidx1(predSamples[i][j] + Sign(resSamples[i][j]) *((Abs(resSamples[i][j])*varScale+(1<<(shiftC - 1)))>>shiftC)) i = 0..nCurrSw - 1, j = 0..nCurrSh - 1 - Otherwise (tu_cbf_cIdx[xCurr][yCurr] is equal to 0) recSamples[xCurr + i][yCurr + j] = ClipCidx1(predSamples[i][j]) i = 0..nCurrSw - 1, j = 0..nCurrSh - 1

[0187] This application describes various aspects including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described and are often described in a way that may seem limiting in order to show at least individual characteristics. However, this is for the purpose of clarifying the description and does not limit the application or scope of those aspects. In fact, all of the various aspects can be combined and exchanged to result in further aspects. Additionally, aspects can also be combined and exchanged with the aspects described in previous applications.

[0188] The aspects described and contemplated herein can be implemented in many different forms. While Figures 26, 27, and 28 below illustrate some embodiments, other embodiments are contemplated, and the discussion of Figures 26, 27, and 28 is not intended to limit the scope of implementations. At least one aspect generally relates to encoding and decoding video, and at least one other aspect generally relates to transmitting generated or encoded bitstreams. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.

[0189] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for a method to operate properly, the order and / or use of specific steps and / or actions can be modified or combined.

[0190] Various methods and other aspects described herein can be used to modify modules, such as the image segmentation and scaling modules (102, 151, 235, and 251) of video encoder 100 and decoder 200 shown in Figures 26 and 27. Furthermore, aspects of the present invention are not limited to VVC or HEVC, but can be applied, for example, to other standards and recommendations, existing or developed in the future, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, aspects described herein can be used individually or in combination.

[0191] Various numerical values are used herein, such as Wmax, Hmax, WmaxC, and HmaxC. The specific values are for illustrative purposes and the described aspects are not limited to those specific values.

[0192] 26 shows an encoder 100. Although variations of this encoder 100 are possible, encoder 100 is described below for the sake of clarity without listing all contemplated variations.

[0193] Before being encoded, the video sequence may be subjected to pre-encoding processing (101), for example, applying a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or remapping the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and may be added to the bitstream.

[0194] In encoder 100, a picture is coded by the encoder elements as described below. The picture to be coded is divided (102) into, for example, CU units and processed. Each unit is coded, for example, using intra mode or inter mode. If a unit is coded in intra mode, the encoder performs intra prediction (160). In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) whether to use intra mode or inter mode to code the unit and indicates the intra / inter decision, for example, with a prediction mode flag. The encoder may perform forward mapping (191) applied to luma samples to obtain a predicted luma block. For chroma samples, forward mapping is not applicable. For example, a prediction residual is calculated by subtracting (110) the predicted block from the original image block. For chroma samples, chroma residual scaling may be applied to the chroma residual (111).

[0195] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.

[0196] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The chroma residual is then processed by inverse scaling (151), which reverses the scaling process (111). The decoded prediction residual and the predicted block are combined (155) to reconstruct an image block. For luma samples, an inverse mapping (190), which is the inverse of the forward mapping step (191), may be applied. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / sample adaptive offset (SAO) or adaptive loop filter (ALF) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).

[0197] Figure 27 shows a block diagram of a video decoder 200, in which the bitstream is decoded by elements of the decoder as described below. The video decoder 200 generally performs a decoding pass that is the inverse of the encoding pass described in Figure 26. The encoder 100 also generally performs video decoding as part of encoding the video data.

[0198] Specifically, the decoder input includes a video bitstream, such as might be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coding information. Picture partition information indicates how the picture is partitioned. The decoder can then divide the picture according to the decoded picture partition information (235). To decode the prediction residual, the transform coefficients are dequantized (240) and inverse transformed (250). For chroma samples, the chroma residual is processed by inverse scaling (251), similar to the inverse scaling (151) of the encoder. The decoded prediction residual is combined (255) with the predicted block to reconstruct an image block. The predicted block can result from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (270). After prediction, forward mapping (295) can be applied to the luma samples. An inverse mapping (296) similar to the inverse mapping (190) of the encoder can be applied to the reconstructed luma samples. An in-loop filter (265) is then applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).

[0199] The decoded picture can be further subjected to post-decoding processing (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that reverses the remapping process performed in the pre-encoding processing (101). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0200] FIG. 28 illustrates a block diagram of an example system in which various aspects and embodiments can be implemented. System 1000 can be implemented as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and servers. Elements of system 1000 can be implemented singly or in combination within a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or separate components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or by dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described herein.

[0201] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a nonvolatile memory device). The system 1000 includes storage 1040, which may include nonvolatile and / or volatile memory including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. The storage 1040 may include, by way of non-limiting example, internal storage, additional storage (including removable and non-removable storage), and / or network-accessible storage.

[0202] System 1000 includes an encoder / decoder module 1030 configured to process data to provide, for example, encoded video or decoded video, which may include its own processor and memory. The encoder / decoder module 1030 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software, as is known to those skilled in the art.

[0203] Program code loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in storage 1040 and then loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, memory 1020, storage 1040, and encoder / decoder module 1030 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and processing of computational logic.

[0204] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing unit (e.g., the processing unit may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage 1040, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used, for example, to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, an emerging standard being developed by JVET, the Joint Video Experts Team).

[0205] Input to the elements of system 1000 may be provided by various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF (radio frequency) signals transmitted wirelessly, for example by a broadcaster, (ii) a component (COMP) input (or a set of COMP inputs), (iii) a universal serial bus (USB) input, and / or (iv) a high-definition multimedia interface (HDMI) input. Other examples not shown in FIG. 28 include composite video.

[0206] In various embodiments, the input devices of block 1130 have associated individual input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel in certain embodiments (for example), (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band-limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner for performing various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In some set-top box embodiments, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of them, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0207] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, e.g., Reed-Solomon error correction, may be implemented as needed, for example, in a separate input processing IC or within processor 1010. Similarly, aspects of USB or HDMI interface processing may be implemented as needed in a separate interface IC or within processor 1010. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030 operating in combination with memory and storage elements, to process the data stream as needed for presentation on an output device.

[0208] The various elements of the system 1000 may be provided within a unitary housing, in which the various elements may be interconnected and data may be transmitted therebetween using a suitable connection arrangement 1140, such as an internal bus known in the art, including an Inter-IC (I2C) bus, wiring, and a printed circuit board. The system 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0209] In various embodiments, data is streamed or provided to system 1000 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received over communication channel 1060 and communication interface 1050, which are adapted for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that delivers data over the HDMI connection of input block 1130. Still other embodiments provide streamed data to system 1000 using the RF connection of input block 1130. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0210] System 1000 can provide output signals to various output devices, including a display 1100, speakers 1110, and other peripherals 1120. Display 1100 in various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 1100 can be for a television, a tablet, a laptop, a cell phone, or other device. Display 1100 can be integrated with other components (e.g., as in a smartphone) or can be separate (e.g., an external monitor for a laptop). In various example embodiments, other peripherals 1120 include one or more of a stand-alone digital video disc (or digital versatile disc) (DVR in both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripherals 1120 to provide functionality based on the output of system 1000. For example, a disc player performs the function of playing the output of system 1000.

[0211] In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow inter-device control with or without user intervention. Output devices may be communicatively coupled to system 1000 via dedicated connections through individual interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speakers 1110 may be integrated into a single unit with other components of system 1000 in an electronic device such as a television. In various embodiments, display interface 1070 includes a display driver, such as a timing controller (T Con) chip.

[0212] For example, if the RF portion of input 1130 is part of a separate set-top box, display 1100 and speakers 1110 can alternatively be separate from one or more of the other components. In various embodiments in which display 1100 and speakers 1110 are external components, the output signal can be provided by a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.

[0213] The embodiments may be performed by computer software implemented by the processor 1010, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, including, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technical environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0214] Various implementations include decoding. As used herein, "decoding" may encompass all or part of the processes performed on a received encoded sequence to, for example, produce a final output suitable for display. In various embodiments, such processes include one or more of processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or instead include processes performed by decoders in various implementations described herein that enable inter-component dependency tools for a chroma block depending on the size of the chroma block and the size of at least one luma block that colocates with at least one sample of the chroma block.

[0215] As a further example, in some embodiments, "decoding" refers to entropy decoding only, in other embodiments "decoding" refers to differential decoding only, and in other embodiments "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.

[0216] Various implementations include encoding. Similar to the above discussion regarding "decoding," "encoding," as used herein, may encompass all or part of the processes performed on an input video sequence, for example, to result in a coded bitstream. In various embodiments, such processes include one or more of processes typically performed by an encoder, such as partitioning, differential coding, transforming, quantizing, and entropy coding. In various embodiments, such processes also or instead include processes performed by the encoder in various implementations described herein that enable inter-component dependency tools to be used on a chroma block, for example, depending on the size of the chroma block and depending on the size of at least one luma block that colocates with at least one sample of the chroma block.

[0217] As a further example, in some embodiments, "encoding" refers only to entropy encoding, in other embodiments "encoding" refers only to differential encoding, and in other embodiments "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.

[0218] It should be noted that the syntax elements used herein, e.g., chroma scale factor indices, are descriptive terms, and as such, they do not preclude the use of other syntax element names.

[0219] Where a drawing is shown as a flow diagram, it should be understood that the drawing also provides a block diagram of the corresponding apparatus. Similarly, where a drawing is shown as a block diagram, it should be understood that the drawing also provides a flow diagram of the corresponding method / process.

[0220] Various embodiments have referred to rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often given a computational complexity constraint. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are various approaches to solving the rate-distortion optimization problem. For example, these approaches may involve a thorough evaluation of the coding cost and the associated distortion of the reconstructed signal after coding and decoding, and may be based on a comprehensive testing of all coding options, including all modes or coding parameter values considered. Faster approaches can also be used to reduce coding complexity, particularly by calculating approximate distortion based on a predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some of the coding options and full distortion for other coding options. Other approaches evaluate only a subset of the coding options. More generally, many approaches use any of a variety of techniques to perform optimization, but the optimization does not necessarily involve a thorough evaluation of both the coding cost and the associated distortion.

[0221] The implementations and aspects described herein may be implemented by, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single type of implementation (e.g., only discussed as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented by, for example, appropriate hardware, software, and firmware. A method may be implemented by, for example, a processor, which refers generally to processing devices, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0222] Reference to "one embodiment," or "an embodiment," or "one implementation," or "an implementation," as well as other variants thereof, means that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment," or "in an embodiment," or "in one implementation," or "in an implementation," as well as any other variants, in various places throughout this application do not necessarily all refer to the same embodiment.

[0223] Additionally, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0224] Additionally, this application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, replicating information, calculating information, determining information, predicting information, or estimating information.

[0225] Additionally, this application may refer to "receiving" various pieces of information. Receiving is intended to be broad, similar to "accessing." Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves some form of operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0226] For example, the use of " / ", "and / or", and "at least one of" in the cases of "A / B", "A and / or B", and "at least one of A and B" is intended to encompass selecting only the first listed (A) option, or selecting only the second listed (B) option, or selecting both (A and B) options. As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to encompass selecting only the first listed (A) option, or selecting only the second listed (B) option, or selecting only the third listed (C) option, or selecting only the first and second listed options (A and B), or selecting only the first and third listed options (A and C), or selecting only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this representation can be expanded to include as many items as are listed.

[0227] Furthermore, as used herein, the term "signaling" refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a particular chroma scale factor index. In this manner, the same parameters are used at both the encoder and decoder sides in one embodiment. Thus, for example, an encoder can transmit a particular parameter to a decoder (explicit signaling), allowing the decoder to use the same particular parameter. Conversely, if the decoder already has that particular parameter along with other parameters, signaling can be used without transmission simply to allow the decoder to know and select that particular parameter (implicit signaling). By avoiding transmitting any actual functionality, bit savings are realized in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. While the above content relates to the verb form of the word "signal," the word "signal" can also be used as a noun herein.

[0228] As will be apparent to those skilled in the art, implementations can result in various signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0229] Several embodiments have been described. Features of these embodiments may be provided alone or in any combination across various claim categories and types. Furthermore, embodiments may include one or more of the following features, devices, or aspects across various claim categories and types, alone or in any combination: Enabling the inter-component dependency tool depending on the size of the chroma block and possibly depending on the size of at least one luma block that is co-located with the chroma block. Enabling the inter-component dependency tool when there are chroma blocks within a single chroma rectangle, where the chroma rectangle is defined by dividing the chroma components of a picture into non-overlapping rectangular regions; enabling an inter-component dependency tool when there is at least one co-located luma block in a luma rectangular area that is co-located with a first chroma rectangular area, the luma rectangular area being defined by dividing the luma of the picture into non-overlapping rectangular areas; A bitstream or signal containing one or more of the syntax elements described or variations thereof. A bitstream or signal containing syntax conveying information generated according to any of the described embodiments. Inserting syntax elements into the signaling that allow the decoder to enable / disable inter-component dependency tools in a manner corresponding to that used by the encoder. Creating and / or transmitting, and / or receiving and / or decoding bitstreams or signals that include one or more of the described syntax elements or variations thereof. · Creation and / or transmission and / or reception and / or decoding according to any of the described embodiments. A method, process, apparatus, instruction storage medium, data storage medium, or signal according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that enables / disables a cross-component dependency tool according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that enables / disables a cross-component dependency tool according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). A TV, set-top box, mobile phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal containing encoded images and enables / disables an inter-component dependency tool according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that wirelessly receives (e.g., using an antenna) a signal containing the encoded image and enables / disables the inter-component dependency tool according to any of the described embodiments.

[0230] Furthermore, embodiments may include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types.

[0231] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block co-located with the chroma block; decoding said chroma blocks in response to said enabling of said inter-component dependency tool; A decoding method is presented, including:

[0232] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block that colocates with the chroma block; and decoding said chroma blocks in response to said enabling of said inter-component dependency tool; A decoding device is presented that includes one or more processors configured to:

[0233] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block co-located with the chroma block; encoding the chroma blocks in response to the enabling of the inter-component dependency tool. A coding method is presented, including:

[0234] According to a general aspect of at least one embodiment, - enabling an inter-component dependency tool for use on a chroma block of a picture depending on the size of the chroma block and depending on the size of at least one luma block that colocates with the chroma block; and encoding the chroma blocks in response to the enabling of the inter-component dependency tool. An encoding device is presented that includes one or more processors configured to:

[0235] One or more embodiments of the present invention also provide a computer-readable storage medium storing instructions for encoding or decoding video data according to at least a portion of any of the above methods. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the above encoding methods. One or more embodiments also provide methods and apparatus for transmitting or receiving a bitstream generated according to the above encoding methods. One or more embodiments also provide a computer program product including instructions for performing at least a portion of any of the above methods.

[0236] In one embodiment, the inter-component dependency tool is luma-dependent chroma residual scaling.

[0237] In an embodiment, enabling said luma-dependent chroma residual scaling for use on said chroma blocks of a picture according to a size of the chroma blocks and according to a size of at least one luma block co-located with the chroma blocks comprises: -The luma-dependent chroma residual scaling mentioned above the chroma blocks are within a single chroma rectangular region, the chroma rectangular region being defined by dividing the chroma components of the picture into non-overlapping rectangular regions; and At least one luma block co-located with a chroma block within a single luma rectangular area co-located with said chroma rectangular area, said luma rectangular area being defined by dividing the luma of said picture into non-overlapping rectangular areas. to enable it, and - otherwise disable the luma-dependent chroma residual scaling mentioned above Includes.

[0238] In an embodiment, enabling said luma-dependent chroma residual scaling for use on said chroma blocks of a picture according to a size of the chroma blocks and according to a size of at least one luma block co-located with the chroma blocks comprises: -The luma-dependent chroma residual scaling mentioned above the chroma blocks are within a single chroma rectangular region, the chroma rectangular region being defined by dividing the chroma components of the picture into non-overlapping rectangular regions; and All luma blocks that are co-located with a chroma block are within a single luma rectangular area that is co-located with said chroma rectangular area, said luma rectangular area being defined by dividing the luma of said picture into non-overlapping rectangular areas. to enable it, and - otherwise disable the luma-dependent chroma residual scaling mentioned above Includes.

[0239] In an embodiment, enabling said luma-dependent chroma residual scaling for use on said chroma blocks of a picture according to a size of the chroma blocks and according to a size of at least one luma block co-located with the chroma blocks comprises: enabling the luma-dependent chroma residual scaling when there is at least one luma block co-located with the chroma block in a single luma rectangular area co-located with a first chroma rectangular area, the chroma rectangular area being defined by dividing chroma components of the picture into non-overlapping rectangular areas, and the luma rectangular area being defined by dividing luma of the picture into non-overlapping rectangular areas; and - otherwise disable the luma-dependent chroma residual scaling mentioned above Includes.

[0240] In one embodiment, when the luma-dependent chroma residual scaling is disabled, the scale factor for the chroma block is determined by making a prediction from a chroma scale factor associated with at least one decoded (respectively encoded) chroma block adjacent to the chroma block, or from an index identifying such a chroma scale factor.

[0241] In one embodiment, if said luma-dependent chroma residual scaling is disabled, then delta quantization parameters are decoded from (respectively encoded within) the bitstream for said chroma blocks.

[0242] In another embodiment, the inter-component dependency tool is inter-component linear model prediction.

Claims

1. For chroma blocks of a picture within a single non-overlapping rectangular region, - determining a luma non-overlapping rectangular region that includes a luma block that is co-located with the position of said chroma block; determining a scale factor based on an average of reconstructed luma samples belonging to a line below and / or a column to the right of luma samples belonging to adjacent luma non-overlapping rectangular areas above and / or to the left of said luma non-overlapping rectangular area containing the current chroma block; and - scaling the residuals of said chroma blocks according to said scale factors; A decoding method comprising:

2. The method of claim 1 , wherein the chroma residual scaling tool is enabled according to size and position constraints of the co-located chroma and luma blocks.

3. 3. The method of claim 2, wherein the chroma block has a size of 16x16 pixels or less and at least one luma block co-located with the chroma block has a size of 32x32 pixels or less.

4. For chroma blocks of a picture within a single non-overlapping rectangular region, - determining a luma non-overlapping rectangular region that includes a luma block that is co-located with the position of said chroma block; determining a scale factor based on an average of reconstructed luma samples belonging to a line below and / or a column to the right of luma samples belonging to adjacent luma non-overlapping rectangular areas above and / or to the left of said luma non-overlapping rectangular area containing the current chroma block; and - scaling the residuals of said chroma blocks according to said scale factors; 10. An encoding method comprising:

5. The method of claim 4 , wherein the chroma residual scaling tool is enabled according to size and position constraints of the collocated chroma and luma blocks.

6. The method of claim 5 , wherein the chroma block has a size of 16x16 pixels or less and at least one luma block co-located with the chroma block has a size of 32x32 pixels or less.

7. For chroma blocks of a picture within a single non-overlapping rectangular region, - determining a luma non-overlapping rectangular region that includes a luma block that is co-located with the position of said chroma block; determining a scale factor based on an average of reconstructed luma samples belonging to a line below and / or a column to the right of luma samples belonging to adjacent luma non-overlapping rectangular areas above and / or to the left of said luma non-overlapping rectangular area containing the current chroma block; and - scaling the residuals of said chroma blocks according to said scale factors; 11. A decoding device comprising: one or more processors configured to:

8. The apparatus of claim 7 , wherein the chroma residual scaling tool is enabled according to size and position constraints of the collocated chroma and luma blocks.

9. 9. The device of claim 8, wherein the chroma block has a size of 16x16 pixels or less and at least one luma block co-located with the chroma block has a size of 32x32 pixels or less.

10. For chroma blocks of a picture within a single non-overlapping rectangular region, - determining a luma non-overlapping rectangular region that includes a luma block that is co-located with the position of said chroma block; determining a scale factor based on an average of reconstructed luma samples belonging to a line below and / or a column to the right of luma samples belonging to adjacent luma non-overlapping rectangular areas above and / or to the left of said luma non-overlapping rectangular area containing the current chroma block; and - scaling the residuals of said chroma blocks according to said scale factors; 1. An encoding device comprising: one or more processors configured to:

11. The apparatus of claim 10 , wherein the chroma residual scaling tool is enabled according to size and position constraints of the collocated chroma and luma blocks.

12. 12. The device of claim 11, wherein the chroma block has a size of 16x16 pixels or less and at least one luma block co-located with the chroma block has a size of 32x32 pixels or less.

13. A computer program comprising instructions for carrying out the method of any one of claims 1 to 3 when executed by one or more processors.

14. A non-transitory computer readable medium containing a computer program comprising instructions for performing the method of any one of claims 1 to 3 when executed by one or more processors.

Citation Information

Patent Citations

  • Linear model prediction mode with sample access for video coding

    JP2020502925A

  • Image shaping in video coding using rate-distortion optimization.

    JP2021513284A

  • Method, apparatus, and program for small block prediction and transformation

    JP2021518088A

  • Linear model prediction mode with sample accessing for video coding

    US20180176594A1