Parameter derivation in cross-component mode
By optimizing chroma residual scaling and cross-component linear model processing through selective enabling/disabling and conditional syntax signaling, the latency and complexity issues in video coding are addressed, enhancing processing efficiency in dual/separate tree structures.
Patent Information
- Application Number
- JP2021559969
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-18
- Filing Date
- 2020-04-20
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2040-04-20
AI Technical Summary
Existing video coding standards face challenges with high latency and computational complexity due to cross-component dependencies in luma-dependent chroma residual scaling and cross-component linear model prediction, particularly in dual/separate tree structures, leading to inefficiencies in processing chroma samples.
Implement methods to reduce cross-component dependencies by deriving chroma residual scaling factors based on available neighboring luma blocks and selectively enabling/disabling cross-component linear models, using conditional syntax signaling to optimize processing.
Reduces latency and computational complexity by enabling efficient processing of chroma samples, improving video coding efficiency and reducing processing delays in dual/separate tree structures.
Smart Images

Figure 0007732898000032 
Figure 0007732898000033 
Figure 0007732898000034
Abstract
Description
[Technical Field]
[0001] [Related Applications] Book This application is a subsidiary of International Patent Application No. PCT / CN2019 / 083320, filed April 18, 2019. rights Based on claimed International Patent Application No. PCT / CN2020 / 085674, filed April 20, 2020. All of the foregoing patent applications are incorporated herein by reference in their entirety.
[0002] [Technical field] TECHNICAL FIELD This disclosure relates to video coding and decoding techniques, devices, and systems. [Background technology]
[0003] Despite advances in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth requirements for digital video usage are expected to continue to grow. Summary of the Invention
[0004] Apparatuses, systems, and methods relate to digital video coding / decoding, and in particular, a simplified linear model derivation for cross-component linear model (CCLM) prediction modes in video coding / decoding is described. The described methods are applicable to both existing video coding standards (e.g., High Efficiency Video Coding (HEVC)) and future video coding standards (e.g., Versatile Video Coding (VVC)) or codecs.
[0005] In one representative aspect, a method of video media processing is disclosed that includes performing a conversion between a current chroma video block of video media data and a bitstream representation of the current chroma video block, wherein during the conversion, a chroma residual of the current chroma video block is scaled based on a scaling factor, the scaling factor being derived based at least on a luma sample located at a predetermined location.
[0006] In one representative aspect, a method of video media processing is disclosed that includes performing a conversion between a current video block of video media data and a bitstream representation of the current video block, wherein during the conversion, a second set of color component values for the current video block are derived from a first set of color component values of the video media data using a cross-component linear model (CCLM) and / or luma mapping with chroma scaling (LMCS) mode processing step.
[0007] In another representative aspect, a method of video media processing is disclosed that includes performing a conversion between a current video block of video media data and a bitstream representation of the current video block, wherein during the conversion, one or more reconstructed samples associated with a current frame of the video media data are used to derive a chroma residual scaling factor in a luma mapping with chroma scaling (LMCS) mode processing step.
[0008] In another representative aspect, a method of video media processing is disclosed that includes performing a conversion between a current video block of video media data and a bitstream representation of the current video block, wherein during the conversion, one or more luma prediction samples or luma reconstructed samples in a current frame other than a reference frame are used to derive a chroma residual scaling factor in a luma mapping with chroma scaling (LMCS) mode processing step.
[0009] In another representative aspect, a method of video media processing is disclosed, the method comprising: checking availability of one or more neighboring luma blocks of a corresponding luma block that cover a top-left sample of a co-located luma block during conversion between a current chroma video block and a bitstream representation of the current chroma video block; determining whether to search for neighboring luma samples of the corresponding luma block based on the availability of one or more neighboring luma blocks; deriving a scaling factor based on said determining step; scaling a chroma residual of the current chroma video block based on the scaling factor to generate a scaled chroma residual; performing the transformation based on the scaled chroma residual.
[0010] In another representative aspect, a method of video media processing is disclosed, the method comprising: During a conversion between a current video block of video media data and a bitstream representation of the current video block, the method includes using a model associated with a processing step to derive a second set of color component values for the current video block from a first set of color component values of the video media data, the first set of color component values being neighboring samples of a corresponding luma block that cover a top-left sample of the co-located luma block.
[0011] In another representative aspect, a method of video media processing is disclosed that includes determining, during conversion between a current chroma video block of video media data and a bitstream representation of the current chroma video block, to selectively enable or disable application of a cross-component linear model (CCLM) and / or chroma residual scaling (CRS) to the current chroma video block based at least in part on one or more conditions associated with a co-located luma block of the current chroma video block.
[0012] In another representative aspect, a method of video media processing is disclosed, the method comprising: selectively enabling or disabling application of luma-dependent chroma residual scaling (CRS) to chroma components of a current video block in a video domain of the video media data for encoding the current video block into a bitstream representation of the video media data; determining to include or exclude a field in the bitstream representation of the video media data, the field indicating the selective enabling or disabling, and if included, signaled outside a first syntax level associated with the current video block.
[0013] In another representative aspect, a method of video media processing is disclosed, the method comprising: Parsing a field within a bitstream representation of video media data, the field being included in a level other than a first syntax level currently associated with the video block; and selectively enabling or disabling application of luma-dependent chroma residual scaling (CRS) to chroma components of the current video block of video media data to generate a decoded video region from the bitstream representation based on the field.
[0014] In another representative aspect, a method of video media processing is disclosed, the method comprising: selectively enabling or disabling application of a cross-component linear model (CCLM) to the current video block of video media data for encoding the current video block into a bitstream representation of the video media data; determining to include or exclude a field in the bitstream representation of the video media data, the field indicating the selective enabling or disabling, and if included, signaled outside a first syntax level associated with the current video block.
[0015] In another representative aspect, a method of video media processing is disclosed, the method comprising: Parsing a field within a bitstream representation of video media data, the field being included in a level other than a first syntax level currently associated with the video block; and selectively enabling or disabling application of a cross-component linear model (CCLM) to the current video block of video media data based on the field to generate a decoded video region from the bitstream representation.
[0016] In yet another exemplary aspect, a video encoder or decoder device is disclosed that includes a processor configured to perform the above-described method.
[0017] According to another exemplary aspect, a computer-readable program medium is disclosed, the medium storing code embodying processor-executable instructions for performing one of the disclosed methods.
[0018] In yet another exemplary aspect, the above-described methods are embodied in the form of processor-executable code and stored on a computer-readable program medium.
[0019] In yet another exemplary aspect, an apparatus configured or operative to perform the above-described method is disclosed. The apparatus may include a processor programmed to implement the method.
[0020] These and other aspects and features of the disclosed technology are explained in further detail in the drawings, description, and claims. [Brief explanation of the drawings]
[0021] [Figure 1] 1 shows an example of an angular intra prediction mode in HEVC. [Figure 2] Here is an example of a non-HEVC directional mode: [Figure 3] An example related to CCLM mode is shown below. [Figure 4] 1 shows an example of luma mapping with a chroma scaling architecture. [Figure 5] 1 shows examples of luma and chroma blocks for different color formats. [Figure 6] An example of a luma block and a chroma block of the same color format is shown. [Figure 7] 1 shows an example of a co-located luma block covering multiple formats. [Figure 8] 1 shows an example of a luma block within a larger luma block. [Figure 9] 1 shows an example of a luma block within a larger luma block and within a bounding box. [Figure 10]FIG. 1 is a block diagram of an example hardware platform for implementing the video media decoding or video media encoding techniques described herein. [Figure 11] 1 shows a flowchart of an exemplary method for deriving a linear model for cross-component prediction in accordance with the disclosed technique. [Figure 12] FIG. 1 is a block diagram of an exemplary video processing system in which the disclosed techniques may be implemented. [Figure 13] 1 shows a flowchart of an exemplary method for video media processing. [Figure 14] 1 shows a flowchart of an exemplary method for video media processing. [Figure 15] 1 shows a flowchart of an exemplary method for video media processing. [Figure 16] 1 shows a flowchart of an exemplary method for video media processing. [Figure 17] 1 shows a flowchart of an exemplary method for video media processing. [Figure 18] 1 shows a flowchart of an exemplary method for video media processing. [Figure 19] 1 shows a flowchart of an exemplary method for video media processing. [Figure 20] 1 shows a flowchart of an exemplary method for video media processing. [Figure 21] 1 shows a flowchart of an exemplary method for video media processing. [Figure 22] 1 shows a flowchart of an exemplary method for video media processing. [Figure 23] 1 shows a flowchart of an exemplary method for video media processing. DETAILED DESCRIPTION OF THE INVENTION
[0022] 2.1 Overview of HEVC 2.1.1 Intra Prediction in HEVC / H.265 Intra prediction involves generating samples for a given TB (transform block) using previously reconstructed samples in the color channel under consideration. Intra prediction modes are signaled separately for luma and chroma channels, and the chroma channel intra prediction mode optionally depends on the luma channel intra prediction via "DM_CHROMA". Although the intra prediction mode is signaled at the PB (prediction block) level, the intra prediction process follows the residual quadtree hierarchical structure of the CU and is applied at the TB level. This allows the coding of one TV to influence the coding of the next TB within the CU, thus reducing the distance to the samples used as reference values.
[0023] HEVC includes 35 intra-prediction modes: DC mode, planar mode, and 33 directional or "angular" intra-prediction modes. The 33 angular intra-prediction modes are shown in FIG.
[0024] For PBs associated with chroma color channels, intra prediction modes are specified as planar, DC, horizontal, vertical, "DM_CHROMA" mode, or sometimes as diagonal mode "34".
[0025] Note that in chroma formats 4:2:2 and 4:2:0, a chroma PB may overlap with two or four (respectively) luma PBs, in which case the luma direction of DM_CHROMA is taken from the top-left of these luma PBs.
[0026] The DM_CHROMA mode indicates that the intra prediction mode of the luma color channel PB is applied to the chroma color channel PB. Since this is relatively common, the most-probable-mode coding scheme of intra_chroma_pred_mode is biased to select this mode.
[0027] 2.2 Description of the VVC (Versatile Video Coding) Algorithm 2.2.1 VVC Coding Architecture The Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015 to develop future coding technologies beyond HEVC. JVET meetings are currently held quarterly, and the new coding standard aims to achieve a 50% bitrate reduction compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the April 2018 JVET meeting, and the first version of the VVC test model (VTM) was published at that time. Continuous efforts to contribute to the VVC standardization have led to the adoption of new coding techniques in the VVC standard at each JVET meeting. The VVC working draft and test model VTM are therefore updated after each meeting. The VVC project is currently aiming for technical finalization (FDIS) at the July 2020 meeting.
[0028] Like most previous standards, VVC has a block-based hybrid coding architecture, joint inter-picture and intra-picture prediction, and transform coding with entropy coding. The picture partition structure divides the input video into blocks called coding tree units (CTUs). The CTUs are divided into coding units (CUs) using a quadtree with a nested multi-type tree structure, with leaf coding units (CUs) defining regions that share the same prediction mode (e.g., intra or inter). In this specification, the term "unit" defines a region of an image covering all color components, and the term "block" is used to define a region covering a specific color component (e.g., luma), which may differ in spatial location when considering chroma sampling formats such as 4:2:0.
[0029] 2.2.2 Dual / Individual Tree Partitions in VVC The luma and chroma components can have separate partition trees for I slices. The separate tree partitions are below the 64x64 block level, not at the CTU level. The VTM software has an SPS flag to control dual tree on and off.
[0030] 2.2.3 Intra Prediction in VVC 2.2.3.1 Intra Prediction Mode To capture any edge direction present in natural video, the number of directional intra modes in VTM4 has been expanded from 33 in HEVC to 65. The new directional modes not present in HEVC are indicated by the red dotted arrows in Figure 2, while the planar and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and to both luma and chroma intra prediction.
[0031] 2.2.3.2 Cross-component linear model prediction (CCLM) To reduce the redundancy between components, a cross-component linear model (CCLM) prediction mode is used in VTM4. For this, a chroma sample is predicted based on the reconstructed luma sample of the same CU using a linear model as follows:
number
number
[0032] The division operation to calculate the parameter α is performed by a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are expressed in exponential notation. For example, diff is approximated by a 4-bit significant part and an exponent. Thus, the table for 1 / diff is reduced to 16 elements for 16 mantissa values as follows:
number
[0033] This may have the advantage of reducing both the computational complexity and the memory size required to store the necessary tables.
[0034] In addition to the top and left templates being able to be used together to calculate the linear model coefficients, they can alternatively be used in other 2LM modes, referred to as LM_A and LM_L modes.
[0035] In LM_A mode, only the top template is used to calculate the linear model coefficients. To obtain more samples, the top template is extended to (W+H). In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H+W).
[0036] For non-square blocks, the top template is expanded to W+W and the left template is stored at H+H.
[0037] To match the chroma sample positions for 4:2:0 video sequences, two types of downsampling filters are applied to the luma samples to achieve two downsampling ratios in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type-0" and "type-2" content, respectively:
number
[0038] Note that when the upper reference line is at a CTU boundary, only one luma line (general purpose line buffer in intra prediction) is used to generate the downsampled luma samples.
[0039] This parameter calculation is performed as part of the decoding process and is not simply an encoder search operation. As a result, there is no syntax available to communicate the α and β values to the decoder.
[0040] In chroma intra mode coding, a total of eight intra modes are allowed for chroma intra mode coding. These modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). Chroma mode coding directly depends on the intra prediction mode of the corresponding luma block. Since separate block partition structures for luma and chroma components are possible in an I slice, one chroma block may correspond to multiple luma blocks. Therefore, the chroma DM mode directly inherits the intra prediction mode of the corresponding luma block that covers the center position of the current chroma block.
[0041] 2.2.3.2.1 Corresponding Modified Working Draft (JVET-N0271) The following specification is based on the modified working draft of JVET-M1001 and its adoption in JVET-N0271. The adopted changes from JVET-N0220 are shown in bold and underlined.
[0042] Syntax Table [Table 1] Semantics sps_cclm_enabled_flag equal to 0 specifies that cross-component linear model intra prediction from luma components to chroma components is disabled. sps_cclm_enabled_flag equal to 1 specifies that cross-component linear model intra prediction from luma components to chroma components is enabled.
[0043] Decryption process 8.4.4.2.8 Specification of INTRA_LT_CCLM, INTRA_L_CCLM and INTRA_T_CCLM intra prediction mode, The inputs to this process are: Intra prediction mode predModeIntra, The sample position ( xTbC, yTbC ) of the top-left sample of the current transform block relative to the top-left sample of the current picture, · Variable nTbW that specifies the transformation block width, · Variable nTbH that specifies the transformation block height, Neighborhood samples p[ x ][ y ], where x = -1, y = 0..2 * nTbH - 1 and x = 0..2 * nTbW - 1, y = - 1, The output of this process is the predicted samples predSamples[ x ][ y ], where x = 0..nTbW - 1, y = 0..nTbH - 1.
[0044] The current luma position ( xTbY, yTbY ) is derived as follows:
number
[0045] Clause 6.4.X [Ed.(BB): Neighboring blocks availability checking processtbd] The process of deriving the availability of the left neighboring sample for a block as specified in is invoked with the current chroma position (xCurr, yCurr) set equal to (xTbC, yTbC) and the neighboring chroma position (xTbC - 1, yTbC) as input, and the output is assigned to availL.
[0046] Clause 6.4.X [Ed.(BB): Neighboring blocks availability checking processtbd] The process of deriving the availability of the above neighboring sample for a block as specified in is invoked with the current chroma position (xCurr, yCurr) set equal to (xTbC, yTbC) and the neighboring chroma position (xTbC, yTbC - 1) as input, and the output is assigned to availT.
[0047] Clause 6.4.X [Ed.(BB): Neighboring blocks availability checking processtbd]The process of deriving the availability of the left neighboring sample for the block as specified in is invoked with the current chroma position (xCurr, yCurr) set equal to (xTbC, yTbC) and the neighboring chroma position (xTbC - 1, yTbC - 1) as input, and the output is assigned to availTL.
[0048] The number of available top-right neighboring chroma samples, numTopRight, is derived as follows:
[0049] The variable numTopRight is set equal to 0 and availTR is set equal to TRUE.
[0050] When -predModeIntra is equal to INTRA_T_CCLM, the following applies for x = nTbW..2 * nTbW - 1 until availTR is equal to FALSE or x is equal to 2 * nTbW - 1:
[0051] Clause 6.4.X [Ed.(BB): Neighboring blocks availability checking processtbd] The availability derivation process for the block as specified in is invoked with the current chroma position (xCurr, yCurr) set equal to (xTbC, yTbC) and the neighbor chroma position (xTbC + x, yTbC - 1) as input, and the output is assigned to availableTR.
[0052] · When availableTR is equal to TRUE, numTopRight is incremented by 1.
[0053] The number of available bottom-left neighboring chroma samples, numLeftBelow, is derived as follows:
[0054] The variable numLeftBelow is set equal to 0 and availLB is set equal to TRUE.
[0055] When -predModeIntra is equal to INTRA_L_CCLM, the following applies for y = nTbH..2 * nTbH - 1, until availLB is equal to FALSE or y is equal to 2 * nTbH - 1.
[0056] Clause 6.4.X [Ed.(BB): Neighboring blocks availability checking processtbd] The availability derivation process for the block as specified in is invoked with the current chroma position (xCurr, yCurr) set equal to (xTbC, yTbC) and the neighboring chroma position (xTbC - 1, yTbC + y) as input, and the output is assigned to availableLB.
[0057] When availableLB is equal to TRUE, numLeftBelow is incremented by 1.
[0058] The number of available neighboring chroma samples above and to the right, numTopSamp, and the number of available neighboring chroma samples to the left and to the left, nLeftSamp, are derived as follows:
[0059] If predModeIntra is equal to INTRA_LT_CCLM, the following applies:
number
number
number
number
number
number
[0060] 67 intra modes with wide-angle mode extension - 4-tap interpolation filter depending on block size and mode Position-dependent intra prediction combination (PDPC) Cross-component linear model intra-prediction Multi-reference line intra-prediction Intra-subpartition 2.2.4 Inter Prediction in VVC 2.2.4.1 Combined Inter and Intra Prediction (CIIP) In VTM4, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., CU width × CU height is 64 or more), an additional flag is signaled to indicate whether combined inter / intra prediction (CIIP) mode is currently applied to the CU.
[0061] To form the CIIP prediction, the intra-prediction mode is first derived from two additional syntax elements. Up to four possible intra-prediction modes are available: DC, planar, horizontal, or vertical. Next, the inter-prediction and intra-prediction signals are derived using normal intra- and inter-decoding processes. Finally, a weighted average of the inter- and intra-prediction signals is performed to obtain the CIIP prediction.
[0062] 2.2.4.2 Various Inter Prediction Modes VTM4 includes many inter-coding tools that differ from HEVC, for example, the following features are included in VCC Test Model 3 on top of the block tree structure: Affine motion inter-prediction Sub-block based temporal motion vector prediction Adaptive motion vector resolution 8x8 block-based motion compression for temporal motion estimation High-precision (1 / 16 pel) motion vector storage and motion compensation with an 8-tap interpolation filter for the luma component and a 4-tap interpolation filter for the chrominance component Triangular partition Combined intra and inter prediction Merge with MVD (MMVD) Symmetric MVD coding Bidirectional optical flow Decoder-side motion vector refinement Bi-predictive weighted average
[0063] 2.2.5 In-loop filter There are three in-loop filters in total in VTM4. In addition to the two loop filters in HEVC, the deblocking filter and the sample adaptive offset filter (SAO), an adaptive loop filter (ALF) is applied in VTM4. The filtering process order in VTM4 is the deblocking filter, SAO, and ALF.
[0064] In VTM4, the SAO and deblocking filtering processes are almost the same as those in HEVC.
[0065] VTM4 adds a new process called luma mapping with chroma scaling (previously known as adaptive in-loop reshaper). This new process is performed before deblocking.
[0066] 2.2.6 Luma mapping with chroma scaling (LMCS), also known as in-loop reshaping In VTM4, a coding tool called luma mapping with chroma scaling (LMCS) is added as a new processing block before the loop filter. LMCS has two main components: 1) in-loop mapping of the luma component based on an adaptive piecewise linear model, and 2) for the chroma component, luma-dependent chroma residual scaling is applied. Figure 4 shows the LMCS architecture from the decoder's perspective. In Figure 4, the lightly shaded blocks indicate where processing is applied in the mapped domain, including inverse quantization, inverse transform, luma intra prediction, and summation of the luma prediction with the luma residual. In Figure 4, the white blocks indicate where processing is applied in the original (i.e., unmapped) domain, including deblocking, loop filters such as ALF and SAO, motion-compensated prediction, chroma intra prediction, summation of the chroma prediction with the chroma residual, and storing the decoded picture as a reference picture. In Figure 4, the darkly shaded block is the new LMCS functional block, which includes forward and backward mapping of the luma signal and luma-dependent chroma scaling processing. Like most other tools in VVC, LMCS can be enabled / disabled at the sequence level using the SPS flag.
[0067] 2.2.6.1 Luma mapping with a piecewise linear model In-loop mapping of the luma component adjusts the dynamic range of the input signal by redistributing codewords across the dynamic range, improving compression efficiency. Luma mapping uses a forward mapping function FwdMap and a corresponding inverse mapping function InvMap. The FwdMap function is signaled using a piecewise linear model with 16 equal pieces. The InvMap function does not need to be signaled, but is instead derived from the FwdMap function.
[0068] The luma mapping model is signaled at the tile group level. A presence flag is signaled first. If a luma mapping model is currently present in the tile group, the corresponding piecewise linear model parameters are signaled. The piecewise linear model partitions the dynamic range of the input signal into 16 equal partitions, and for each partition, its linear mapping parameters are expressed using the number of codewords assigned to that partition. Take a 10-bit input as an example. Each of the 16 partitions has 64 codewords assigned to it by default. The signaled number of codewords is used to calculate the scaling factor and adjust the mapping function for that partition accordingly. At the tile group level, another LMCS enable flag is signaled to indicate whether LMCS processing as shown in Figure 4 is currently applied to the tile group.
[0069] Each ith, i=0...15, piece of the FwdMap piecewise linear model is defined by two input pivot points InputPivot[] and two output (mapped) pivot points MappedPivot[].
[0070] InputPivot[] and MappedPivot[] are calculated as follows (assuming 10-bit video):
number
number
number
[0071] The luma mapping process (forward and / or backward mapping) can be performed using a look-up-table (LUT) or using on-the-fly calculations. If a LUT is used, the FwdMapLUT and InvMapLUT can be pre-computed and pre-stored for use at the tile group level, and the backward mapping can be simply performed as follows, respectively:
number
[0072] Alternatively, on-the-fly calculations may be used. Take the forward mapping function FwdMap as an example. To find the partition to which a luma sample belongs, the sample value is right-shifted by 6 bits (which corresponds to 16 equal partitions). Then, the linear model parameters of that partition are looked up and applied on-the-fly to calculate the mapped luma value. If i is the partition index, a1 and a2 are respectively:
number
number
number
[0073] The InvMap function can be computed on the fly in a similar way, but when determining which partition a sample value belongs to, a condition check needs to be applied instead of a simple right bit shift, since the partitions in the mapped domain are not of equal size.
[0074] 2.2.6.2 Luma-dependent chroma residual scaling Chroma residual scaling is designed to compensate for the interaction between a luma signal and its corresponding chroma signal. Whether chroma residual scaling is enabled is also signaled at the tile group level. If luma mapping is enabled and dual tree partitioning (also known as individual chroma trees) is not currently applied to the tile group, an additional flag is signaled to indicate whether luma-dependent chroma residual scaling is enabled. When dual tree partitioning is not used in the current tile group or when luma mapping is not used, luma-dependent chroma residual scaling is disabled. Furthermore, luma-dependent chroma residual scaling is always disabled for chroma blocks with an area of 4 or less.
[0075] The chroma residual scaling depends on the average value of the corresponding luma prediction block (for both intra- and inter-coding blocks). avgY' denotes the average of the luma prediction block. C ScaleInv The value of is calculated in the following steps:
number
[0076] If the current block is coded in intra, CIIP, or intra block copy (IBC, also known as current picture referencing or CPR) mode, avgY' is calculated as the average of the intra, CIIP, or IBC predicted luma values; otherwise, avgY' is calculated as the average of the forward-mapped inter-predicted luma values (Y' in Figure 4). pred ). Unlike luma mapping, which is performed per sample, C ScaleInve is a constant value for the entire chroma block. ScaleInve Then, chroma residual scaling is applied as follows:
number
[0077] 2.2.6.3 Adoption in JVET-N0220 and corresponding working draft in JVET-M1001_v7 The following specification is based on the modified working draft of JVET-M1001 and its adoption in JVET-N0220. Changes in the adopted JVET-N0220 are shown in bold and underlined.
[0078] Syntax Table [Table 2] [Table 3] [Table 4]
[0079] Semantics 7.4.3.1 Sequence parameter set RBSP In the mandix, sps_lmcs_enabled_flag equal to 1 specifies that luma mapping with chroma scaling is used in CVS. sps_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is not used in CVS.
[0080] tile_group_lmcs_model_present_flag equal to 1 specifies the presence of lmcs_data() in the tile group header. tile_group_lmcs_model_present_flag equal to 0 specifies the absence of lmcs_data() in the tile group header. When tile_group_lmcs_model_present_flag is not present, it is inferred to be equal to 0.
[0081] tile_group_lmcs_enabled_flag equal to 1 specifies that luma mapping with chroma scaling is enabled for the current tile group. tile_group_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is not enabled for the current tile group. When tile_group_lmcs_enabled_flag is not present, it is inferred to be equal to 0.
[0082] tile_group_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for the current tile group. tile_group_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current tile group. When tile_group_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.
[0083] 7.4.5.4 Data Semantics of Luma Mapping with Chroma Scaling lmcs_min_bin_idx specifies the minimum bin index used in the luma mapping composition process by chroma scaling. The value of lmcs_min_bin_idx should be in the range 0 to 15 inclusive.
[0084] lmcs_delta_max_bin_idx specifies the delta value between 15 and the maximum bin index LmcsMaxBinIdx used in the luma mapping configuration process with chroma scaling. The value of lmcs_delta_max_bin_idx should be in the range of 0 to 15, inclusive. The value of LmcsMaxBinIdx is set equal to 15 - lmcs_delta_max_bin_idx. The value of LmcsMaxBinIdx should be greater than or equal to lmcs_min_bin_idx.
[0085] lmcs_delta_cw_prec_minus1 plus 1 specifies the number of bits used for the representation of the syntax lmcs_delta_abs_cw[ i ]. The value of lmcs_delta_cw_prec_minus1 should be in the range 0 to BitDepthY - 2, inclusive.
[0086] lmcs_delta_abs_cw[ i ] specifies the absolute delta codeword value for the ith bin.
[0087] lmcs_delta_sign_cw_flag[ i ] specifies the sign of the variable lmcsDeltaCW[ i ] as follows:
number
[0088] 3. Shortcomings of existing implementations The current design of LMCS / CCLM may have the following problems: (1) In the LMCS coding tool, the chroma residual scaling factor is derived by the average value of the co-located luma prediction block, which causes a delay for processing chroma samples in the LMCS chroma residual scaling. a) In the case of a single / shared tree, the delay is caused by (a) waiting for all prediction samples for the entire available luma block, and (b) averaging all luma prediction samples obtained by (a). b) In the case of dual / separate trees, the delay is even worse because separate block partition structures are enabled for luma and chroma components in an I slice. Thus, one chroma block may correspond to multiple luma blocks, and one 4x4 chroma block may correspond to a 64x64 luma block. Therefore, the worst case is that the chroma residual scaling factor for the current 4x4 chroma block needs to wait until all prediction samples for the entire 64x64 luma block are available. In other words, the delay problem in dual / separate trees is much more severe. (2) In CCLM coding tools, CCLM model calculation for intra-chroma prediction depends on the left and top reference samples of both luma and chroma blocks, and CCLM prediction of a chroma block depends on the co-located luma reconstructed sample of the same CU, which may cause high latency in dual / separate trees. In the case of dual / separate trees, one 4x4 chroma block may correspond to a 64x64 luma block. Therefore, the worst case scenario is that CCLM processing of the current chroma block must wait until the entire corresponding 64x64 luma block is reconstructed. This latency issue is similar to LMCS chroma scaling in dual / separate trees.
[0089] 4. Exemplary Techniques and Embodiments To address the problem, we propose several methods to remove / reduce / limit cross-component dependencies in luma-dependent chroma residual scaling, CCLM, and other coding tools that rely on information from different color components.
[0090] The detailed embodiments described below should be considered as examples to illustrate the general concept. These embodiments should not be construed in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0091] It should be noted that although the bullets mentioned below explicitly refer to LMCS / CCLM, the methods may also be applicable to other coding tools that rely on information from different color components. Furthermore, the terms "luma" and "chroma" mentioned below may be replaced by "first color component" and "second color component", respectively, such as "G component" and "B / R component" in the RGB color format.
[0092] In the following discussion, the definition of "collocated sample / block" is consistent with the definition of collocated sample / block in VVC Working Draft JVET-M1001. More specifically, in 4:2:0 color format, the top-left sample of a chroma block is assumed to be at position (xTbC,yTbC), and therefore the top-left sample of the co-located luma block position (xTbY,yTbY) is derived as follows:
number
[0093] In the following discussion, a "corresponding block" may have a different location than the current block. For example, there may be a motion shift between the current block and its corresponding block in the reference frame. As shown in Figure 6, if the current block is located at (x, y) in the current frame and has a motion vector (mv x ,mv y ), the corresponding block of the current block is (x+mv x ,y+mv y) in the current frame. Also, in an IBC coding block, the co-located luma block (pointed to by a zero vector) and the corresponding luma block (pointed to by a non-zero BV) may be located at different locations in the current frame. In another example, when the partition of a luma block (in the dual-tree partition of an I slice) does not align with the partition of a chroma block, the co-located luma block of the current chroma block may belong to a larger luma block depending on the partition size of the overlapping luma coding block that covers the upper-left sample of the co-located luma block. As shown in Figure 5, assuming that the bold rectangle indicates the partition of a block, a 64x64 luma block is first divided by BT, and then the right part of the 64x64 luma block is further divided by TT, resulting in three luma blocks with sizes equal to 32x16, 32x32, and 32x16, respectively. Therefore, the top-left sample (x=32, y=32) of the co-located luma block of the current chroma block belongs to the central 32x32 luma block of the TT partition. In this case, we call the corresponding luma block that covers the top-left sample of the co-located luma block the "corresponding luma block." Therefore, in this example, the top-left sample of the corresponding luma block is located at (x=32, y=16).
[0094] Hereinafter, DMVD (decoder-side motion vector derivation) is used to represent BDOF (also known as BIO) or / and DMVR (decode-side motion vector refinement) or / and FRUC (frame rate up-conversion) or / and other methods of refining motion vectors or / and predicted sample values on the decoder side.
[0095] Removal of LMCS chroma scaling delay and CCLM model calculation (1) For inter-coding blocks, it is proposed that one or more reference samples of the current block in the reference frame may be used to derive the chroma residual scaling factor in LMCS mode. a) In one example, the reference luma samples may be used directly to derive the chroma residual scaling factors. i. Alternatively, interpolation may be applied to the reference samples first, and the interpolated samples used to derive the chroma residual scaling factors. ii. Alternatively, reference samples in a different reference frame may be utilized to derive the final reference samples used for chroma residual scaling factor derivation. 1) In one example, for bi-predictive coding blocks, the above method may be applied. iii. In one example, the intensities of the reference samples may be transformed into a reshaping domain before being used to derive the chroma residual scaling factor. iv. In one example, a linear combination of the reference samples may be used to derive the chroma residual scaling factor. 1) For example, a×S+b may be used to derive the chroma residual scaling factor, where S is a reference sample and a and b are parameters. In one example, a and b may be derived by Localized Illumination Compensation (LIC). b) In one example, the location of the reference luma sample in the reference frame may depend on the motion vector of the current block. In one example, the reference sample belongs to a reference luma block in the reference picture that has the same width and height as the current luma block. The position of the reference luma sample in the reference picture may be calculated as the position of its corresponding luma sample in the current picture, plus a motion vector. ii. In one example, the location of the reference luma sample is derived from the location of the top-left (or center, or bottom-right) sample of the current luma block and the motion vector of the current block, and may be referred to as the corresponding luma sample in the reference frame. 1) In one example, an integer motion vector may be used to derive a corresponding luma sample in a reference frame. In one example, a motion vector associated with a block may be rounded toward or away from zero to derive the integer motion vector. 2) Alternatively, fractional motion vectors may be used to derive corresponding luma samples in a reference frame, resulting in an interpolation process being required to derive the fractional reference samples. iii. Alternatively, the position of the reference luma sample may be derived by the position of the top-left (or center, or bottom-right) sample of the current luma block. iv. Alternatively, multiple corresponding luma samples at some predetermined positions in the reference frame may be chosen to calculate the chroma residual scaling factor. c) In one example, the median or average value of multiple reference luma samples may be used to derive the chroma residual scaling factor. d) In one example, a reference luma sample in a given reference frame may be used to derive a chroma residual scaling factor. i. In one example, the given reference frame may be the one in reference picture list 0 with a reference index equal to 0. ii. Alternatively, the reference index and / or reference picture list for a given reference frame may be signaled at the sequence / picture / tile group / slice / tile / CTU row / video unit level. iii. Alternatively, reference luma samples in multiple reference frames may be derived and an average or weighted average may be utilized to obtain the chroma residual scaling factor. (2) It is proposed that whether and how to calculate chroma residual scaling factors from luma samples in LMCS mode may depend on whether the current block is bi-predictive. a) In one example, the chroma residual scaling factor is derived separately for each prediction direction. (3) It is proposed that whether and how to extract chroma residual scaling factors from luma samples in LMCS mode may depend on whether the current block is subject to sub-block-based prediction. a) In one example, the sub-block based prediction is an affine prediction. b) In one example, the sub-block based prediction is Alternative Temporal Motion Vector Prediction (ATMVP). c) In one example, the chroma residual scaling factors are derived for each sub-block separately. d) In one example, the chroma residual scaling factor is derived for the entire block, even if predicted by sub-blocks. i. In one example, the motion vector of one selected sub-block (eg, the top-left sub-block) may be used to identify a reference sample for the current block, as indicated by bullet 1. (4) It is proposed that the luma predicted value used to derive the chroma residual scaling factor may be an intermediate luma predicted value instead of the final luma predicted value. a) In one example, luma prediction values before Bi-Directional Optical Flow (BDOF, also known as BIO) processing may be used to derive chroma residual scaling factors. b) In one example, the luma prediction value before decoder-side motion vector refinement (DMVR) processing may be used to derive the chroma residual scaling factor. c) In one example, the luma prediction value before LIC processing may be used to derive the chroma residual scaling factor. d) In one example, the luma prediction value before Prediction Refinement Optical Flow (PROF) processing as proposed in JVET-N0236 may be used to derive the chroma residual scaling factor. (5) The intermediate motion vector may be used to identify a reference sample. a) In one example, the motion vectors before processing of BDOF or / and DMVR or / and other DMVD methods may be used to identify the reference samples. b) In one example, the motion vectors before Prediction Refinement Optical Flow (PROF) processing as proposed in JVET-N0236 may be used to identify the reference samples. (6) The above method may be applicable when the current block is coded in inter mode. (7) For IBC coded blocks, it is proposed that one or more reference samples in a reference block of the current frame may be used to derive a chroma residual scaling factor in LMCS mode. When a block is IBC coded, if the reference picture is set as the current picture, the term "motion vector" may also be called "block vector." a) In one example, the reference sample belongs to a reference block in the current picture that has the same width and height as the current block. The position of the reference sample may be calculated as the position of its corresponding luma sample plus a motion vector. b) In one example, the position of the reference luma sample may be derived by the position of the top-left (or center, or bottom-right) sample of the current luma block, plus a motion vector. c) Alternatively, the position of the reference luma sample may be derived by the position of the top-left (or center, or bottom-right) sample of the current luma block, plus the block vector of the current block. d) Alternatively, multiple corresponding luma samples at some predetermined positions within the reference region of the current luma block may be chosen to calculate the chroma residual scaling factor. e) In one example, multiple corresponding luma samples may be calculated by a function to derive a chroma residual scaling factor. i. For example, the median or average of multiple corresponding luma samples may be calculated to derive a chroma residual scaling factor. f) In one example, the intensities of the reference samples may be transformed into a reshaping domain before being used to derive the chroma residual scaling factors. i. Alternatively, the intensities of the reference samples may be transformed back to the original domain before being used to derive the chroma residual scaling factor. (8) It is proposed that one or more predicted / reconstructed samples located at the identified position of the current luma block in the current frame may be used to derive a chroma residual scaling factor for the current chroma block in LMCS mode. a) In one example, if the current block is inter-coded, the luma predicted (or reconstructed) sample located in the center of the current block may be chosen to derive the chroma residual scaling factor. b) In one example, the average value of the first M×N luma predicted (or reconstructed) samples may be taken to derive the chroma residual scaling factor, where M×N may be smaller than the co-located luma block size width×height. (9) It is proposed that the whole or part of the procedure used to calculate the CCLM model may be used for deriving the chroma residual scaling factor of the current chroma block in LMCS mode. a) In one example, reference samples located at identified positions of neighboring luma samples of the co-located luma block in the CCLM model parameter derivation process may be utilized to derive chroma residual scaling factors. In one example, the reference samples may be used directly. ii. Alternatively, downsampling may be applied to the reference samples and the downsampled reference samples may be applied. b) In one example, K of the S reference samples selected for CCLM model calculation may be used for chroma residual scaling factor derivation in LMCS mode, where K is equal to 1 and S is equal to 4, for example. c) In one example, the mean / min / max of reference samples of the co-located luma block in CCLM mode may be used for chroma residual scaling factor derivation in LMCS mode. (10) How samples are selected for the derivation of the chroma residual scaling factor may depend on the coding information of the current block. a) The coding information may include QP, coding mode, POC, intra prediction mode, motion information, etc. b) In one example, the method of selecting samples may be different for IBC coded or non-IBC coded blocks. c) In one example, the method of selecting samples may differ based on reference picture information such as the POC difference between the reference picture and the current picture. (11) The chroma residual scaling factor and / or model calculation of CCLM may depend on neighboring samples of the corresponding luma block that cover the top-left sample of the co-located luma block. a) A "corresponding luma coding block" may be defined as a coding block that covers the upper-left position of the co-located luma coding block. i. Figure 5 shows an example in which the CTU partition of a chroma component may differ from the CTU partition of a luma component for an intra-coding chroma block in the dual-tree case. First, a "corresponding luma coding block" that covers the upper-left sample of the co-located luma block of the current chroma block is searched for. Next, the upper-left sample of the "corresponding luma coding block" can be derived using block size information of the "corresponding luma coding block," and the upper-left luma sample of the "corresponding luma coding block" that covers the upper-left sample of the co-located luma block is identified as being located at (x=32, y=16). b) In one example, reconstructed samples that are not within the "corresponding luma coding block" may be used to derive the chroma residual scaling factor and / or model calculation for CCLM. i. In one example, reconstructed samples adjacent to the "corresponding luma coding block" may be used to derive the chroma residual scaling factor and / or model calculation for CCLM. 1) In one example, N samples located in the left-neighboring column and / or the top-neighboring row of the "corresponding luma coding block" may be used to derive the chroma residual scaling factor and / or model calculation for CCLM, where N=1...2W+2H, where W and H are the width and height of the "corresponding luma coding block." a) If the top-left sample of the "corresponding luma coding block" is at (xCb, yCb), then in one example, the top-neighboring luma sample may be located at (xCb+W / 2, yCb-1) or (xCb-1, yCb-1). In an alternative example, the left-neighboring luma sample may be located at (xCb+W-1, yCb-1). b) In one example, the locations of the neighboring samples may be fixed and / or in a predetermined checking order. 2) In one example, one neighboring sample out of N may be selected to derive the chroma residual scaling factor and / or model calculation for CCLM. Given N=3 and a check order of three neighboring samples as (xCb-1, yCb-H-1), (xCb+W / 2, yCb-1), (xCb-1, yCb-1), the first available neighboring sample in the check list may be selected to derive the chroma residual scaling factor. 3) In one example, the median or average value of N samples located in the left-neighboring column and / or the top-neighboring row of the "corresponding luma coding block" may be used to derive the chroma residual scaling factor and / or model calculation for CCLM, where N=1...2W+2H, where W and H are the width and height of the "corresponding luma coding block". c) In one example, whether to perform chroma residual scaling may depend on the "available" neighboring samples of the corresponding luma block. i. In one example, the "availability" of neighboring samples may depend on the coding mode of the current block / sub-block or / and the coding mode of the neighboring samples. 1) In one example, for a block coded in inter mode, neighboring samples coded in intra mode or / and IBC mode or / and CIIP mode or / and LIC mode may be considered "unavailable". 2) In one example, for blocks coded in inter mode, neighboring samples utilizing a diffusion filter or / and a bilateral filter or / and a Hadamard transform filter may be considered "unavailable." ii. In one example, the "availability" of neighboring samples may depend on the width and / or height of the current picture / tile / tile group / VPDU / slice. 1) In one example, if a neighboring block is located outside the current picture, it is treated as "unavailable." iii. In one example, when there are no "available" neighboring samples, chroma residual scaling may not be allowed. iv. In one example, when the number of "available" neighboring samples is less than K (K>=1), chroma residual scaling may not be allowed. v. Alternatively, unavailable neighboring samples may be filled, padded, or substituted with a default fixed value, so that chroma residual scaling is always applied. 1) In one example, if no neighboring samples are available, it may be filled with 1<<(bitDepth-1), where bitDepth specifies the bit depth of the luma / chroma component samples. 2) Alternatively, if a neighboring sample is not available, it may be filled by padding from surrounding samples located in the left / right / top / bottom neighborhood. 3) Alternatively, if a neighboring sample is not available, it may be substituted by the first available neighboring sample in a predetermined check order. d) In one example, the neighboring filtered / mapped reconstructed samples of the "corresponding luma coding block" may be used to derive the chroma residual scaling factor and / or model calculation for CCLM. i. In one example, the filtering / mapping process may include post-filtering such as reference smoothing filters for intra blocks, bilateral filters, Hadamard transform based filters, forward mapping of the reshaper domain, etc.
[0096] Constraints on whether chroma residual scaling and / or CCLM is applied (12) It is proposed that whether chroma residual scaling or CCLM is applied may depend on the partition of the corresponding and / or co-located luma block. a) In one example, whether to enable or disable a tool with cross-component information may depend on the number of CUs / PUs / TUs in the co-located luma (eg, Y or G component) block. i. In one example, if the number of CUs / PUs / TUs in a co-located luma (e.g., Y or G component) block exceeds a threshold, such tools may be disabled. ii. Alternatively, enabling or disabling tools with cross-component information may depend on the depth of the partition tree. 1) In one example, if the maximum (or minimum or average or other variable) quadtree depth of CUs in the co-located luma block exceeds a threshold, such tools may be disabled. 2) In one example, if the maximum (or minimum or average or other variable) BT and / or TT depth of CUs in the co-located luma block exceeds a threshold, such tools may be disabled. iii. Alternatively, further enabling or disabling of tools with cross-component information may depend on the block dimensions of the chroma blocks. iv. Alternatively, further enabling or disabling of tools with cross-component information may depend on whether the co-located luma spans multiple VPDUs / predetermined region size. v. The thresholds in the above discussion may be fixed numbers, or may be signaled, or may depend on the standard profile / level / tier. b) In one example, if the co-located luma block of the current chroma block is divided by multiple partitions (eg, FIG. 7), chroma residual scaling and / or CCLM may be prohibited. i. Alternatively, if the co-located luma block of the current chroma block is not split (eg, within one CU / TU / PU), chroma residual scaling and / or CCLM may be applied. c) In one example, if the co-located luma block of the current chroma block contains more than M CUs / PUs / TUs, chroma residual scaling and / or CCLM may be prohibited. In one example, M may be an integer greater than 1. ii. In one example, M may depend on whether it is a CCLM or chroma residual scaling process. iii. M may be a fixed number, or may be signaled, or may depend on the standard profile / level / tier. d) The above-mentioned CUs in the co-located luma block may be interpreted as all CUs in the co-located luma block. Alternatively, the CUs in the co-located luma block may be interpreted as some CUs in the co-located luma block, such as CUs along the boundary of the co-located luma block. e) The above-mentioned CUs within the co-located luma block may be interpreted as sub-CUs or sub-blocks. i. For example, sub-CUs or sub-blocks may be used in ATMVP. ii. For example, a sub-CU or sub-block may be used in affine prediction. iii. For example, the sub-CU or sub-block may be used in Intra Sub-Partitions (ISP) mode. f) In one example, if the CU / PU / TU covering the top-left luma sample of the co-located luma block is larger than a predetermined luma block size, chroma residual scaling and / or CCLM may be prohibited. i. An example is shown in Figure 8, where a co-located luma block is 32x32 but within a corresponding luma block having a size equal to 64x64, and the given luma block size is 32x64, chroma residual scaling and / or CCLM are prohibited in this case. ii. Alternatively, if the co-located luma block of the current chroma block is not split and the corresponding luma block covering the top-left luma sample of the co-located luma block is completely contained within a predetermined bounding box, chroma residual scaling and / or CCLM may be applied. As shown in Figure 9, the bounding box may have a width W and a height H, and may be defined as a rectangle denoted by W x H. Here, the corresponding luma block has a width of 32 and a height of 64, and the bounding box has a width of 40 and a height of 70. 1) In one example, the size W×H of the bounding box may be defined according to the CTU width and / or height, or according to the CU width and / or height, or according to any arbitrary value. g) In one example, if the co-located luma block of the current chroma block is divided into multiple partitions, only predicted samples (or reconstructed samples) within a given partition of the co-located luma block are used to derive the chroma residual scaling factor in LMCS mode. i. In one example, the average of all prediction samples (or reconstructed samples) in the first partition of the co-located luma block is used to derive the chroma residual scaling factor in LMCS mode. ii. Alternatively, the top-left predicted sample (or reconstructed sample) in the first partition of the co-located luma block is used to derive the chroma residual scaling factor in LMCS mode. iii. Alternatively, the central prediction sample (or reconstructed sample) in the first partition of the co-located luma block is used to derive the chroma residual scaling factor in LMCS mode. h) It is proposed that whether and how cross-component tools such as CCLM and LMCS are applied may depend on the coding mode of one or more luma CUs that cover at least one sample of the co-located luma block. i. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded in affine mode, the cross-component tool is disabled. ii. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded bi-predictively, the cross-component tool is disabled. iii. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded with BDOF, the cross-component tool is disabled. iv. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded with DMVR, the cross-component tool is disabled. v. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded in the matrix affine prediction mode proposed in JVET-N0217, the cross-component tool is disabled. vi. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded in inter mode, the cross-component tool is disabled. vii. For example, if one or more luma CUs covering at least one sample of the co-located luma block are coded in ISP mode, the cross-component tool is disabled. viii. In one example, "one or more luma CUs covering at least one sample of a co-located luma block" may refer to the corresponding luma block. i) When CCLM / LMCS is prohibited, signaling of instructions to use CCLM / LMCS may be skipped. j) In this disclosure, CCLM may refer to any deformation mode of CCLM, including LM mode, LM-T mode, and LM-L mode. (13) It is proposed that whether and how to apply cross-component tools such as CCLM and LMCS may be performed on portions of chroma blocks. a) In one example, whether and how cross-component tools such as CCLM and LMCS are applied is at the level of the chroma sub-blocks. i. In one example, a chroma sub-block is defined as a 2x2 or 4x4 block within a chroma CU. ii. In one example, for a chroma sub-block, CCLM may be applied when the corresponding luma coding block of the current chroma CU covers all samples of the corresponding block of the sub-block. iii. In one example, for a chroma sub-block, CCLM is not applied when not all samples of the corresponding block are covered by the corresponding luma coding block of the current chroma CU. iv. In one example, CCLM or LMCS parameters are derived for each chroma sub-block when treating the sub-block as a chroma CU. v. In one example, when CCLM or LMCS is applied on a chroma sub-block, samples of the co-located block may be used.
[0097] Applicability of Chroma Residual Scaling in LMCS Mode (14) It is proposed that whether luma-dependent chroma residual scaling is applicable may be signaled at other syntax levels in addition to the tile group header as specified in JVET-M1001. a) For example, chroma_residual_scale_flag may be signaled at the sequence level (e.g., in an SPS), picture level (e.g., in a PPS or picture header), slice level (e.g., in a slice header), tile level, CTU row level, CTU level, or CU level. chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for CUs below the signaled syntax level. chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for CUs below the signaled syntax level. When chroma_residual_scale_flag is not present, it is inferred to be equal to 0. b) In one example, if chroma residual scaling is not limited at the partition node level, chroma_residual_scale_flag may not be signaled and may be inferred to be 0 for the CUs covered by the partition node. In one example, the partition node may be a CTU (a CTU is treated as the root node of a quadtree partition). c) In one example, if chroma residual scaling is limited for chroma block sizes of 32x32 or less, chroma_residual_scale_flag may not be signaled and may be inferred to be 0 for chroma block sizes of 32x32 or less.
[0098] Applicability of CCLM mode (15) It is proposed that whether CCLM mode is applicable may be signaled at other syntax levels in addition to the SPS level as specified in JVET-M1001. a) For example, it may be signaled at the picture level (e.g., in the PPS or picture header), the slice level (e.g., in the slice header), the tile group level (e.g., in the tile group header), the tile level, the CTU row level, the CTU level, or the CU level. b) In one example, cclm_flag may not be signaled and may be inferred to be 0 if CCLM is not applicable. i. In one example, if chroma residual scaling is restricted for chroma block sizes of 8x8 or less, cclm_flag may not be signaled and may be inferred to be 0 for chroma block sizes of 8x8 or less.
[0099] Unifying the derivation of chroma residual scaling factors for intra and inter modes (16) The chroma residual scaling factor may be derived after encoding / decoding a luma block and may be stored and used for subsequent coding blocks. a) In one example, a particular predicted sample or / and intermediate predicted sample or / and reconstructed sample within a luma block or / and reconstructed sample before loop filtering (e.g., before being processed by a deblocking filter or / and SAO filter or / and bilateral filter or / and Hadamard transform filter or / and ALF filter) may be used to derive a chroma residual scaling factor. i. For example, some samples in the bottom row or / and right column of the luma block may be used for deriving the chroma residual scaling factor. b) In the case of a single tree, when encoding a block coded in intra mode or / and IBC mode or / and inter mode, the derived chroma residual scaling factors of neighboring blocks may be used to derive the scaling factor of the current block. i. In one example, certain neighboring blocks may be checked in order, and the first available chroma residual scaling factor may be used for the current block. ii. In one example, a particular neighboring block may be checked in turn, and a scaling factor may be derived based on the first K available neighboring chroma residual scaling factors. iii. In one example, for a block coded in inter mode or / and CIIP mode, if a neighboring block is coded in intra mode or / and IBC mode or / and CIIP mode, the chroma residual scaling factor of the neighboring block may be considered "unavailable." iv. In one example, neighboring blocks may be checked in order from left (or top left) to top (or top right). 1) Alternatively, the neighboring blocks may be checked in order from top (or top right) to left (or top left). c) In the case of a separate tree, when encoding a chroma block, the corresponding luma block may be first identified, and then the derived chroma residual scaling factors of its neighboring blocks (e.g., of the corresponding luma block) may be used to derive a scaling factor for the current block. i. In one example, certain neighboring blocks may be checked in order, and the first available chroma residual scaling factor may be used for the current block. ii. In one example, a particular neighboring block may be checked in turn, and a scaling factor may be derived based on the first K available neighboring chroma residual scaling factors. d) Neighboring blocks may be checked in a predetermined order. i. In one example, the neighboring blocks are checked in order from left (or top left) to top (or top right). ii. In one example, neighboring blocks may be checked in order from top (or top right) to left (or top left). iii. In one example, the neighboring blocks may be checked in the following order: bottom left → left → top right → top → top left. iv. In one example, neighboring blocks may be checked in the following order: left → top → top right → bottom left → top left. e) In one example, whether to apply chroma residual scaling may depend on the "availability" of neighboring blocks. i. In one example, when there are no "available" neighboring blocks, chroma residual scaling may not be allowed. ii. In one example, when the number of "available" neighboring blocks is less than K (K>=1), chroma residual scaling may not be allowed. iii. Alternatively, when there are no "available" neighboring blocks, the chroma residual scaling factor may be derived by a default value. 1) In one example, a default value of 1<<(BitDepth-1) may be used to derive the chroma residual scaling factor.
[0100] 5. Exemplary Implementations of the Disclosed Techniques FIG. 10 is a block diagram of a video processing device 1000. The device 1000 may be used to implement one or more of the methods described herein. The device 1000 may be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 1000 may include one or more processors 1002, one or more memories 1004, and video processing hardware 1006. The processor 1002 may be configured to implement one or more of the methods described herein (including, but not limited to, methods 800 and 900). The memory(s) 1004 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 106 is a hardware circuit that may be used to implement some of the techniques described herein.
[0101] In some embodiments, the video coding method may be performed using equipment implemented on a hardware platform such as that described with respect to FIG.
[0102] 11 shows a flowchart of an example method 1100 of deriving a linear model for cross-component prediction in accordance with the disclosed techniques. The method 1100 includes, at step 1110, performing a conversion between a current video block and a bitstream representation of the current video block, where during the conversion, a second set of color component values for the current video block are derived from a first set of color component values included in one or more reference frames, the first set of color component values being usable in a linear model of a video coding step.
[0103] Some embodiments may be described using the following clause-based format.
[0104] (1) A method of video processing, comprising: 1. A method comprising: performing a conversion between a current video block and a bitstream representation of the current video block, wherein during the conversion, a second set of color component values of the current video block are derived from a first set of color component values included in one or more reference frames, the first set of color component values being usable in a linear model of a video coding step.
[0105] (2) The method of item 1, wherein the first set of color component values is interpolated before being used in the linear model of the video coding step.
[0106] (3) The method according to one or more of items 1 to 2, wherein a linear combination of the first set of color component values is usable as a parameter in the linear model.
[0107] (4) The method of clause 1, wherein the location of the first set of color component values included in the one or more reference frames is selected based at least in part on motion information of the current video block.
[0108] (5) The method of clause 4, wherein the positions of luma component values in the one or more reference frames are calculated from the positions of corresponding luma component values in the current video block and the motion information of the current video block.
[0109] (6) The method of clause 5, wherein the location of the corresponding luma component value is the top-left sample, the center sample, or the bottom-right sample of the current video block.
[0110] (7) The method of clause 6, wherein the motion information of the current video block corresponds to an integer motion vector or a fractional motion vector.
[0111] (8) The method of clause 7, wherein the fractional motion vector is derived using fractional luma component values in the one or more reference frames.
[0112] (9) The method of clause 7, wherein the integer motion vector is derived by rounding towards or away from zero.
[0113] (10) The method according to item 1, wherein the position of the first color component value set included in the one or more reference frames is a predetermined position.
[0114] (11) The method according to one or more of clauses 1 to 10, wherein a median or mean of the first set of color component values is used to derive the second set of color component values for the current video block.
[0115] (12) The method according to one or more of items 1 to 11, wherein the one or more reference frames are predetermined reference frames.
[0116] (13) The method according to item 12, wherein the predetermined reference frame includes a frame having a reference index in a reference picture list.
[0117] (14) The method according to item 13, wherein the reference index is 0 and the reference picture list is 0.
[0118] (15) The method of clause 13, wherein the reference index and / or the reference picture list are signaled in the bitstream representation associated with one or more of the following: a sequence, a picture, a tile, a group, a slice, a tile, a coding tree unit row, or a video block.
[0119] (16) The method of clause 1, wherein the second set of color component values for the current video block are derived from an arithmetic average or a weighted average of the first set of color component values included in the one or more reference frames.
[0120] (17) The method described in paragraph 1, wherein the second color component value set of the current video block is selectively derived from the first color component value set included in the one or more reference frames based on whether the current video block is a bi-predictive coding block.
[0121] (18) The method of clause 17, wherein the second color component value set for the current video block is derived separately for each prediction direction of the first color component value set.
[0122] (19) The method of claim 1, wherein the second color component value set of the current video block is selectively derived from the first color component value set included in the one or more reference frames based on whether the current video block is associated with sub-block-based prediction.
[0123] (20) The method according to item 1, wherein the sub-block based prediction corresponds to affine prediction or alternative temporal motion vector prediction (ATMVP).
[0124] (21) The method according to one or more of clauses 19 to 20, wherein the second color component value set of the current video block is derived for each individual sub-block.
[0125] (22) The method of one or more of clauses 19 to 21, wherein the second color component value set for the current video block is derived for the entire current video block, independent of the sub-block-based prediction.
[0126] (23) The method of one or more of clauses 19-22, wherein the first set of color component values included in one or more reference frames is selected based at least in part on motion vectors of sub-blocks of the current video block.
[0127] (24) The method according to one or more of items 1 to 23, wherein the first set of color component values included in one or more reference frames are intermediate color component values.
[0128] (25) The method according to one or more of clauses 1 to 24, wherein the video coding step precedes another video coding step.
[0129] (26) The method described in clause 25, wherein the first color component value set included in the one or more reference frames is selected based at least in part on an intermediate motion vector of the current video block or a sub-block of the current video block, the intermediate motion vector being calculated before the further video coding step.
[0130] (27) The method according to one or more of clauses 24 to 26, wherein the video coding step comprises one or a combination of the following steps: a Bi-Directional Optical Flow (BDOF) step, a decoder-side motion vector refinement (DMVR) step, a prediction refinement optical flow (PROF) step; (28) The method according to one or more of clauses 1 to 27, wherein the first set of color component values included in the one or more reference frames corresponds to M×N luma component values associated with a corresponding luma block.
[0131] (29) The method of clause 28, wherein the corresponding luma block is a co-located luma block of the current video block.
[0132] (30) The method of clause 29, wherein the product of M and N is less than the product of the block width and block height of the co-located luma block of the current video block.
[0133] (31) The method according to one or more of clauses 27 to 30, wherein the first color component value set included in the one or more reference frames corresponds to at least a portion of reference samples identified at the positions of neighboring luma samples of the co-located luma block.
[0134] (32) The method according to one or more of items 1 to 31, wherein the first set of color component values is downsampled before being used in the linear model of the video coding step.
[0135] (33) The method of clause 1, wherein the second color component value set of the current video block is selected based at least in part on one or more of the following information of the current video block: a quantization parameter, a coding mode, or a picture order count (POC).
[0136] (34) The method according to clause 31, wherein the positions of the neighboring luma samples are such that the top-left sample of the co-located luma block is covered.
[0137] (35) The method described in clause 228, wherein the first color component value set included in the one or more reference frames corresponds to at least a portion of reference samples identified at positions outside the corresponding luma block.
[0138] (36) The method of clause 28, wherein the second set of color component values for the current video block are selectively derived from the first set of color component values included in the one or more reference frames based on the availability of neighboring samples of the corresponding luma block.
[0139] (37) The method of clause 28, wherein the availability of the neighboring samples of the corresponding luma block is based on one or more of: use of a coding mode of the current video block; use of a coding mode of the neighboring samples of the corresponding luma block; use of a type of filter associated with the neighboring samples of the corresponding luma block; and position of the neighboring samples of the corresponding luma block relative to the current block or its sub-blocks.
[0140] (38) The method of clause 28, further comprising, in response to the lack of availability of the neighboring samples of the corresponding luma block, replacing, filtering, or padding the unavailable samples with other samples.
[0141] (39) The method according to clause 28, further comprising the step of applying a smoothing filter to neighboring samples of the corresponding luma block.
[0142] (40) A method of video processing, comprising: performing a conversion between a current video block and a bitstream representation of the current video block, wherein during the conversion, a second set of color component values of the current video block are derived from a first set of color component values included in one or more reference frames, the first set of color component values being usable in a linear model of a video coding step; selectively enabling or disabling derivation of the second color component value set for the current video block based on one or more conditions associated with the co-located luma block of the current video block in response to determining that the first color component value set included in the one or more reference frames is a co-located luma block of the current video block; A method comprising:
[0143] (41) The method of clause 40, wherein the one or more conditions associated with the co-located luma block of the current video block include a partition size of the co-located luma block, a number of coding units of the co-located luma block reaching a threshold, an upper-left luma sample of the co-located luma block reaching a threshold size, a partition tree depth of the co-located luma block, or a corresponding luma block covering the upper-left luma sample of the co-located luma block, or a corresponding luma block covering the upper-left luma sample of the co-located luma block and being contained within a bounding box of a predetermined size.
[0144] (42) The method of clause 40, wherein information indicating selectively enabling or disabling the derivation is included in the bitstream representation.
[0145] (43) The method according to clause 28, wherein the availability of neighboring samples of the corresponding luma block is associated with a check on the neighboring samples according to a predetermined order.
[0146] (44) The method of one or more of clauses 1 to 43, wherein the second color component value set of the current video block is stored for use in conjunction with one or more other video blocks.
[0147] (45) The method according to one or more of clauses 1 to 44, wherein the linear model corresponds to a cross-component linear model (CCLM) and the video coding step corresponds to a luma mapping with chroma scaling (LMCS) mode.
[0148] (46) The method of one or more of clauses 1 to 45, wherein the current video block is an inter-coding block, a bi-predictive coding block, or an intra block copy (IBC) coding block.
[0149] (47) The method according to one or more of paragraphs 1 to 46, wherein the first set of color component values corresponds to luma sample values and the second set of color component values corresponds to chroma scaling factors.
[0150] (48) An apparatus in a video system including a processor and a non-transitory memory having instructions, the instructions, when executed by the processor, causing the processor to perform a method according to any one of clauses 1 to 47.
[0151] (49) A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for executing the method according to any one of paragraphs 1 to 47.
[0152] 12 is a block diagram illustrating an example video processing system 1200 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1200. System 1200 may include an input 1202 that receives video content. The video content may be received in raw or uncompressed format, e.g., 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 1202 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0153] System 1200 may include a coding component 1204 that may implement various coding or encoding methods described herein. Coding component 1204 may reduce the average bitrate of video from input 1202 to the output of coding component 1204 to generate a coded representation of the video. Coding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of coding component 1204 may be stored or transmitted over a communication connection, as represented by component 1206. The bitstream (or coded) representation received at input 1202, stored, or communicated, may be used by component 1208 to generate pixel values or displayable video that are transmitted to display interface 1210. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as "coding" operations or tools, it is understood that the coding tools or operations are used in an encoder and that corresponding decoding tools or operations that reverse the results of the coding are performed by a decoder.
[0154] Examples of peripheral bus interfaces or display interfaces may include universal serial bus (USB), high definition multimedia interface (HDMI), DisplayPort, etc. Examples of storage interfaces include serial advanced technology attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video displays.
[0155] 13 shows a flowchart of an exemplary method for video media processing. The steps of this flowchart are discussed in connection with exemplary embodiments 7d and 7e8 in Chapter 4 of this specification. In step 1302, the process performs a conversion between a current chroma video block of video media data and a bitstream representation of the current chroma video block, where during the conversion, a chroma residual of the current chroma video block is scaled based on a scaling factor, the scaling factor being derived based at least on a luma sample located at a predetermined location.
[0156] 14 shows a flowchart of an exemplary method for video media processing. The steps of this flowchart are discussed in connection with exemplary embodiment 1 in Section 4 of this specification. At step 1402, the process performs a conversion between a current video block of video media data and a bitstream representation of the current video block, where, during the conversion, a second set of color component values for the current video block are derived from a first set of color component values of the video media data using a cross-component linear model (CCLM) and / or luma mapping with chroma scaling (LMCS) mode processing step.
[0157] 15 shows a flowchart of an exemplary method for video media processing. The steps of this flowchart are discussed in connection with exemplary embodiment 7 in Section 4 of this specification. In step 1502, the process performs a conversion between a current video block of video media data and a bitstream representation of the current video block, where, during the conversion, one or more reconstructed samples associated with a current frame of the video media data are used to derive a chroma residual scaling factor in a luma mapping with chroma scaling (LMCS) mode processing step.
[0158] 16 shows a flowchart of an exemplary method for video media processing. The steps of this flowchart are discussed in connection with exemplary embodiment 8 in Chapter 4 of this specification. In step 1602, the process performs a conversion between a current video block of video media data and a bitstream representation of the current video block, where during the conversion, one or more luma prediction samples or luma reconstructed samples in a current frame other than a reference frame are used to derive a chroma residual scaling factor in a luma mapping with chroma scaling (LMCS) mode processing step.
[0159] FIG. 17 shows a flowchart of an example method for video media processing. The steps of this flowchart are discussed in connection with example embodiments 11a, 11b, 11c, and 11d in Section 4 of this specification. In step 1702, the process checks the availability of one or more neighboring luma blocks of a corresponding luma block that cover a top-left sample of the co-located luma block during conversion between a current chroma video block and a bitstream representation of the current chroma video block. In step 1704, the process determines whether to search for neighboring luma samples of the corresponding luma block based on the availability of one or more neighboring luma blocks. In step 1706, the process derives a scaling factor based on the determination. In step 1708, the process scales the chroma residual of the current chroma video block based on the scaling factor to generate a scaled chroma residual. In step 1710, the process performs a conversion based on the scaled chroma residual.
[0160] 18 shows a flowchart of an exemplary method for video media processing. The steps of this flowchart are discussed in connection with exemplary embodiment 11 in Section 4 of this specification. In step 1802, the process derives a second set of color component values for a current video block from a first set of color component values of the video media data using a model associated with the processing step during a conversion between the current video block of the video media data and a bitstream representation of the current video block, where the first set of color component values are neighboring samples of a corresponding luma block that cover the top-left sample of the co-located luma block.
[0161] 19 shows a flowchart of an exemplary method for video media processing. The steps of this flowchart are discussed in connection with exemplary embodiment 12 in Section 4 of this specification. At step 1902, the process determines to selectively enable or disable application of a cross-component linear model (CCLM) and / or chroma residual scaling (CRS) to a current chrominance video block during conversion between the current chrominance video block and a bitstream representation of the current chrominance video block based at least in part on one or more conditions associated with a co-located luma block of the current chrominance video block.
[0162] 20 shows a flowchart of an exemplary method for video media coding. The steps of this flowchart are discussed in connection with exemplary embodiment 14 in Chapter 4 of this specification. At step 2002, the process selectively enables or disables application of luma-dependent chroma residual scaling (CRS) to chroma components of a current video block in a video domain of the video media data for encoding the current video block into a bitstream representation of the video media data. At step 2004, the process determines whether to include or exclude a field in the bitstream representation of the video media data, where the field indicates selective enabling or disabling and, if included, is signaled outside the first syntax level associated with the current video block.
[0163] 21 shows a flowchart of an example method for video media decoding. The steps of this flowchart are discussed in connection with example embodiment 14 in Section 4 of this specification. At step 2102, the process parses fields in a bitstream representation of video media data, where the fields are included in a level other than a first syntax level associated with a current video block. At step 2104, the process selectively enables or disables application of luma-dependent chroma residual scaling (CRS) to chroma components of the current video block of the video media data based on the fields to generate a decoded video region from the bitstream representation.
[0164] 22 shows a flowchart of an exemplary method for video media coding. The steps of this flowchart are discussed in connection with exemplary embodiment 15 in Section 4 of this specification. At step 2202, the process selectively enables or disables application of a cross-component linear model (CCLM) to a current video block of video media data for encoding the current video block into a bitstream representation of the video media data. At step 2204, the process makes a determination of whether to include or exclude a field in the bitstream representation of the video media data, where the field indicates selective enabling or disabling and, if included, is signaled outside the first syntax level associated with the current video block.
[0165] 23 shows a flowchart of an exemplary method for video media decoding. The steps of this flowchart are discussed in connection with exemplary embodiment 15 in Section 4 of this specification. At step 2302, the process parses fields in a bitstream representation of video media data, where the fields are included in a level other than a first syntax level associated with the current video block. At step 2304, the process selectively enables or disables application of a cross-component linear model (CCLM) to the current video block of the video media data based on the fields to generate a decoded video region from the bitstream representation.
[0166] Some embodiments discussed herein are presented below in a section-based format.
[0167] (Y1) A method for processing video media, checking availability of one or more neighboring luma blocks of a corresponding luma block that cover a top-left sample of a co-located luma block during conversion between a current chroma video block and a bitstream representation of the current chroma video block; determining whether to search for neighboring luma samples of the corresponding luma block based on the availability of one or more neighboring luma blocks; deriving a scaling factor based on said determining step; scaling a chroma residual of the current chroma video block based on the scaling factor to generate a scaled chroma residual; performing the transformation based on the scaled chroma residual; A method comprising:
[0168] (Y2) The method according to item Y1, wherein the one or more neighboring luma blocks include a left neighboring luma block and an above neighboring luma block.
[0169] (Y3) The method according to item Y1, wherein the neighboring luma samples include one or more left neighboring sample columns and / or one or more above neighboring sample rows of the corresponding luma block.
[0170] (Y4) The method according to item Y1 or Y3, wherein, if the one or more neighboring luma blocks are available, the neighboring luma samples are searched and averaged based on rounding to derive the scaling factor.
[0171] (Y5) The method described in terms Y1 or Y3, wherein the neighboring luma samples are searched and a median value of the neighboring luma samples is used to derive the scaling factor if the one or more neighboring luma blocks are available.
[0172] (Y6) The method according to item Y3, wherein the number of neighboring luma samples is N, where 1<=N<=2W+2H, and W and H are the width and height of the corresponding luma block.
[0173] (Y7) The method described in item Y1 or Y2, wherein the availability of one or more neighboring luma blocks is determined based on a current picture, tile, tile group, virtual pipeline data unit (VPDU), or slice.
[0174] (Y8) The method described in clause Y7, wherein the one or more neighboring luma blocks are unavailable if the one or more neighboring blocks are located in different pictures, different tiles, different tile groups, different VPDUs, or different slices.
[0175] (Y9) The method according to item Y1, wherein the corresponding neighboring luma sample is skipped from the search if the one or more neighboring luma blocks are unavailable.
[0176] (Y10) The method according to item Y9, wherein the scaling factor is derived when the corresponding neighboring luma sample is skipped from the search.
[0177] (Y11) The method described in item Y10, wherein the scaling coefficient is derived based on a default value.
[0178] (Y12) The method of claim Y11, wherein the default value is based on the bit depth of the current chroma video block and the co-located luma block.
[0179] (Y13) The method according to clause Y12, wherein the default value is expressed as 1<<(bitDepth-1), where bitDepth indicates the bit depth of the current chroma video block and the co-located luma block.
[0180] (Y14) The method according to item Y1, wherein the neighboring luma samples are reconstructed based on forward mapping.
[0181] (X1) A method of video media processing, comprising: performing a conversion between a current chroma video block of video media data and a bitstream representation of the current chroma video block; during the conversion, a chroma residual of the current chroma video block is scaled based on a scaling factor, the scaling factor being derived based at least on a luma sample located at a predetermined position.
[0182] (X2) The method according to item X1, wherein the scaling factor is calculated using a function applied to the luma sample at the predetermined position.
[0183] (X3) The method according to item X1, wherein the function is a median function or a mean function based on rounding.
[0184] (X4) The method according to item X1, wherein the predetermined position is determined based on a co-located luma block corresponding to the current chroma video block.
[0185] (A1) A method of processing video media, comprising: 1. A method, comprising: performing a conversion between a current video block of video media data and a bitstream representation of the current video block, wherein during the conversion, a second set of color component values of the current video block are derived from a first set of color component values of the video media data using a cross-component linear model (CCLM) and / or a luma mapping with chroma scaling (LMCS) mode processing step.
[0186] (A2) The method of claim A1, wherein the first set of color component values are reference samples of the current video block, and the second set of color component values are chroma residual scaling factors in the LMCS mode.
[0187] (A3) The method according to item A2, wherein the reference sample is a reference luma sample that is interpolated before deriving a chroma residual scaling factor.
[0188] (A4) The method according to item A1, wherein the first set of color component values are reference samples contained in different reference frames.
[0189] (A5) The method according to any one of paragraphs A1 to A4, wherein the position of the reference sample is calculated from the position of the corresponding luma component value in the current video block and motion information of the current video block.
[0190] (A6) The method of claim A5, wherein the location of the corresponding luma component value is a top-left sample, a center sample, or a bottom-right sample of the current video block.
[0191] (A7) The method of clause A6, wherein the motion information of the current video block corresponds to an integer motion vector or a fractional motion vector.
[0192] (A8) The method of claim A7, wherein the fractional motion vector is derived using a fractional luma component value in a reference frame.
[0193] (A9) The method according to paragraph A7, wherein the integer motion vector is derived by rounding towards or away from zero.
[0194] (A10) The method according to one or more of paragraphs A1 to A2, wherein the first set of color component values is included in a predetermined reference frame of the video media data.
[0195] (A11) The method according to one or more of paragraphs A1-A10, wherein a median or mean value of the first set of color component values is used to derive the second set of color component values for the current video block.
[0196] (A12) The method according to clause A10, wherein the predetermined reference frame includes a frame having a reference index in a reference picture list.
[0197] (A13) The method according to item A12, wherein the reference index is 0 and the reference picture list is 0.
[0198] (A14) A method according to one or more of paragraphs A1-A2, wherein the first set of color component values is included in multiple reference frames of the video media data, and a weighted combination of the first set of color component values is used to derive the second set of color component values.
[0199] (A15) The method described in clause A13, wherein the reference index and / or the reference picture list are signaled as fields in the bitstream representation associated with one or more of the following: a sequence, a group of pictures, a picture, a tile, a tile group, a slice, a subpicture, a coding tree unit row, a coding tree unit, a virtual pipeline data unit (VPDU), or a video block.
[0200] (A16) The method of claim A1, wherein the second set of color component values for the current video block are selectively derived from the first set of color component values based on whether the current video block is a bi-predictive coding block.
[0201] (A17) The method of clause A16, wherein the second set of color component values for the current video block are derived separately for each prediction direction associated with the first set of color component values.
[0202] (A18) The method of claim A1, wherein the second set of color component values for the current video block are selectively derived from the first set of color component values based on whether the current video block is associated with sub-block based prediction.
[0203] (A19) The method according to clause A18, wherein the sub-block based prediction corresponds to affine prediction or alternative temporal motion vector prediction (ATMVP).
[0204] (A20) The method according to one or more of paragraphs A18-A19, wherein the second set of color component values for the current video block are derived for individual sub-blocks.
[0205] (A21) The method according to one or more of paragraphs A18 to A19, wherein the second set of color component values for the current video block is derived for the entire current video block, independent of prediction based on the sub-blocks.
[0206] (A22) The method of one or more of paragraphs A18-A21, wherein the first set of color component values is selected based at least in part on motion vectors of sub-blocks of the current video block.
[0207] (A23) The method according to one or more of paragraphs A18 to A21, wherein a motion vector associated with a sub-block of the current video block is used to select the first set of color component values.
[0208] (A24) The method according to one or more of paragraphs A1 to A23, wherein the first set of color component values are intermediate color component values.
[0209] (A25) The method according to one or more of paragraphs A1 to A24, wherein the LMCS mode processing step precedes another subsequent processing step.
[0210] (A26) The method described in paragraph A25, wherein the first color component value set is selected based at least in part on an intermediate motion vector of the current video block or a sub-block of the current video block, the intermediate motion vector being calculated before the further video coding step.
[0211] (A27) The method described in clause A26, wherein the further processing step includes one or a combination of the following: a Bi-Directional Optical Flow (BDOF) step, a decoder-side motion vector refinement (DMVR) step, or a prediction refinement optical flow (PROF) step.
[0212] (A28) A method for processing video media, comprising: 1. A method, comprising: performing a conversion between a current video block of video media data and a bitstream representation of the current video block, wherein during the conversion, one or more reconstructed samples associated with a current frame of the video media data are used to derive a chroma residual scaling factor in a luma mapping with chroma scaling (LMCS) mode processing step.
[0213] (A29) The method of clause A28, wherein the current video block is intra block copy (IBC) coded.
[0214] (A30) The method according to item A28, wherein the one or more reconstructed samples are reference samples in a reference block associated with the current frame.
[0215] (A31) The method described in paragraph A28, wherein the one or more reconstituted samples are predetermined.
[0216] (A32) The method of claim A31, wherein the one or more reconstructed samples are reconstructed luma samples located in an upper row and left column adjacent to a block that covers a corresponding luma block of the current video block.
[0217] (A33) The method of clause A28, wherein the position of the one or more reconstructed samples is based on a position of a corresponding luma block of the current video block and motion information of the current video block.
[0218] (A34) A method for processing video media, comprising: 1. A method, comprising: performing a transformation between a current video block of video media data and a bitstream representation of the current video block, wherein during the transformation, one or more luma prediction samples or luma reconstructed samples in a current frame other than a reference frame are used to derive a chroma residual scaling factor in a luma mapping with chroma scaling (LMCS) mode processing step.
[0219] (A35) The method according to item A34, wherein the one or more luma prediction samples or luma reconstructed samples are located in a neighboring region of an M×N luma block covering a corresponding luma block.
[0220] (A36) The method according to clause A35, wherein the corresponding luma block is a co-located luma block of the current video block.
[0221] (A37) The method of clause A36, wherein the product of M and N is less than the product of the block width and block height of the co-located luma block of the current video block.
[0222] (A38) The method according to clause A36, wherein M and N are a predetermined width and a predetermined height of a video block covering the co-located luma block of the current video block.
[0223] (A39) The method according to clause A36, wherein M and N are the width and height of a virtual pipeline data unit (VPDU) covering the co-located luma block of the current video block.
[0224] (A40) The method according to one of paragraphs A1 to A39 above, wherein during the transformation, the reference samples are used directly or are downsampled before being used in the derivation.
[0225] (A41) A method according to one or more of clauses A1 to A40, wherein the samples used in deriving the chroma residual scaling factor are selected based at least in part on one or more of the following information of the current video block: a quantization parameter, a coding mode, or a picture order count (POC).
[0226] (A42) The method of one or more of clauses A1-A41, wherein the current video block is an inter-coding block, an intra-coding block, a bi-predictive coding block, or an intra block copy (IBC) coding block.
[0227] (A43) The method according to one or more of paragraphs A1 to A42, wherein the first set of color component values corresponds to luma sample values and the second set of color component values corresponds to chroma scaling factors of the current video block.
[0228] (A44) The method according to one or more of paragraphs A1 to A43, wherein the transforming includes generating the bitstream representation from the current video block.
[0229] (A45) The method according to one or more of paragraphs A1 to A43, wherein the converting includes generating pixel values of the current video block from the bitstream representation.
[0230] (A46) A video encoder device including a processor configured to implement the method according to one or more of paragraphs A1 to A45.
[0231] (A47) A video decoder device including a processor configured to implement the method according to one or more of paragraphs A1 to A45.
[0232] (A48) A computer-readable medium having code stored thereon, the code embodying processor-executable instructions for carrying out the methods described in one or more of paragraphs A1 to A38.
[0233] (B1) A method of processing video media, comprising: deriving a second set of color component values for a current video block from a first set of color component values for the video media data using a model associated with a processing step during a conversion between a current video block of the video media data and a bitstream representation of the current video block, the first set of color component values being neighboring samples of a corresponding luma block that cover a top-left sample of the co-located luma block; A method comprising:
[0234] (B2) The method of clause B1, wherein the current video block is one of an intra-coded video block with a dual tree partition, an intra-coded video block with a single tree partition, or an inter-coded video block with a single tree partition.
[0235] (B3) The method according to one or more of paragraphs B1 to B2, wherein the first color component value set corresponds to at least a reference sample position identified at a position outside the corresponding luma block.
[0236] (B4) The method of claim B3, wherein the portion of reference samples identified at locations outside the corresponding luma block includes samples adjacent to the corresponding luma coding block.
[0237] (B5) The method described in clause B4, wherein the samples adjacent to the corresponding luma coding block include N samples located in a left neighboring column and / or an upper neighboring row of the corresponding luma coding block, where N=1...2W+2H, where W and H are the width and height of the corresponding luma coding block.
[0238] (B6) The method according to clause B5, wherein the upper neighboring luma sample is located at (xCb+W / 2, yCb-1) or (xCb-1, yCb-1) when the top-left sample of the co-located luma block is located at (xCb, yCb).
[0239] (B7) The method according to clause B5, wherein the left neighboring luma sample is located at (xCb+W-1, yCb-1) when the top-left sample of the co-located luma block is located at (xCb, yCb).
[0240] (B8) The method according to paragraph B4, wherein the portion of the reference samples identified at positions outside the corresponding luma block are at predetermined positions.
[0241] (B9) The method described in paragraph B4, wherein the second color component value set of the current video block is derived based on the median or arithmetic mean of N samples located in a left neighboring column and / or an upper neighboring row of the corresponding luma coding block.
[0242] (B10) The method according to one or more of paragraphs B1-B2, wherein the second set of color component values for the current video block are selectively derived from the first set of color component values based on the availability of neighboring samples of the corresponding luma block.
[0243] (B11) The method described in clause B10, wherein the availability of the neighboring samples of the corresponding luma block is based on one or more of: use of the coding mode of the current video block; use of the coding mode of the neighboring samples of the corresponding luma block; use of the type of filter associated with the neighboring samples of the corresponding luma block; position of the neighboring samples of the corresponding luma block relative to the current block or its sub-blocks; width of the current picture / subpicture / tile / tile group / VPDU / slice; and / or height of a current picture / subpicture / tile / tile group / VPDU / slice / coding tree unit (CTU) row.
[0244] (B12) The method of clause B10, further comprising, in response to determining the lack of availability of the neighboring samples of the corresponding luma block, replacing, filtering, or padding the unavailable samples with other samples.
[0245] (B13) The method described in clause B12, wherein the lack of availability of the neighboring samples of the corresponding luma block is based at least in part on a determination that when the coding mode of the current video block is an inter mode, the coding mode of the neighboring samples is an intra mode and / or an intra block copy (IBC) mode and / or a joint inter-intra prediction (CIIP) mode and / or a local luminance compensation (LIC) mode.
[0246] (B14) The method described in clause B12, wherein the lack of availability of the neighboring samples of the corresponding luma block is based at least in part on a determination that the neighboring samples are subjected to a diffusion filter and / or a bilateral filter and / or a Hadamard transform filter when the coding mode of the current video block is inter mode.
[0247] (B15) The method described in clause B12, wherein the lack of availability of the neighboring samples of the corresponding luma block is based at least in part on a determination that the neighboring block is located outside a current picture / subpicture / tile / tile group / VPDU / slice / coding tree unit (CTU) row associated with the current video block.
[0248] (B16) The method according to one or more of clauses B12 to B15, further comprising a step of disabling derivation of the second color component value set of the current video block from the first color component value set in response to determining the lack of availability of the neighboring samples of the corresponding luma block.
[0249] (B17) The method of clause B10, further comprising: disabling derivation of the second color component value set of the current video block from the first color component value set in response to determining that the number of available neighboring samples is less than a threshold.
[0250] (B18) The method according to item B17, wherein the threshold value is 1.
[0251] (B19) The method described in clause B12, wherein if it is determined that a neighboring sample is not available, the neighboring sample is filled with 1<<(bitDepth-1) samples, where bitDepth indicates the bit depth of the samples of the first color component value set or the samples of the second color component value set.
[0252] (B20) The method according to clause B12, wherein if a neighboring sample is determined to be unavailable, the neighboring sample is replaced by the first available neighboring sample according to a predetermined checking order.
[0253] (B21) The method of clause B13, wherein if it is determined that a neighboring sample is not available, the neighboring sample is padded with one or more of a left neighboring sample, a right neighboring sample, an upper neighboring sample, or a lower neighboring sample.
[0254] (B22) The method of clause B10, further comprising applying a smoothing filter to neighboring samples of the corresponding luma block used in deriving the second set of color component values for the current video block.
[0255] (B23) The method of claim B22, wherein the smoothing filter includes one or more of the following: a bilateral filter, a Hadamard transform filter, or a forward mapping of a reshaper domain.
[0256] (B24) The method according to one or more of clauses B1 to B23, wherein the second color component value set of the current video block is stored for use in conjunction with one or more other video blocks.
[0257] (B25) The method according to one or more of paragraphs B1 to B23, wherein the model corresponds to a cross-component linear model (CCLM) and / or the processing step corresponds to a luma mapping with chroma scaling (LMCS) mode.
[0258] (B26) The method of one or more of clauses B1-B23, wherein the current video block is an intra-coded block, an inter-coded block, a bi-predictively coded block, or an intra-block copy (IBC) coded block.
[0259] (B27) The method according to one or more of paragraphs B1 to B23, wherein the first set of color component values corresponds to luma sample values and the second set of color component values corresponds to chroma scaling factors of the current video block.
[0260] (B28) The method according to one or more of clauses B1 to B23, wherein the converting includes generating the bitstream representation from the current video block.
[0261] (B29) The method according to one or more of clauses B1 to B23, wherein the converting includes generating pixel values of the current video block from the bitstream representation.
[0262] (B30) A video encoder device including a processor configured to implement the method according to one or more of paragraphs B1 to B23.
[0263] (B31) A video decoder device including a processor configured to implement the method according to one or more of paragraphs B1 to B23.
[0264] (B32) A computer-readable medium having code stored thereon, said code embodying processor-executable instructions for carrying out the methods described in one or more of paragraphs B1 to B23.
[0265] (C1) A method of processing video media, comprising: 1. A method, comprising: determining, during a conversion between a current chrominance video block of video media data and a bitstream representation of the current chrominance video block, to selectively enable or disable application of a cross-component linear model (CCLM) and / or chroma residual scaling (CRS) to the current chrominance video block based at least in part on one or more conditions associated with a co-located luma block of the current chrominance video block.
[0266] (C2) The method of clause C1, wherein the one or more conditions associated with the co-located luma block of the current chroma video block include a partition size of the co-located luma block, a number of coding units of the co-located luma block reaching a threshold, a top-left luma sample of the co-located luma block reaching a threshold size, a depth of a partition tree of the co-located luma block, a corresponding luma block covering the top-left luma sample of the co-located luma block, a corresponding luma block covering the top-left luma sample of the co-located luma block and being contained within a bounding box of a predetermined size, a coding mode of one or more coding units (CUs) covering at least one sample of the co-located luma block, and / or dimensions of the current chroma video block.
[0267] (C3) The method of one or more of clauses C1-C2, wherein application of the CCLM and / or CRS to the current chroma video block is disabled in response to determining that the co-located luma block of the current chroma video block is divided into multiple partitions.
[0268] (C4) The method of one or more of clauses C1 to C2, wherein application of the CCLM and / or CRS to the current chroma video block is enabled in response to determining that the co-located luma block of the current chroma video block is not divided into multiple partitions.
[0269] (C5) The method of one or more of clauses C1-C2, wherein application of the CCLM and / or CRS to the current chroma video block is disabled in response to determining that the co-located luma block of the current chroma video block includes more than one of a threshold number of coding units and / or a threshold number of partition units and / or a threshold number of transform units.
[0270] (C6) The method according to item C5, wherein the threshold number is 1.
[0271] (C7) The method of claim C5, wherein the threshold number is based at least in part on whether the CCLM and / or the CRS applies.
[0272] (C8) The method of claim C5, wherein the threshold number is fixed or included in the bitstream representation.
[0273] (C9) The method of clause C5, wherein the threshold number is based at least in part on a profile / level / tier associated with the current chroma video block.
[0274] (C10) The method according to clause C5, wherein the coding unit and / or the partition unit and / or the transform unit are all located within the co-located luma block.
[0275] (C11) The method according to claim C5, wherein the coding unit and / or the partition unit and / or the transform unit are partially located within the co-located luma block.
[0276] (C12) The method according to clause C11, wherein the coding unit and / or the partition unit and / or the transform unit are located, in part, along the boundary of the co-located luma block.
[0277] (C13) The method according to clause C5, wherein the coding unit and / or the partition unit and / or the transform unit are associated with sub-block based prediction.
[0278] (C14) The method according to clause C13, wherein the sub-block based prediction corresponds to Intra Sub-Partitions (ISP) or Affine Prediction or Alternative Temporal Motion Vector Prediction (ATMVP).
[0279] (C15) A method according to one or more of clauses C1 to C2, wherein application of the CCLM and / or the CRS to the current chroma video block is disabled in response to determining that a coding unit and / or partition unit and / or transform unit covering the top-left luma sample of the co-located luma block is larger than a predetermined block size.
[0280] (C16) The method according to clause C15, wherein the co-located luma block is 32x32 in size and is contained within a corresponding luma block of size 64x64, and the predetermined luma block size is 32x64.
[0281] (C17) The method described in clause C2, wherein application of the CCLM and / or the CRS to the current chroma video block is enabled in response to determining that the co-located luma block of the current chroma video block is not split and that the corresponding luma block covering the top-left luma sample of the co-located luma block is completely contained within the bounding box of a predetermined size.
[0282] (C18) The method according to claim C17, wherein the corresponding blocks are 32x64 in size and the bounding boxes are 40x70 in size.
[0283] (C19) The method described in clause C17, wherein the predetermined size of the bounding box is based in part on the size of a coding tree unit (CTU) associated with the current chroma video block and / or the size of a coding unit (CU) associated with the current chroma video block.
[0284] (C20) A method according to one or more of clauses C1 to C2, wherein the co-located luma block of the current chrominance video block is divided into a plurality of partitions, and predicted or reconstructed samples within one or more of the plurality of partitions are used to derive a value associated with the CRS of the current chrominance video block.
[0285] (C21) The method described in clause C20, wherein an average of the predicted samples or the reconstructed samples within the first partition of the co-located luma block of the current chroma video block is used to derive the value associated with the CRS of the current chroma video block.
[0286] (C22) The method described in clause C20, wherein an upper-left predicted sample or an upper-left reconstructed sample within the first partition of the co-located luma block of the current chroma video block is used to derive the value associated with the CRS of the current chroma video block.
[0287] (C23) The method described in clause C20, wherein a central predicted sample or a central reconstructed sample within the first partition of the co-located luma block of the current chroma video block is used to derive the color component value of the current chroma video block.
[0288] (C24) The method described in clause C2, wherein application of the CCLM and / or the CRS to the current chroma video block is disabled in response to determining that the coding mode of the one or more coding units (CUs) covering the at least one sample of the co-located luma block is one of an affine mode, a bi-prediction mode, a Bi-Directional Optical Flow (BDOF) mode, a DMVR mode, a matrix affine prediction mode, an inter mode, or an Intra Sub-Partitions (ISP) mode.
[0289] (C25) The method according to claim C2, wherein the one or more coding units (CUs) covering the at least one sample of the co-located luma block are the corresponding luma blocks.
[0290] (C26) A method according to one or more of clauses C1 to C25, further comprising a step of indicating, based on a field in the bitstream representation, that the CCLM and / or the CRS are selectively enabled or disabled for the current chroma video block.
[0291] (C27) A method according to one or more of clauses C1 to C26, wherein selectively enabling or disabling the application of the CCLM and / or the CRS to the current chroma video block is performed on one or more sub-blocks of the current chroma video block.
[0292] (C28) The method of claim C27, wherein the one or more sub-blocks of the current chroma video block are 2x2 or 4x4 in size.
[0293] (C29) The method described in clause C27, wherein application of the CCLM and / or the CRS is enabled for the sub-block of the current chroma video block when the corresponding luma coding block of the current chroma video block covers all samples of the corresponding block of the sub-block.
[0294] (C30) The method described in clause C27, wherein application of the CCLM and / or the CRS is disabled for the sub-block of the current chroma video block when all samples of the corresponding block of the sub-block are not covered by the corresponding luma coding block.
[0295] (C31) The method of claim C27, wherein the CCLM and / or CRS parameters are associated with each sub-block of the current chroma video block.
[0296] (C32) The method described in clause C27, wherein selectively enabling or disabling the application of the CCLM and / or the CRS to sub-blocks of the current chroma video block is based on samples contained within the co-located luma block.
[0297] (C33) The method according to one or more of clauses C1 to C32, wherein the current chroma video block is an inter-coding block, an intra-coding block, a bi-predictive coding block, or an intra-block copy (IBC) coding block.
[0298] (C34) The method according to one or more of clauses C1 to C33, wherein the conversion includes generating the bitstream representation from the current chroma video block.
[0299] (C35) The method according to one or more of clauses C1 to C33, wherein the converting includes generating pixel values of the current chroma video block from the bitstream representation.
[0300] (C36) A video encoder device including a processor configured to implement the method according to one or more of clauses C1 to C33.
[0301] (C37) A video decoder device including a processor configured to implement the method according to one or more of clauses C1 to C33.
[0302] (C38) A computer-readable medium having code stored thereon, the code embodying processor-executable instructions for performing the methods described in one or more of clauses C1 to C33.
[0303] (D1) A method of video media encoding, comprising: selectively enabling or disabling application of luma-dependent chroma residual scaling (CRS) to chroma components of a current video block in a video domain of the video media data for encoding the current video block into a bitstream representation of the video media data; determining to include or exclude a field in the bitstream representation of the video media data, the field indicating the selective enabling or disabling, and if included, signaled outside a first syntax level associated with the current video block.
[0304] (D2) A method of decoding video media, comprising: Parsing a field within a bitstream representation of video media data, the field being included in a level other than a first syntax level currently associated with the video block; and selectively enabling or disabling application of luma-dependent chroma residual scaling (CRS) to chroma components of the current video block of video media data to generate a decoded video region from the bitstream representation based on the field.
[0305] (D3) The method of one or more of clauses D1-D2, wherein the first syntax level is a tile group header level, and the field is included in one of a sequence parameter set (SPS) associated with the current video block, a tile associated with the current video block, a coding tree unit (CTU) row associated with the current video block, a coding tree unit (CTU) associated with the current video block, a virtual pipeline data unit (VPDU) associated with the current video block, or a coding unit (CU) associated with the current video block.
[0306] (D4) A method according to one or more of items D1 to D3, wherein the field is a flag indicated as chroma_residual_scale_flag.
[0307] (D5) The method described in one or more of items D1 to D4, wherein the field is associated with a syntax level, and if the field is 1, application of the luma-dependent chroma residual scaling below the syntax level is enabled, and if the field is 0, application of the luma-dependent chroma residual scaling below the syntax level is disabled.
[0308] (D6) The method described in item D5, wherein the field is associated with a partition node level, and when the field is 1, application of the luma-dependent chroma residual scaling below the partition node level is enabled, and when the field is 0, application of the luma-dependent chroma residual scaling below the partition node level is disabled.
[0309] (D7) A method according to one or more of paragraphs D1 to D4, wherein the field is associated with a threshold dimension, and if the field is 1, application of the luma-dependent chroma residual scaling is enabled for video blocks equal to or greater than the threshold dimension, and if the field is 0, application of the luma-dependent chroma residual scaling is disabled for video blocks below the threshold dimension.
[0310] (D8) The method according to item D7, wherein the threshold dimension is 32x32.
[0311] (D9) A method according to one or more of items D1 to D8, wherein the field is not signaled in the bitstream representation, and the absence of the field in the bitstream representation is used to estimate that application of the luma-dependent chroma residual scaling is disabled, and the field is estimated to be 0.
[0312] (D10) The method of one or more of paragraphs D1 to D9, wherein a value associated with the luma-dependent CRS of the current video block is stored for use in connection with one or more other video blocks.
[0313] (D11) The method according to item D10, wherein the value associated with the luma-dependent CRS is derived after encoding or decoding of a luma block.
[0314] (D12) The method described in paragraph D11, wherein within the luma block, predicted samples and / or intermediate predicted samples and / or reconstructed samples and / or reconstructed samples before loop filtering are used to derive the value associated with the luma-dependent CRS.
[0315] (D13) The method according to clause D12, wherein the loop filtering includes the use of a deblocking filter and / or a sample adaptive offset (SAO) filter and / or a bilateral filter and / or a Hadamard transform filter and / or an adaptive loop filter (ALF).
[0316] (D14) The method described in item D11, wherein samples in the row below and / or the column to the right of the luma block are used to derive the value associated with the luma-dependent CRS.
[0317] (D15) The method described in item D11, wherein samples associated with neighboring blocks are used to derive the value associated with the luma-dependent CRS.
[0318] (D16) The method of clause D15, wherein the current video block is an intra-coding block, an inter-coding block, a bi-predictive coding block, or an intra block copy (IBC) coding block.
[0319] (D17) The method according to item D15, wherein the availability of samples associated with the neighboring blocks is checked according to a predetermined order.
[0320] (D18) The method of claim D17, wherein the predetermined order for the current video block is one of: left to right, top left to top right, left to top, top left to top right, top to left, top right to top left.
[0321] (D19) The method of claim D17, wherein the predetermined order for the current video block is one of bottom left to left to top right to top to top left.
[0322] (D20) The method according to clause D17, wherein the predetermined order for the current video block is one of: left to top, top right, bottom left, top left.
[0323] (D21) The method according to item D17, wherein the predetermined order is associated with samples in a first available subset of the neighboring blocks.
[0324] (D22) The method of clause D15, wherein if the current video block is an inter-coding block and the neighboring block is an intra-coding block, an IBC coding block, or a CIIP coding block, it is determined that the sample associated with the neighboring block is unavailable.
[0325] (D23) The method of clause D15, wherein if the current video block is a CIIP coding block and the neighboring block is an intra-coding block, an IBC coding block, or a CIIP coding block, it is determined that the sample associated with the neighboring block is unavailable.
[0326] (D24) The method of clause D15, further comprising disabling derivation of the luma-dependent CRS in response to determining that the number of neighboring blocks is less than a threshold.
[0327] (D25) The method according to item D24, wherein the threshold value is 1.
[0328] (D26) The method described in item D24, wherein if it is determined that no samples from neighboring blocks are available, the samples are filled with 1<<(bitDepth-1) samples, where bitDepth indicates the bit depth of the chroma or luma component.
[0329] (E1) A method of video media coding, comprising: selectively enabling or disabling application of a cross-component linear model (CCLM) to the current video block of video media data for encoding the current video block into a bitstream representation of the video media data; determining to include or exclude a field in the bitstream representation of the video media data, the field indicating the selective enabling or disabling, and if included, signaled outside a first syntax level associated with the current video block.
[0330] (E2) A method of decoding video media, comprising: Parsing a field within a bitstream representation of video media data, the field being included in a level other than a first syntax level currently associated with the video block; and selectively enabling or disabling application of a cross-component linear model (CCLM) to the current video block of video media data based on the field to generate a decoded video region from the bitstream representation.
[0331] (E3) The method of one or more of clauses E1-E2, wherein the first syntax level is a sequence parameter set (SPS) level, and the field is included in one of: a picture parameter set (PPS) associated with the current video block, a slice associated with the current video block, a picture header associated with the current video block, a tile associated with the current video block, a tile group associated with the current video block, a coding tree unit (CTU) row associated with the current video block, a coding tree unit (CTU) associated with the current video block, a virtual pipeline data unit (VPDU) associated with the current video block, or a coding unit (CU) associated with the current video block.
[0332] (E4) The method according to one or more of clauses E1 to E3, wherein the field is a flag indicated as cclm_flag.
[0333] (E5) The method according to one or more of clauses E1 to E4, wherein the absence of the bitstream representation is used to infer that application of the CCLM is disabled.
[0334] (E6) The method according to one or more of clauses E1 to E4, wherein the presence of the field in the bitstream representation is used to infer that application of the CCLM is enabled.
[0335] (E7) The method described in clause E5, wherein if the dimensions of the current video block are less than or equal to a threshold dimension, the field is excluded from the bitstream representation, whereby the exclusion of the field is used to infer that application of the CCLM is disabled.
[0336] (E8) The method described in item E7, wherein the threshold dimension is 8x8.
[0337] (F1) A video encoder device including a processor configured to implement the methods described in one or more of paragraphs X1-E8.
[0338] (F2) A video decoder device including a processor configured to implement the methods described in one or more of paragraphs X1-E8.
[0339] (F3) A computer-readable medium having code stored thereon, the code embodying processor-executable instructions for performing the methods described in one or more of paragraphs X1-E8.
[0340] As used herein, the terms "video processing" or "visual media processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion of video from a pixel representation to a corresponding bitstream representation or vice versa. The bitstream representation of a current video block may correspond to bits that are co-located or spread across different locations in the bitstream, as defined by the syntax. For example, a macroblock may be coded in terms of a converted coded error residual value, using bits from the header and other fields in the bitstream as well. Furthermore, during conversion, a decoder may parse the bitstream with the knowledge that some fields may be present or absent based on a decision, as described in the solutions above. Similarly, an encoder may determine whether a particular syntax field is included or not and generate the coded representation accordingly by including or excluding the syntax field in the coded representation.
[0341] From the foregoing, it will be understood that, although specific embodiments of the technology of the present disclosure have been described herein for purposes of illustration, various modifications may be made without departing from the scope of the invention. Accordingly, the technology of the present disclosure is not to be limited except as by the appended claims.
[0342] Implementations of the subject matter and functional operations described herein may be implemented in various systems, digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed herein, and structural equivalents thereof, or any combination of one or more of these. Implementations of the subject matter described herein may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions, encoded on a tangible, non-transitory computer-readable medium for execution by or to control the operation of a data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing unit" or "data processing device" encompasses any device, apparatus, and machine that processes data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, a device may include code that creates an execution environment for a subject computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.
[0343] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including a stand-alone program or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple associated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0344] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs that perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and devices may be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0345] Processors suitable for the execution of a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or be operatively coupled to receive data from or transfer data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all types of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0346] The present specification, together with the drawings, are intended to be exemplary only, and exemplary means illustrative. As used herein, the use of "or" is intended to include "and / or" unless the context clearly indicates otherwise.
[0347] While the present specification contains numerous specificities, these should not be considered as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features inherent in particular embodiments of particular inventions. Certain features described herein in the context of separate implementations may also be combined in a single implementation. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments. Furthermore, while features may be described above as working in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be separated from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0348] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring that such operations be performed in the particular order shown, or sequentially, or that all of the illustrated operations be performed, to achieve desirable results. Further, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all embodiments.
[0349] Only a few implementations and examples are described; other implementations, extensions, and variations may be made based on what is described and shown herein.
Claims
1. 1. A method for processing video data, comprising: determining that a scaling operation is applied to chroma residual samples of a current chrominance video block during conversion between the current chrominance video block and a bitstream of the video; performing the transformation by applying the scaling operation to the chroma residual samples; Including, the scaling process scales the chroma residual samples based on a scaling factor before using them to reconstruct the current chroma video block; the scaling factor is derived based on an average luma variance calculated based on neighboring luma samples of a video unit of the video determined based on a luma sample corresponding to a top-left sample of the current chroma video block; The scaling factor is checking the availability of each of one or more neighboring luma blocks of the video unit, each of the one or more neighboring luma blocks including at least one sample of the neighboring luma samples; determining whether to search for the neighboring luma samples of the video unit based on the availability of each of the one or more neighboring luma blocks; deriving the scaling factor based on the average luma variable calculated using the neighboring luma samples by a rounding-based averaging operation in response to the neighboring luma samples being available; deriving the scaling factor by setting the average luma variable to 1<<(bitDepth-1), in response to determining that one or more left-neighboring sample columns and one or more above-neighboring sample rows of the video unit are unavailable, where bitDepth is a bit depth of the video. The method is derived by:
2. The method of claim 1 , wherein the neighboring luma samples are located at predetermined positions in the vicinity of the video unit.
3. The method of claim 2 , wherein the neighboring luma samples located at the predetermined locations in the neighborhood of the video unit include reconstructed luma samples outside the video unit.
4. The method of claim 2 , wherein the neighboring luma samples located at the predetermined positions in the vicinity of the video unit include reconstructed luma samples that neighbor the video unit.
5. The method of claim 4 , wherein the reconstructed luma samples neighboring the video unit include at least one of the one or more left-neighboring sample columns or the one or more above-neighboring sample rows of the video unit.
6. The method of any one of claims 1 to 5, wherein the video unit is a virtual pipeline data unit, and the position of a top-left luma sample of the video unit is derived using size information of the virtual pipeline data unit.
7. The method according to any one of claims 1 to 6, wherein the total number of neighboring samples of the video unit is N, where N is an integer greater than 1, and the range of N depends on size information of the video unit.
8. 8. The method of claim 1, wherein checking the availability of each of the one or more neighboring luma blocks comprises checking the availability of each of the one or more neighboring luma blocks based on at least one of a width and a height of the video region.
9. For the luma video block in the video region: 1) a forward mapping operation of the luma video block, in which predicted samples of the luma video block are transformed from the original domain to a reshaped domain; or 2) a backward mapping process, which is the inverse operation of the forward mapping process, in which reconstructed samples of the luma video block in the reshaped domain are transformed to the original domain; At least one of the following is performed: The method of claim 8 , wherein the neighboring luma samples comprise reconstructed samples in the reshaped domain.
10. 10. The method of claim 8 or 9, wherein the video region is a picture.
11. The method of any of claims 1 to 10, wherein the scaling process is based on a piecewise linear model, an index identifying the piece to which the mean luma variable belongs, and the scaling factor is derived based on the index.
12. The method of any of claims 1 to 11, wherein the transforming comprises encoding the current chroma video block into the bitstream.
13. The method of any of claims 1 to 11, wherein the converting comprises decoding the current chroma video block from the bitstream.
14. 1. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions that, when executed by the processor, cause the processor to: determining that a scaling operation is applied to chroma residual samples of a current chrominance video block during conversion between the current chrominance video block and a bitstream of the video; performing the transformation by applying the scaling operation to the chroma residual samples; the scaling process scales the chroma residual samples based on a scaling factor before using them to reconstruct the current chroma video block; the scaling factor is derived based on an average luma variance calculated based on neighboring luma samples of a video unit of the video determined based on a luma sample corresponding to a top-left sample of the current chroma video block; The scaling factor is checking the availability of each of one or more neighboring luma blocks of the video unit, each of the one or more neighboring luma blocks including at least one sample of the neighboring luma samples; determining whether to search for the neighboring luma samples of the video unit based on the availability of each of the one or more neighboring luma blocks; deriving the scaling factor based on the average luma variable calculated using the neighboring luma samples by a rounding-based averaging operation in response to the neighboring luma samples being available; deriving the scaling factor by setting the average luma variable to 1<<(bitDepth-1), in response to determining that one or more left-neighboring sample columns and one or more above-neighboring sample rows of the video unit are unavailable, where bitDepth is a bit depth of the video. The equipment is derived by:
15. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to: determining that a scaling operation is applied to chroma residual samples of a current chrominance video block during conversion between the current chrominance video block and a bitstream of the video; performing the transformation by applying the scaling operation to the chroma residual samples; Including, the scaling process scales the chroma residual samples based on a scaling factor before using them to reconstruct the current chroma video block; the scaling factor is derived based on an average luma variance calculated based on neighboring luma samples of a video unit of the video determined based on a luma sample corresponding to a top-left sample of the current chroma video block; The scaling factor is checking the availability of each of one or more neighboring luma blocks of the video unit, each of the one or more neighboring luma blocks including at least one sample of the neighboring luma samples; determining whether to search for the neighboring luma samples of the video unit based on the availability of each of the one or more neighboring luma blocks; deriving the scaling factor based on the average luma variable calculated using the neighboring luma samples by a rounding-based averaging operation in response to the neighboring luma samples being available; deriving the scaling factor by setting the average luma variable to 1<<(bitDepth-1), in response to determining that one or more left-neighboring sample columns and one or more above-neighboring sample rows of the video unit are unavailable, where bitDepth is a bit depth of the video. A non-transitory computer-readable storage medium derived by
16. 1. A method for storing a video bitstream, the method comprising: determining a scaling operation to be applied to chroma residual samples of a current chroma video block; generating the bitstream by applying the scaling operation to the chroma residual samples; storing the bitstream on a non-transitory computer readable medium; Including, the scaling process scales the chroma residual samples based on a scaling factor before using them to reconstruct the current chroma video block; the scaling factor is derived based on an average luma variance calculated based on neighboring luma samples of a video unit of the video determined based on a luma sample corresponding to a top-left sample of the current chroma video block; The scaling factor is checking the availability of each of one or more neighboring luma blocks of the video unit, each of the one or more neighboring luma blocks including at least one sample of the neighboring luma samples; determining whether to search for the neighboring luma samples of the video unit based on the availability of each of the one or more neighboring luma blocks; deriving the scaling factor based on the average luma variable calculated using the neighboring luma samples by a rounding-based averaging operation in response to the neighboring luma samples being available; deriving the scaling factor by setting the average luma variable to 1<<(bitDepth-1), in response to determining that one or more left-neighboring sample columns and one or more above-neighboring sample rows of the video unit are unavailable, where bitDepth is a bit depth of the video. The method is derived by:
Citation Information
Patent Citations
Apparatus and method for scalable and multiview / 3d coding of video information
JP2016506683A
Method and apparatus for encoding and decoding pictures
JP2022532277A
Intra prediction techniques for video coding
WO2018132475A1