Method and apparatus for inheriting shared cross-component linear model and history table in video coding system
By adopting an inherited cross-component model with historical table design in the video encoding and decoding system, the problems of low cross-component prediction accuracy and low encoding and decoding efficiency in the prior art are solved, and more efficient video encoding and decoding are achieved.
Patent Information
- Application Number
- CN202380080068.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-10-27
- Publication Date
- 2025-06-27
AI Technical Summary
When existing video encoding and codec systems process cross-component prediction, it is difficult to effectively utilize historical table information, resulting in low prediction accuracy and codec efficiency.
Using an inherited cross-component model with a history table design, a list of prediction candidates is determined by receiving the first and second color blocks in the input data, including the cross-component prediction candidates inherited from the cross-component model history table. Based on the selected inherited model parameter set, the target model parameter set is determined and used for video encoding and decoding.
By effectively utilizing historical table information, the accuracy and codec efficiency of cross-component prediction are improved, and the overall performance of the video codec system is improved.
Smart Images

Figure CN120226353A_ABST
Abstract
Description
[0001]
Cross - reference
[0002] This invention is a non - provisional application of U.S. Provisional Patent Application No. 63 / 384,241, and claims its priority. The U.S. Provisional Patent Application was filed on November 18, 2022. This U.S. Provisional Patent Application is incorporated herein by reference in its entirety.
Technical Field
[0003] This invention relates to video coding and decoding systems. In particular, this invention relates to using a history table to inherit cross - component models in a video coding and decoding system.
Background Art
[0004] Versatile video coding (VVC) is the latest international video coding and decoding standard developed by the Joint Video Team (JVET) of the ITU - T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). This standard has been published as an ISO standard: ISO / IEC 23090 - 3:2021, Information technology - Coding representation of immersive media - Part 3: Versatile video coding, which was released in February 2021. VVC is developed based on its predecessor, High Efficiency Video Coding (HEVC), by adding more coding and decoding tools to improve coding and decoding efficiency and handle various types of video sources, including three - dimensional (3D) video signals.
[0005] Figure 1AAn exemplary intra / inter-frame video codec system incorporating loop processing is shown. For intra prediction, the prediction data is derived based on previously encoded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed at the encoder side, and motion compensation (MC) is performed based on the result of ME to provide prediction data derived from other pictures and motion data. Switch 114 selects either intra prediction 110 or inter prediction 112 and supplies the selected prediction data to adder 116 to form a prediction error, also known as a residual. The prediction error is then processed by a transform (T) 118, followed by quantization (Q) 120. The transformed and quantized residual is then encoded by an entropy encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed together with side information, such as motion and codec modes related to intra and inter prediction, and other information such as parameters related to loop filters applied to underlying image regions. The side information related to intra prediction 110, inter prediction 112, and loop filter 130 is provided to entropy encoder 122 as shown in Figure 1A When the inter prediction mode is used, the reference picture or pictures must also be reconstructed at the encoder side. Therefore, the transformed and quantized residual is processed by inverse quantization (IQ) 124 and inverse transformation (IT) 126 to recover the residual. The residual is then added back to the prediction data 136 at reconstruction (REC) 128 to reconstruct the video data. The reconstructed video data can be stored in a reference picture buffer 134 and used for prediction of other frames.
[0006] As Figure 1AAs shown, the input video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to the series of processes. Therefore, the loop filter 130 is typically applied to the reconstructed video data before it is stored in the reference picture buffer 134 to improve the video quality. For example, a deblocking filter (DF), a Sample Adaptive Offset (SAO), and an Adaptive Loop Filter (ALF) can be used. The loop filter information may need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to the entropy encoder 122 to be included in the bitstream. In Figure 1A the loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. Figure 1A The system in
[0007] is designed to show an exemplary structure of a typical video encoder. It may correspond to a High Efficiency Video Coding (HEVC) system, VP8, VP9, H.264, or VVC. Figure 1B As shown, the decoder can use the same or partially the same functional blocks as the encoder, except for the transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses an EntropyDecoder 140 to decode the video bitstream into quantized transform coefficients and the required coding and decoding information (such as loop filter information (ILPF information), intra prediction information, and inter prediction information). Intra prediction 150 at the decoder side does not need to perform a mode search. Instead, the decoder only needs to generate an intra prediction based on the intra prediction information received from the entropy decoder 140. In addition, for inter prediction, the decoder only needs to perform motion compensation (MC152) based on the inter prediction information received from the entropy decoder 140, without performing motion estimation.
[0008] According to VVC, the input picture is divided into non - overlapping square regions called Coding Tree Units (CTUs), similar to HEVC. Each CTU can be divided into one or more coding units (CUs) of smaller sizes. The generated CU partitions can be square or rectangular in shape. In addition, VVC divides the CTU into prediction units (PUs) as the units to which prediction processes, such as inter prediction, intra prediction, etc., are applied.
[0009] The VVC standard incorporates various new coding and decoding tools to further improve the coding and decoding efficiency beyond the HEVC standard. Some of the new tools relevant to the present invention are described below.
[0010] Using a tree structure to partition CTUs
[0011] In HEVC, a CTU is partitioned into CUs using a quaternary-tree (QT) structure, called a coding tree, to adapt to various local characteristics. The decision on whether to use inter (temporal) or intra (spatial) prediction to encode a picture region is made at the leaf CU level. Each leaf CU can be further partitioned into one, two, or four PUs according to the PU split type. Inside a PU, the same prediction process is applied, and the relevant information is transmitted to the decoder on a PU basis. After obtaining the residual block after applying the prediction process according to the PU split type, the leaf CU can be partitioned into transform units (TUs) according to another quaternary-tree structure similar to the coding tree of the CU. A key feature of the HEVC structure is that it has a multiple partitioning concept that includes CUs, PUs, and TUs.
[0012] In VVC, a quaternary tree with a multi-type tree structure nested with binary and ternary splits replaces the concept of multiple partition unit types, that is, it removes the separation of the CU, PU, and TU concepts, unless the size of the CU is too large to exceed the maximum transform length, and supports more flexibility in the CU partition shape. In the coding tree structure, a CU can be square or rectangular in shape. A coding tree unit (CTU) is first partitioned by a quaternary tree (also known as a quadtree) structure. Then the quaternary tree leaf nodes can be further partitioned by a multi-type tree structure. As Figure 2 shown, there are four split types in the multi-type tree structure, vertical binary split (SPLIT_BT_VER 210), horizontal binary split (SPLIT_BT_HOR 220), vertical ternary split (SPLIT_TT_VER 230), and horizontal ternary split (SPLIT_TT_HOR 240). The multi-type tree leaf nodes are called coding units (CUs), and this split is used for prediction and transform processing without any further splitting unless the CU is too large to exceed the maximum transform length. This means that in most cases, the CUs, PUs, and TUs have the same block size in the quaternary tree nested multi-type tree coding block structure. Exceptions occur when the maximum supported transform length is less than the width or height of the color component of the CU.
[0013] Figure 3Shows the signaling mechanism for the partition splitting information in the quadtree and nested multi-type tree codec tree structures. The codec tree unit (CTU) is regarded as the root node of the quadtree and is first partitioned by the quadtree structure. Each quadtree leaf node (allowed when large enough) is then further partitioned by the multi-type tree structure. In the quadtree and nested multi-type tree codec tree structures, for each CU node, first a flag (split_cu_flag) is signaled to indicate whether the node is further partitioned. If the current CU node is a quadtree CU node, a second flag (split_qt_flag) is signaled to indicate whether it is in the QT partition or MTT partition mode. When the node is partitioned in the MTT partition mode, a third flag (mtt_split_cu_vertical_flag) is signaled to indicate the split direction, and then a fourth flag (mtt_split_cu_binary_flag) is signaled to indicate whether the split is a binary split or a ternary split. Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree split mode (MttSplitMode) of the CU is derived, as shown in Table 1.
[0014] Table 1 – Derivation of MttSplitMode based on multi-type tree syntax elements
[0015] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1
[0016] Figure 4 Shows a CTU divided into multiple CUs, with a quadtree and nested multi-type tree codec block structure, where the bold block edges represent quadtree partitions and the remaining edges represent multi-type tree partitions. The quadtree and nested multi-type tree partitions provide a content-adaptive codec tree structure composed of CUs. The size of the CU can be as large as the CTU or as small as 4×4 luminance sample units. For the 4:2:0 chroma format, the maximum chroma CB size is 64×64, and the minimum size chroma CB consists of 16 chroma samples.
[0017] In VVC, the maximum supported luminance transform size is 64×64, and the maximum supported chroma transform size is 32×32. When the width or height of the CB is greater than the maximum transform width or height, the CB is automatically split in the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0018] The following parameters are defined for the quadtree and nested multi-type tree codec tree scheme. These parameters are specified by SPS syntax elements and can be further refined by picture header syntax elements.
[0019] – CTU size: the size of the quadtree root node
[0020] –MinQTSize: Minimum allowable quadtree leaf node size
[0021] –MaxBtSize: Maximum allowable binary tree root node size
[0022] –MaxTtSize: Maximum allowable ternary tree root node size
[0023] –MaxMttDepth: Maximum allowable multi-type tree split level depth starting from quadtree leaf nodes
[0024] –MinCbSize: Minimum allowable decoded block node size
[0025] In an example of the quadtree and nested multi-type tree codec tree structure, the CTU size is set to 128×128 luma samples, with two corresponding 64×64 blocks of 4:2:0 chroma samples. MinQTSize is set to 16×16, MaxBtSize is set to 128×128, MaxTtSize is set to 64×64, MinCbsize (width and height) is set to 4×4, and MaxMttDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf QT node is 128×128, since its size exceeds MaxBtSize and MaxTtSize (i.e., 64×64), it will not be further split by the binary tree. Otherwise, the leaf quadtree node can be further partitioned by the multi-type tree. Thus, the quadtree leaf node is also the root node of the multi-type tree, and its multi-type tree depth (mttDepth) is 0. When the multi-type tree depth reaches MaxMttDepth (i.e., 4), further splitting is no longer considered. When the width of the multi-type tree node equals MinCbsize, further horizontal splitting is no longer considered. Similarly, when the height of the multi-type tree node equals MinCbsize, further vertical splitting is no longer considered.
[0026] In VVC, the codec tree scheme supports separate block tree structures for luma and chroma. For P and B slices, the luma and chroma CTBs in a CTU must share the same codec tree structure. However, for I slices, the luma and chroma can have separate block tree structures. When the separate block tree mode is applied, the luma CTB is partitioned into CUs by one codec tree structure, and the chroma CTB is partitioned into chroma CUs by another codec tree structure. This means that the CUs in I slices may consist of decoded blocks of the luma component or decoded blocks of both chroma components, while the CUs in P or B slices always consist of decoded blocks of all three color components, unless the video is monochrome.
[0027] Virtual Pipeline Data Units (VPDUs)
[0028] Virtual Pipeline Data Units (VPDUs) are defined as non - overlapping units in a picture. In a hardware decoder, multiple pipeline stages process consecutive VPDUs simultaneously. The VPDU size is roughly proportional to the buffer size in most pipeline stages, so it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) and binary tree (BT) partitioning may cause an increase in VPDU size.
[0029] To keep the VPDU size at 64x64 luma samples, the following canonical partitioning restrictions (with syntax signaling modifications) are applied in VTM, as Figure 5 shown:
[0030] For a CU with width or height equal to 128, TT splitting is not allowed (as shown by "X" in Figure 5 ).
[0031] For a 128xN CU, where N ≤ 64 (i.e., width equals 128 and height is less than 128), horizontal BT is not allowed.
[0032] For an Nx128 CU, where N ≤ 64 (i.e., height equals 128 and width is less than 128), vertical BT is not allowed. In Figure 5 the luma block size is 128x128. The dashed lines indicate a block size of 64x64. According to the above constraints, Figure 5 the non - allowed partitioning examples in various examples (510 - 580) in are indicated by "X".
[0033] Intra - frame Chroma Partitioning and Prediction Limitations
[0034] In typical hardware video encoders and decoders, due to the sample - processing data dependencies between adjacent blocks within a frame, when the image has smaller intra - frame blocks, the processing throughput decreases. The generation of predictors for intra - frame blocks requires reconstructed samples from the top and left boundaries of adjacent blocks. Therefore, intra - frame prediction must be processed block - by - block sequentially.
[0035] In HEVC, the smallest intra CU is an 8x8 luma sample. The luma component of the smallest intra CU can be further divided into four 4x4 luma intra prediction units (PUs), but the chroma component of the smallest intra CU cannot be further divided. Therefore, the worst hardware processing throughput occurs when processing 4x4 chroma intra blocks or 4x4 luma intra blocks. In VVC, to improve the worst throughput, by restricting the splitting of chroma intra CBs, chroma intra CBs smaller than 16 chroma samples (sizes 2x2, 4x2, and 2x4) and chroma intra CBs with a width less than 4 chroma samples (size 2xN) are prohibited.
[0036] In a single codec tree, a smallest chroma intraprediction unit (SCIPU) is defined, where the chroma block size of the codec tree node is greater than or equal to 16 chroma samples and at least one sub-luma block is less than 64 luma samples, or the chroma block size of the codec tree node is not 2xN and at least one sub-luma block is 4xN luma samples. It is required that in each SCIPU, all CBs are inter or all CBs are non-inter, i.e., intra or intra block copy (IBC). In the case of a non-inter SCIPU, it is also required that the chroma of the non-inter SCIPU cannot be further divided, while the luma of the SCIPU can be further divided. In this way, small chroma intra CBs smaller than 16 chroma samples or of size 2xN are removed. Additionally, in the case of a non-inter SCIPU, chroma scaling is not applied. There is no additional syntax signal, and whether the SCIPU is non-inter can be deduced from the prediction mode of the first luma CB in the SCIPU. If the current slice is an I slice or the current SCIPU has a 4x4 luma partition after being further divided once (since inter 4x4 is not allowed in VVC), the type of the SCIPU is inferred as non-inter; otherwise, the type of the SCIPU (inter or non-inter) is indicated by a flag before parsing the CUs in the SCIPU.
[0037] For the dual tree in an intra picture, 2xN intra chroma blocks are removed by disabling the vertical binary and vertical ternary splitting of 4xN and 8xN chroma partitions. Small chroma blocks of sizes 2x2, 4x2, and 2x4 are also removed by the splitting restrictions.
[0038] Furthermore, considering that the image width and height are multiples of max(8,MinCbSizeY), restrictions on the image size are considered to avoid 2x2 / 2x4 / 4x2 / 2xN intra chroma blocks at the image corners.
[0039] Intra Mode Coding with 67 Intra Prediction Modes
[0040] To capture any edge direction presented in natural videos, the number of directional intra modes in VVC is extended from 33 used in HEVC to 65. The new direction modes not available in HEVC are shown by dashed arrows in Figure 6 and the Planar and DC modes remain unchanged. These denser directional intra prediction modes apply to all block sizes and are applicable to both luma and chroma intra prediction.
[0041] In VVC, several traditional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes for non-square blocks.
[0042] In HEVC, each intra-coded block has a square shape with the length of each side being a power of 2. Therefore, no division operation is required to generate an intra predictor using the DC mode. In VVC, blocks can have a rectangular shape and, in general, one division operation per block is required. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks.
[0043] To keep the complexity of the most probable mode (MPM) list generation low, an intra mode coding method with 6 MPMs is used considering two available neighboring intra modes. Three aspects are considered when constructing the MPM list:
[0044] Default intra mode
[0045] Neighboring intra modes
[0046] Deriving intra modes.
[0047] Regardless of whether the MRL and ISP coding tools are applied, a unified 6-MPM list is used for intra blocks. The MPM list is constructed based on the intra modes of the left and above neighboring blocks. Assuming the mode on the left is denoted as Left and the mode of the above block is denoted as Above, the unified MPM list is constructed as follows:
[0048] – When the neighboring blocks are not available, their intra modes are default set to Planar.
[0049] – If both the Left and Above modes are non-angular modes:
[0050] ○ MPM list → {Planar, DC, V, H, V - 4, V + 4}
[0051] – If one of the Left and Above modes is an angular mode and the other is a non-angular mode:
[0052] ○ Set a mode Max to the larger mode between Left and Above
[0053] ○ MPM list → {Planar, Max, Max - 1, Max + 1, Max–2, Max + 2}
[0054] – If both Left and Above are angular modes and they are different:
[0055] ○ Set a mode Max to the larger mode between Left and Above
[0056] ○ If Max – Min equals 1:
[0057] ■ MPM list → {Planar, Left, Above, Min–1, Max + 1, Min–2}
[0058] ○ Otherwise, if Max – Min is greater than or equal to 62:
[0059] ■ MPM list → {Planar, Left, Above, Min + 1, Max – 1, Min + 2}
[0060] ○ Otherwise, if Max – Min equals 2:
[0061] ■ MPM list → {Planar, Left, Above, Min + 1, Min–1, Max + 1}
[0062] ○ Otherwise:
[0063] ■ MPM list → {Planar, Left, Above, Min–1, Min + 1, Max – 1}
[0064] ○ If both Left and Above are angular modes and they are the same:
[0065] ■ MPM list → {Planar, Left, Left - 1, Left + 1, Left–2, Left + 2}
[0066] In addition, the first bin of the MPM index codeword is CABAC context - encoded. A total of three contexts are used, corresponding to whether the MRL is enabled, the ISP is enabled, or it is a normal intra - frame block within the current frame.
[0067] During the 6 - MPM list generation process, pruning is used to remove duplicate modes so that the MPM list contains only unique modes. For the entropy encoding and decoding of 61 non - MPM modes, truncated binary code (TBC) is used.
[0068] Wide-Angle Intra Prediction for Non-Square Blocks
[0069] The traditional angular intra prediction directions are defined as clockwise from 45 degrees to -135 degrees. In VVC, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original mode indices, which are remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode encoding and decoding methods remain unchanged.
[0070] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 7A and Figure 7B shown.
[0071] The number of replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2.
[0072] Table 2 – Intra Prediction Modes Replaced by Wide-Angle Modes
[0073]
[0074] In VVC, 4:2:2 and 4:4:4 chroma formats as well as 4:2:0 are supported. The chroma derived mode (DM) derivation table for the 4:2:2 chroma format initially ported from HEVC extends the number of entries from 35 to 67 to be consistent with the extension of the intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luminance intra prediction mode range from 2 to 5 is mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the mapping table entries to more accurately transform the prediction angles of chroma blocks.
[0075] Cross-Component Linear Model (CCLM) Prediction
[0076] To reduce cross-component redundancy, the cross-component linear model (CCLM) prediction mode is used in VVC, where chroma samples are predicted using a linear model based on the reconstructed luminance samples of the same CU, as follows:
[0077] pred C (i,j) = α·rec L ′(i,j) + β (1)
[0078] where pred C (i,j) represents the predicted chroma samples in the CU, and rec L′(i,j) represents the downsampled reconstructed luma samples of the same CU.
[0079] The CCLM parameters (α and β) are derived from at most four neighboring chroma samples and their corresponding downsampled luma samples. Assuming the size of the current chroma block is W×H, then W' and H' are set to
[0080] – When applying the LM_LA mode, W' = W, H' = H;
[0081] – When applying the LM_A mode, W' = W + H;
[0082] – When applying the LM_L mode, H' = H + W.
[0083] The above neighboring positions are denoted as S[0, -1]…S[W' - 1, -1], and the left neighboring positions are denoted as S[-1, 0]…S[-1, H' - 1]. Then the four samples are selected as
[0084] When applying the LM mode and both the above and left neighboring samples are available, select S[W' / 4, -1], S[3*W' / 4, -1], S[-1, H' / 4], S[-1, 3*H' / 4];
[0085] When applying the LM-A mode or only the above neighboring samples are available, select S[W' / 8, -1], S[3*W' / 8, -1], S[5*W' / 8, -1], S[7*W' / 8, -1];
[0086] When applying the LM-L mode or only the left neighboring samples are available, select S[-1, H' / 8], S[-1, 3*H' / 8], S[-1, 5*H' / 8], S[-1, 7*H' / 8].
[0087] The four neighboring luma samples at the selected positions are downsampled and compared four times to find two larger values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B . Their corresponding chroma sample values are denoted as y 0 A 、y 1 A 、y 0 B and y 1 B . Then X a 、X b 、Y a and Y bExport as:
[0088] X a = (x 0 A + x 1 A + 1) >> 1;
[0089] X b = (x 0 B + x 1 B + 1) >> 1;
[0090] Y a = (y 0 A + y 1 A + 1) >> 1;
[0091] Y b = (y 0 B + y 1 B + 1) >> 1 (2)
[0092] Finally, the linear model parameters α and β are obtained according to the following equations.
[0093]
[0094] β = Y b - α · X b (4)
[0095] Figure 8 Examples of the positions of the left and upper samples and the current block sample involved in the LM_LA mode are shown. Figure 8 The relative sample positions of the N×N chrominance block 810, the corresponding 2N×2N luminance block 820, and its neighboring samples (shown as solid circles) are shown.
[0096] The division operation for calculating the parameter α is implemented through a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented exponentially. For example, diff is approximated with 4 significant bits and an exponent. Thus, the table of 1 / diff is reduced to 16 elements, corresponding to 16 values of the significant part, as follows:
[0097] DivTable[] = {0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} (5)
[0099] This helps reduce the computational complexity and the memory size of the table required for storage.
[0100] In addition to the above templates and the left template that can be used together to calculate the linear model coefficients, they can also be alternately used in two other LM modes, called LM_A and LM_L modes.
[0101] In the LM_A mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to (W + H) samples. In the LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H + W) samples.
[0102] In the LM_LA mode, the left and upper templates are used to calculate the linear model coefficients.
[0103] To match the chrominance sample positions of the 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2-to-1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filter is specified by the SPS level flag. These two downsampling filters are as follows, corresponding to "type-0" and "type-2" content respectively.
[0104] Rec L ′(i,j) = [rec L (2i - 1, 2j - 1) + 2·rec L (2i, 2j - 1) + rec L (2i + 1, 2j - 1) + rec L (2i - 1, 2j) + 2·rec L (2i, 2j) + rec L (2i + 1, 2j) + 4] >> 3 (6)
[0105] Rec L ′(i,j) = rec L (2i, 2j - 1) + rec L (2i - 1, 2j) + 4·rec L (2i, 2j) + rec L (2i + 1, 2j) + rec L (2i, 2j + 1) + 4] >> 3 (7)
[0107] Note that when the upper reference line is at the CTU boundary, only one luma line (the general line buffer in intra prediction) is used to produce the downsampled luma samples.
[0108] This parameter calculation is performed as part of the decoding process, not just the encoder search operation. Therefore, the α and β values are not passed to the decoder using syntax.
[0109] For chrominance intra mode coding and decoding, a total of 8 intra modes are allowed for chrominance intra mode coding and decoding. These modes include five traditional intra modes and three cross-component linear model modes (LM_LA, LM_A, and LM_L). The chrominance mode signaling and derivation process are shown in Table 3. The chrominance mode coding and decoding directly depend on the intra prediction mode of the corresponding luma block. Since the separate block splitting structure of the luma and chrominance components is enabled in I slices, one chrominance block may correspond to multiple luma blocks. Therefore, for the chrominance DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chrominance block is directly inherited.
[0110] Table 3 Deriving Chrominance Prediction Modes from Luma Modes when CCLM is Enabled
[0111]
[0112] Regardless of the value of sps_cclm_enabled_flag, a single binarization table is used, as shown in Table 4.
[0113] Table 4 - Unified Binarization Table for Chrominance Prediction Modes
[0114]
[0115]
[0116] In Table 4, the first binary indication is whether it is a regular mode (0) or a CCLM mode (1). If it is an LM mode, the next binary indication is whether it is LM_LA (0) or others. If it is not LM_LA, the next binary indication is whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first binary of the binarization table of the corresponding intra_chroma_pred_mode can be discarded before entropy coding. In other words, the first binary is inferred as 0 and thus not encoded. This single binarization table is used for the cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binaries in Table 4 are context-encoded using their own context models, and the remaining binaries are bypass-encoded.
[0117] In addition, to reduce the luma-chrominance delay in the dual tree, when the 64x64 luma coding and decoding tree node is not split (and ISP is not used for the 64x64 CU) or QT split, the chrominance CUs in the 32x32 / 32x16 chrominance coding and decoding tree nodes are allowed to use CCLM as follows:
[0118] If the 32x32 chrominance node is not split or split QT split, all chrominance CUs in the 32x32 node can use CCLM
[0119] If the 32x32 chroma nodes are split horizontally by BT and the 32x16 child nodes are not split or split vertically by BT, then all chroma CUs in the 32x16 chroma nodes can use CCLM.
[0120] Under all other luma and chroma codec tree splitting conditions, chroma CUs are not allowed to use CCLM.
[0121] Multiple Model CCLM (MMLM)
[0122] In JEM (J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Video Exploration Team (JVET), Jul. 2017), the Multiple Model CCLM mode (MMLM) was proposed for predicting chroma samples from luma samples using two models. In MMLM, the neighboring luma samples and neighboring chroma samples of the current block are divided into two groups, and each group is used as a training set to derive a linear model (i.e., specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified according to the same rule as the classification of the neighboring luma samples.
[0123] Figure 9 An example of dividing neighboring samples into two groups is shown. The Threshold is calculated as the average of the neighboring reconstructed luma samples. The neighboring sample Rec′ L [x, y] <= Threshold is classified into Group 1; while the neighboring sample Rec′ L [x, y] > Threshold is classified into Group 2.
[0124]
[0125] Therefore, MMLM uses two models according to the sample level of neighboring samples.
[0126] Slope Adjustment of CCLM
[0127] CCLM uses a model with two parameters to map luma values to chroma values, as Figure 10A shown. The slope parameter "a" and the bias parameter "b" define the following mapping:
[0128] chromaVal = a * lumaVal + b
[0129] The adjustment “u” of the slope parameter is signaled to update the model to the following form, as Figure 10B shown:
[0130] chromaVal = a’ * lumaVal + b’
[0131] where
[0132] a’ = a + u,
[0133] b’ = b - u * y r .
[0134] With this selection, the mapping function is tilted or rotated around the point of the luma value y r . The average value of the reference luma samples used for model creation is taken as y r to provide a meaningful modification to the model. Figure 10A And 10B illustrates the process.
[0135] Implementation of CCLM Slope Adjustment
[0136] The slope adjustment parameter is provided as an integer between -4 and 4 and is signaled in the bitstream. The unit of the slope adjustment parameter is the (1 / 8)-th chroma sample value per luma sample value (for 10-bit content).
[0137] The adjustment applies to CCLM models that use the reference samples above and to the left of the block (such as “LM_CHROMA_IDX” and “MMLM_CHROMA_IDX”), but not to the “unilateral” mode. This selection is based on the consideration of codec efficiency and complexity trade-off. “LM_CHROMA_IDX” and “MMLM_CHROMA_IDX” refer to CCLM_LT and MMLM_LT in the present invention. The “unilateral” mode refers to CCLM_L, CCLM_T, MMLM_L, and MMLM_T in the present invention.
[0138] When the slope adjustment is applied to a multi-mode CCLM model, two models can be adjusted, so at most two slope updates are signaled for a single chroma block.
[0139] Encoder Method of CCLM Slope Adjustment
[0140] The proposed encoder method performs a SATD (Sum of Absolute Transformed Differences)-based search to find the best slope update value for Cr and a similar SATD-based search to find the best slope update value for Cb. If one of the results is a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based Cr update, SATD-based Cb update) is included in the RD (Rate-Distortion) check list of the TU.
[0141] Convolutional cross-component model (CCCM) - single model and multi-model
[0142] In the CCCM, a convolutional model is applied to improve chrominance prediction performance. The convolutional model has a 7-tap filter, which consists of a 5-tap plus sign-shaped spatial component, a non-linear term, and a bias term. The input to the spatial 5-tap component of the filter includes the central (C) luminance sample co-located with the chrominance sample to be predicted and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as Figure 11 shown.
[0143] The non-linear term (denoted as P) is expressed as the square of the central luminance sample C and scaled to the sample value range of the content:
[0144] P = (C * C + midVal) >> bitDepth.
[0145] For example, for 10-bit content, the non-linear term is calculated as:
[0146] P = (C * C + 512) >> 10
[0147] The bias term (denoted as B) represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the middle chrominance value (512 for 10-bit content).
[0148] The output of the filter is calculated as the convolution between the filter coefficients c i and the input values, and clipped to the range of valid chrominance samples:
[0149] predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B
[0150] The filter coefficients c i are calculated by minimizing the MSE between the predicted and reconstructed chrominance samples in the reference region. Figure 12 An example of the reference region is illustrated. The region consists of 6 rows of chrominance samples above and to the left of the PU. The reference region is extended one PU width to the right and one PU height downwards. The region is adjusted to only include available samples. The extension of the region (denoted as the "extended region") is used to support Figure 11 the "side samples" of the plus sign-shaped spatial filter in
[0151] MSE minimization is performed by computing the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and the chrominance output. The autocorrelation matrix is subjected to LDL decomposition, and the final filter coefficients are computed using back substitution. This process generally follows the computation of the ALF filter coefficients in ECM, but LDL decomposition is chosen instead of Cholesky decomposition to avoid using square root operations.
[0152] In addition, similar to CCLM, CCCM has an option to use a single model or a multi-model variant. The multi-model variant uses two models, one for samples above the average luminance reference value and the other for the remaining samples (following the spirit of the CCLM design). For PUs for which at least 128 reference samples are available, the multi-model CCCM mode can be selected.
[0153] Gradient Linear Model (GLM)
[0154] Compared with CCLM, instead of using downsampled luminance values, GLM utilizes the luminance sample gradients to derive the linear model. Specifically, when applying GLM, the input to the CCLM process, i.e., the downsampled luminance samples L, is replaced by the luminance sample gradients G. The other parts of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged.
[0155] C = α·G + β
[0156] For signal transmission, when the CCLM mode is enabled for the current CU, two flags are transmitted for the Cb and Cr components respectively to indicate whether GLM is enabled for each component, or one GLM flag is transmitted for the Cb and Cr components, using a shared GLM index. If GLM is enabled for a component, a further syntax element is transmitted to select one of the multiple gradient filters ( Figure 13 from 1310 - 1340 in ) for gradient calculation. GLM can be used in combination with the existing CCLM by transmitting an additional flag in the bitstream. When applying this combination, the filter coefficients for deriving the input luminance samples of the linear model are computed as the combination of the selected gradient filter of GLM and the downsampling filter of CCLM.
[0157] Spatial candidate derivation
[0158] The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Up to four merge candidates (B 0, A 0, B1 and A1) of the current CU 1410 are selected from the candidates at the positions depicted in Figure 14 The derivation order is B 0, A 0, B1, A1 and B2. Location B2 is considered only when one or more neighboring CUs at locations B0, A0, B1, A1 are unavailable (e.g., belong to another slice or tile) or are intra-coded. After a candidate at location A1 is added, the addition of the remaining candidates is subject to a redundancy check to ensure that candidates with the same motion information are excluded from the list, thereby improving the coding and decoding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs linked by the arrows in Figure 15 are considered, and a candidate is added to the list only if the corresponding candidate used for the redundancy check does not have the same motion information.
[0159] Temporal candidate derivation
[0160] In this step, only one candidate is added to the list. In particular, in the derivation of the temporal merge candidate for the current CU 1610, a scaled motion vector is derived based on the co-located CU 1620 belonging to the co-located reference picture as shown in Figure 16 . The reference picture list and reference index used to derive the co-located CU are explicitly transmitted in the slice header. The scaled motion vector 1630 of the temporal merge candidate is scaled from the motion vector 1640 of the co-located CU using the POC (Picture Order Count) distances tb and td as shown by the dashed line in Figure 16 , where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set to zero.
[0161] The location of the temporal candidate is selected between candidates C0 and C1 as shown in Figure 17 . If the CU at location C0 is unavailable, is intra-coded, or is outside the current CTU row, then location C1 is used. Otherwise, location C0 is used in the derivation of the temporal merge candidate.
[0162] Non-adjacent spatial candidates
[0163] During the development of the VVC standard, a codec tool called Non-Adjacent Motion Vector Prediction (NAMVP for short) was proposed. See JVET-L0399 (Yu Han et al., "CE4.4.6: Improvement on Merge / Skip mode", Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC29 / WG 11, 12th meeting: Macau, CN, October 3 - 12, 2018, Document: JVET-L0399). According to the NAMVP technology, non-adjacent spatial merge candidates are inserted after the TMVP (i.e., temporal MVP) in the regular merge candidate list. The mode of spatial merge candidates is as Figure 18 shown. The distance between the non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block. In Figure 18 , each small square corresponds to a NAMVP candidate, and the candidates are sorted according to the distance (as shown by the numbers inside the squares). The row buffer limit does not apply. In other words, NAMVP candidates far from the current block may need to be stored, which may require a large buffer.
[0164] In the present invention, methods and apparatuses for inheriting a history table design of a shared cross-component model are disclosed to improve performance.
Summary of the Invention
[0165] A method and apparatus for video coding and decoding using an inherited cross-component model with a history table design are disclosed. According to the method, input data related to a current block is received, including a first color block and a second color block, where the input data includes pixel data to be encoded at an encoder end or encoded data related to the current block to be decoded at a decoder end. A prediction candidate list is determined, including one or more cross-component prediction candidates inherited from a cross-component model history table, where the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list according to a predefined order. Based on an inherited model parameter set related to a target inherited prediction model selected from the prediction candidate list, a target model parameter set related to the target inherited prediction model is determined. The second color block is encoded or decoded using prediction data, where the prediction data includes cross-color prediction generated by applying the target inherited prediction model and the target model parameter set to a reconstructed first color block.
[0166] In one embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the beginning to the end of the cross-component model history table. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the end to the beginning of the cross-component model history table. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from a predefined position in the cross-component model history table to the end or the beginning of the cross-component model history table. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the cross-component model history table in a staggered manner.
[0167] According to another method, a prediction candidate list is determined, including one or more cross-component prediction candidates inherited from a cross-component model history table, where the cross-component model history table is reset at a specific point related to an image region including non-CTUs.
[0168] In one embodiment, the image region corresponds to the current picture, slice, or tile. In one embodiment, the image region corresponds to every M CTU rows or every N CTUs, where M and N are positive integers. In one embodiment, the specific point related to the image region corresponds to the beginning or the end of the image region.
[0169] According to another method, a prediction candidate list is determined, including one or more cross-component prediction candidates inherited from multiple cross-component model history tables.
[0170] In one embodiment, each image is divided into multiple regions, and a cross-component model history table is maintained for each of the multiple regions. In one embodiment, the sizes of the multiple regions are predefined. In another embodiment, the sizes of the multiple regions correspond to X times Y CTUs, where X and Y are positive integers. In one embodiment, each image is divided into N regions, and the multiple cross-component model history tables correspond to N history tables, where N is an integer greater than 1.
[0171] In one embodiment, cross-component model history table 0 is used to store all previous cross-component models. In one embodiment, cross-component model history table 0 is always updated during the encoding or decoding process. In another embodiment, cross-component model history table 0 and one additional history table among the multiple cross-component model history tables are updated during the encoding or decoding process. In another embodiment, the additional history table is determined according to the current position of the current block.
[0172] In one embodiment, at least two cross-component model history tables are updated at different frequencies. In another embodiment, multiple cross-component model history tables are used to store different types of cross-component models. In another embodiment, different types of cross-component models correspond to single models and multi-models, gradient models and non-gradient models, or simple linear models and complex models. In another embodiment, different types of cross-component models correspond to different reconstructed luminance intensities or different reconstructed chrominance intensities.
[0173] In one embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the beginning to the end of one cross-component model history table, and then inserted from the next cross-component model history table in the same order or the reverse order. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the end to the beginning of one cross-component model history table, and then inserted from the next cross-component model history table in the same order or the reverse order. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from a predefined position to the end or the beginning of one cross-component model history table, and then inserted from the next cross-component model history table in the same order or the reverse order. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from one cross-component model history table in a staggered manner, and then inserted from the next cross-component model history table in the same order or the reverse order. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the beginning to the end of each of multiple cross-component model history tables. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the end to the beginning of each of multiple cross-component model history tables. In another embodiment, the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from a predefined position to the end or the beginning of each of multiple cross-component model history tables. In another embodiment, only a subset of the multiple cross-component model history tables, whose corresponding regions are close to the current region containing the current block, is used to create the prediction candidate list.
[0174] In one embodiment, when one or more cross-component prediction candidates inherited from multiple cross-component model history tables are used to create the prediction candidate list, the range of selecting non-adjacent candidates is narrowed. In another embodiment, the range of selecting non-adjacent candidates is narrowed by measuring the distance from the upper left position of the current block to the target candidate position, and then excluding the cases where the distance between the target candidate and the current block is greater than a predefined threshold.
BRIEF DESCRIPTION OF THE DRAWINGS
[0175] Figure 1A An exemplary adaptive Inter / Intra video codec system including loop processing is illustrated.
[0176] Figure 1B illustrates Figure 1A the corresponding decoder of the encoder in
[0177] Figure 2 illustrates examples of multi-type tree structures corresponding to vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR).
[0178] Figure 3 illustrates an example of a split information signaling mechanism for a multi-type tree encoding and decoding tree structure nested in a quadtree.
[0179] Figure 4 Shows an example of a CTU divided into multiple CUs using a quadtree and a nested multi-type tree encoding and decoding block structure, where the bold block edges represent quadtree splitting and the remaining edges represent multi-type tree splitting.
[0180] Figure 5 Shows some examples of prohibited TT splitting when the width or height of a luma encoding and decoding block is greater than 64.
[0181] Figure 6 Shows the intra prediction modes adopted by the VVC video coding standard.
[0182] Figure 7A -B illustrates an example of wide-angle intra prediction, Figure 7A for a block with a width greater than the height, Figure 7B for a block with a height greater than the width.
[0183] Figure 8 Shows an example of the positions of the left and above samples and the current block samples involved in the LM_LA mode.
[0184] Figure 9 illustrates an example of classifying neighboring samples into two groups.
[0185] Figure 10A illustrates an example of the CCLM model.
[0186] Figure 10B illustrates an example of the effect of the slope adjustment parameter "u" for model update.
[0187] Figure 11 illustrates an example of the spatial part of a convolutional filter.
[0188] Figure 12 illustrates an example of a reference region for an extended region used to derive filter coefficients.
[0189] Figure 13 Illustrates 16 gradient patterns of the Gradient Linear Model (GLM).
[0190] Figure 14 Illustrates examples of neighboring blocks for deriving VVC spatial merge candidates.
[0191] Figure 15 Illustrates examples of possible candidate pairs considering redundancy checking in VVC.
[0192] Figure 16 Illustrates an example of temporal candidate derivation, where scaled motion vectors are derived based on the POC (Picture Order Count) distance.
[0193] Figure 17 Illustrates the positions of temporal candidates selected between candidates C0 and C1.
[0194] Figure 18 Illustrates exemplary patterns of non - adjacent spatial merge candidates.
[0195] Figure 19 Illustrates an example of inheriting temporal neighborhood model parameters.
[0196] Figure 20 Illustrates two search patterns for inheriting non - adjacent spatial neighborhood models.
[0197] Figure 21 Illustrates examples of multiple history tables for storing cross - component models.
[0198] Figure 22 Illustrates an example of a neighboring template for calculating model error.
[0199] Figure 23 Illustrates an example of a neighboring template for calculating model error.
[0200] Figure 24 Illustrates a flowchart of an exemplary video coding and decoding system that includes inheritance of a shared cross - component model using a history table with a predefined insertion order, according to an embodiment of the present invention.
[0201] Figure 25 Illustrates a flowchart of an exemplary video coding and decoding system that includes inheritance of a shared cross - component model using a history table with a specific reset point, according to an embodiment of the present invention.
[0202] Figure 26 Illustrates a flowchart of an exemplary video coding and decoding system that includes inheritance of a shared cross - component model using multiple history tables, according to an embodiment of the present invention.
Detailed Implementation Modes
[0203] The components of the present invention, as described and shown in the figures, can be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of the embodiments of the systems and methods of the present invention, as shown in the figures, is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the invention. References throughout the specification to "an embodiment", "one embodiment", or similar language mean that a particular feature, structure, or characteristic associated with that embodiment may be included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily all refer to the same embodiment.
[0204] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention may be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the invention. The embodiments of the present invention will be best understood by reference to the accompanying drawings, in which like parts are marked with like numerals throughout. The following description is only exemplary and merely illustrates certain selected embodiments of the apparatus and methods consistent with the invention claimed herein.
[0205] To improve the prediction accuracy or coding / decoding performance of cross-component prediction, various schemes related to the inherited cross-component model are disclosed.
[0206] Guided parameter set for refining the cross-component model parameters
[0207] According to this method, the guided parameter set is used to refine the derived model parameters through a specified CCLM mode. For example, the guided parameter set is explicitly signaled in the bitstream, and after the model parameters are derived, the guided parameter set is added to the derived model parameters as the final model parameters. The guided parameter set includes at least one differential scaling parameter (dA), one differential offset parameter (dB), and one differential shift parameter (dS). For example, Equation (1) can be rewritten as:
[0208] pred C (i,j) = ((α′·rec L ′(i,j)) >> s) + β,
[0209] If dA is signalled, the final prediction is:
[0210] pred C (i,j) = (((α′ + dA) · rec L ′(i,j)) >> s) + β。
[0211] Similarly, if dB is signaled, the final prediction is:
[0212] pred C (i,j) = ((α′ · rec L ′(i,j)) >> s) + (β + dB)。
[0213] If dS is signaled, the final prediction is:
[0214] pred C (i,j) = ((α′ · rec L ′(i,j)) >> (s + dS)) + β。
[0215] If dA and dB are signaled, the final prediction is:
[0216] pred C (i,j) = (((α) + dA) · rec L ′(i,j)) >> s) + (β + dB)。
[0217] The set of guiding parameters can be signaled by color component. For example, one set of guiding parameters is signaled for the Cb component and another set of guiding parameters is signaled for the Cr component. Alternatively, one set of guiding parameters can be signaled and shared between color components. The signaled dA and dB can be positive or negative. When dA is signaled, a bin is signaled to indicate the sign of dA. Similarly, when dB is signaled, a bin is signaled to indicate the sign of dB.
[0218] For another embodiment, dA and dB can be the LSB (Least Significant Bit) part of the final scaling and offset parameters. For example, if m bits are required to represent the final scaling parameter, dA is the LSB part of the final scaling parameter and n bits (m > n) are used to represent dA, where the MSB part (m - n bits) of the final scaling parameter is implicitly derived. In other words, for the final scaling parameter, the MSB part of the final scaling parameter is taken from the MSB part of α′ and the LSB part of the final scaling parameter comes from the signaled dA. Similarly, if p bits are required to represent the final offset parameter, dB is the LSB part of the final offset parameter and q bits (p > q) are used to represent dB, where the MSB part (p - q bits) of the final offset parameter is implicitly derived. In other words, for the final offset parameter, the MSB part of the final offset parameter is taken from the MSB part of β and the LSB part of the final offset parameter comes from the signaled dB.
[0219] For another implementation, if dA is transmitted, dB can be implicitly derived from the average of neighboring (e.g., L-shaped) reconstructed samples. For example, in VVC, four neighboring luma and chroma reconstructed samples are selected to derive model parameters. Assuming the averages of the neighboring luma and chroma samples are lumaAvg and chromaAvg respectively, then β is derived as β = chromaAvg - (α′ + dA) · lumaAvg. The average of the neighboring luma samples (i.e., lumaAvg) can be calculated by the average of all selected luma samples, the luma DC mode value of the current luma CB, or the average of the maximum and minimum luma samples (e.g., or ). Similarly, the average of the neighboring chroma samples (i.e., chromaAvg) can be calculated by the average of all selected chroma samples, the chroma DC mode value of the current chroma CB, or the average of the maximum and minimum chroma samples (e.g., or ). Note that for non-4:4:4 color subsampling formats, the selected neighboring luma reconstructed samples can be from the output of the CCLM downsampling process.
[0220] For another implementation, the shift parameter s can be a constant value (e.g., s can be 3, 4, 5, 6, 7, or 8), and dS equals 0 and does not need to be transmitted.
[0221] For another implementation, in MMLM, the set of guiding parameters can also be transmitted for each model. For example, one set of guiding parameters is transmitted for one model, and another set of guiding parameters is transmitted for another model. Or, one set of guiding parameters is transmitted and shared among the linear models. Or only one set of guiding parameters is transmitted for a selected model, and the other model is not further refined by the set of guiding parameters.
[0222] In another implementation, the MSB part of α' is selected according to the cost of possible final scaling parameters. That is, a possible final scaling parameter is derived from the transmitted dA and a possible MSB value of α'. For each possible final scaling parameter, the cost defined by the sum of the absolute differences between the neighboring reconstructed chroma samples generated by the CCLM model and the corresponding chroma values is calculated, and the final scaling parameter is the one with the minimum cost. In one implementation, the cost function is defined as the sum of squared errors.
[0223] Inherit neighbouring model parameters for refining the cross-component model parameters
[0224] The final scaling parameter of the current block is inherited from the neighbouring block and further refined by dA (e.g., dA derivation or transmission can be similar or the same as the method in the previous "set of guiding parameters for refining the cross-component model parameters"). Once the final scaling parameter is determined, the offset parameter (e.g., β in CCLM) is derived based on the inherited scaling parameter and the average of the neighbouring luminance and chrominance samples of the current block. For example, if the final scaling parameter is inherited from a selected neighbouring block and the inherited scaling parameter is α′ nei , then the final scaling parameter is (α′ nei + dA). For another implementation, the final scaling parameter is inherited from the history list and further refined by dA. For example, the history list records the last j final scaling parameter entries of the previous CCLM-encoded blocks. Then, the final scaling parameter is inherited from a selected entry in the history list, α′ list , and the final scaling parameter is (α′ list + dA). For another implementation, the final scaling parameter is inherited from the history list or the neighbouring block, but only the MSB (most significant bit) part of the inherited scaling parameter is taken, and the LSB (least significant bit) of the final scaling parameter comes from dA. For another implementation, the final scaling parameter is inherited from the history list or the neighbouring block, but not further refined by dA.
[0225] For another implementation, after inheriting the model parameter, the offset can be further refined by dB. For example, if the final offset parameter is inherited from a selected neighbouring block and the inherited offset parameter is β′ nei , then the final scaling parameter is (β′ nei + dB). For another implementation, the final offset parameter is inherited from the history list and further refined by dB. For example, the history list records the last j final scaling parameter entries of the previous CCLM-encoded blocks. Then, the final scaling parameter is inherited from a selected entry in the history list, β′ list , and the final scaling parameter is (β′ list + dB).
[0226] For another implementation, if the inherited neighbouring block uses CCCM encoding, the filter coefficient (c i) is inherited. The offset parameter (e.g., c6×B or c6 in CCCM) can be re-derived based on the inherited parameter and the average of the neighboring corresponding position luminance and chrominance samples of the current block. For another embodiment, only some of the filter coefficients are inherited (e.g., only n out of 6 filter coefficients are inherited, where 1 ≤ n < 6), and the remaining filter coefficients are further re-derived using the neighboring luminance and chrominance samples of the current block.
[0227] For another embodiment, if the inherited candidate applies the GLM gradient mode to its luminance reconstruction samples, the current block should also inherit the candidate's GLM gradient mode and apply it to the current luminance reconstruction samples.
[0228] For another example, if the inherited neighboring block is encoded using multiple cross-component models (e.g., MMLM or CCCM using multiple models), the classification threshold is also inherited to classify the neighboring samples of the current block into multiple groups, and the inherited multiple cross-component model parameters are further assigned to each group. For another example, the classification threshold is the average of the neighboring reconstructed luminance samples, and the inherited multiple cross-component model parameters are further assigned to each group. Similarly, once the final scaling parameters for each group are determined, the offset parameter for each group is re-derived based on the inherited scaling parameter and the average of the neighboring luminance and chrominance samples of each group of the current block. For another example, if CCCM using multiple models is used, once the final coefficient parameters for each group (e.g., c0 to c5, except c6 in CCCM) are determined, the offset parameter for each group (e.g., c6×B or c6 in CCCM) is re-derived based on the inherited coefficient parameter and the neighboring luminance and chrominance samples of each group of the current block.
[0229] For another embodiment, the inherited model parameters may depend on the color component. For example, the Cb and Cr components can inherit model parameters or model derivation methods from the same candidate or different candidates. For another example, only one color component inherits the model parameters, and the other color component derives the model parameters based on the inherited model derivation method (e.g., if the inherited candidate is encoded by MMLM or CCCM, the current block also derives the model parameters based on MMLM or CCCM using the current neighboring reconstructed samples). For another example, only one color component inherits the model parameters, and the other color component derives its model parameters using the current neighboring reconstructed samples.
[0230] For another embodiment, after decoding a block, the cross-component model of the current block is derived and stored for later reconstruction of neighboring blocks using the inherited neighbor model parameters. For example, even if the current block is coded by inter prediction, the cross-component model parameters of the current block can be derived by using the reconstructed or predicted samples of the current luma and chroma. Later, if another block is predicted by using the inherited neighbor model parameters, it can inherit the model parameters from the current block. For another example, if the current block is coded by cross-component prediction, the cross-component model parameters of the current block are re-derived by using the reconstructed or predicted samples of the current luma and chroma. For another example, the stored cross-component model can be CCCM, LM_LA (i.e., a single model LM derived using the above and left neighboring samples), or MMLM_LT (i.e., a multi-model LM derived using the above and left neighboring samples).
[0231] Inherit spatial neighbouring model parameters
[0232] For another embodiment, the inherited model parameters can come from an immediately neighboring block. The models of blocks from predefined positions are added to the candidate list in a predefined order. For example, the predefined positions can be Figure 14 the positions shown in, and the predefined order can be B 0, A 0, B 1, A1 and B2, or A 0, B 0, B 1, A1 and B2.
[0233] For another embodiment, the predefined positions include the immediately above position (W>>1) or ((W>>1)-1) if W is greater than or equal to TH, and the immediately left position (H>>1) or ((H>>1)-1) if H is greater than or equal to TH, where W and H are the width and height of the current block, and TH is a threshold that can be 4, 8, 16, 32, or 64.
[0234] For another embodiment, the maximum number of inherited models from spatial neighbors is less than the number of predefined positions. For example, if the predefined positions are as Figure 14 shown and there are 5 predefined positions. If the predefined order is B 0, A 0, B 1, A1 and B2, and the maximum number of inherited models from spatial neighbors is 4, the model of B2 is added to the candidate list only when the previous block is unavailable or not coded in the cross-component model.
[0235] Inheriting temporal neighbouring model parameters
[0236] For another embodiment, if the current slice / image is a non-intra slice / image, the inherited model parameters can come from the blocks in the previous encoded slice / image. For example, as Figure 19 shown, the current block position is (x, y) and the block size is w×h. The inherited model parameters can come from the blocks at positions (x’, y’), (x’, y’+h / 2), (x’+w / 2, y’), (x’+w / 2, y’+h / 2), (x’+w, y’), (x’+w, y’+h) or (x’+w, y’+h) in the previous encoded slice / image, where x’ = x + Δx and y’ = y + Δy. In one embodiment, if the prediction mode of the current block is intra, Δx and Δy are set to 0. If the prediction mode of the current block is inter prediction, Δx and Δy are set to the horizontal and vertical motion vectors of the current block. In another embodiment, if the current block is inter bi-prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference picture list 0. In another embodiment, if the current block is inter bi-prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference picture list 1.
[0237] For another embodiment, if the current block is bi-predicted, the inherited model parameters can come from the blocks in the previously encoded slice / image in the reference list. For example, if the horizontal and vertical motion vectors in reference picture list 0 are Δx L0 and Δy L0 , the motion vectors can be scaled to other reference pictures in reference lists 0 and 1. If the motion vectors are scaled to the i th th reference picture in reference list 0, the reference picture is (Δx L0,i0 , Δy L0,i0 ). The model can come from the blocks in the i th th reference picture in reference list 0, and Δx and Δy are set to (Δx L0,i0 , Δy L0,i0 ). Another example, if the horizontal and vertical motion vectors in reference picture list 0 are Δx L0 and Δy L0 , the motion vectors are scaled to the i th th reference picture in reference list 1, and the reference picture is (Δx L0,i1 , Δy L0,i1 ). The model can come from the blocks in the i th th reference picture in reference list 1, and Δx and Δy are set to (Δx L0,i1 , Δy L0,i1 ).
[0238] Inherit non - adjacent spatial neighbouring models
[0239] For another embodiment, the inherited model parameters can come from spatially neighbouring blocks. The models of blocks from predefined positions are added to the candidate list in a predefined order. For example, the pattern of positions and order can be as Figure 18 shown, where the distance between each position is the width and height of the current coding / decoding block. For another embodiment, the distance between positions closer to the current coding block is less than the distance between positions farther from the current block.
[0240] For another embodiment, the maximum number of inherited models from non - adjacent spatial neighbours is less than the number of predefined positions. For example, if the predefined positions are as Figure 20 shown, where two patterns (2010 and 2020) are shown. If the maximum number of inherited models from non - adjacent spatial neighbours is N, then search pattern 2 is used only if the number of available models from search pattern 1 is less than N.
[0241] Inheriting model parameters from history table
[0242] In one embodiment, the inherited model parameters can come from a cross - component model history table. The cross - component models in the history table can be added to the candidate list in a predefined order. In one embodiment, the addition order of historical candidates can be from the beginning to the end of the table. In another embodiment, the addition order of historical candidates can be from a predefined position to the end of the table. In another embodiment, the addition order of historical candidates can be from the end to the beginning of the table. In another embodiment, the addition order of historical candidates can be from a predefined position to the beginning of the table. In another embodiment, the addition order of historical candidates can be in an interleaved manner (e.g., the first candidate added comes from the beginning of the table, the second candidate added comes from the end of the table, and so on).
[0243] In one embodiment, a single cross - component model history table can be maintained to store previous cross - component models, and the cross - component model history table can be reset at the start of the current picture, current slice, current tile, every M CTU rows, or every N CTUs, where N and M can be any value greater than 0. In another embodiment, the cross - component model history table can be reset at the end of the current picture, current slice, current tile, current CTU row, or current CTU.
[0244] In another embodiment, multiple cross-component model history tables can be maintained to store previous cross-component models. An image can be divided into multiple regions, and a history table is reserved for each region. The size of the region is predefined and can be X by Y CTUs, where X and Y can be any value greater than 0. If there are N regions in an image, a total of N history tables are used, denoted here as History Table 1 to History Table N. There can be another history table for storing all previous cross-component models, denoted here as History Table 0. In one embodiment, History Table 0 will be updated throughout the encoding / decoding process. When reaching the end of a divided region, the history table of that divided region will be updated by History Table 0. Figure 21 Shows an example when the size of the region is 4 by 1 CTU.
[0245] In another embodiment, an image can be divided into several regions, and a history table is reserved for each region. History Table 0 and an additional history table will be updated during the encoding / decoding process. The additional history table can be determined by the current position. For example, if the current CU is in the second region, the additional history table to be updated is History Table 2.
[0246] In another embodiment, multiple history tables are used for different update frequencies. For example, the first history table is updated for each CU, the second history table is updated for every two CUs, the third history table is updated for every four CUs, and so on.
[0247] In another embodiment, multiple history tables are used to store different types of cross-component models. For example, the first history table is used to store a single model, and the second history table is used to store a multi-model. Another example, the first history table is used to store a gradient model, and the second history table is used to store a non-gradient model. Another example, the first history table is used to store a simple linear model (e.g., y = ax + b), and the second history table is used to store a complex model (e.g., CCCM).
[0248] In another embodiment, multiple history tables are used for different reconstructed luminance intensities. For example, if the average value of the reconstructed luminance samples in the current block is greater than a predefined threshold, the cross-component model will be stored in the first history table; otherwise, the cross-component model will be stored in the second history table. In another embodiment, multiple history tables are used for different reconstructed chrominance intensities. For example, if the average value of the neighboring reconstructed chrominance samples in the current block is greater than a predefined threshold, the cross-component model will be stored in the first history table; otherwise, the cross-component model will be stored in the second history table.
[0249] In one embodiment, when adding historical candidates from multiple historical tables to the candidate list, the addition order can be from the beginning to the end of a certain table, and then the next historical table is added in the same order or the reverse order. In another embodiment, the addition order can be from the end to the beginning of a certain table, and then the next historical table is added in the same order or the reverse order. In another embodiment, the addition order can be from a predefined position in a certain table to the end of the table, and then the next historical table is added in the same order or the reverse order. In another embodiment, the addition order can be from a predefined position in a certain table to the beginning of the table, and then the next historical table is added in the same order or the reverse order. In another embodiment, the addition order of historical candidates can be in a staggered manner within a certain historical table (e.g., the first candidate added is from the beginning of a certain historical table, the second candidate added is from the end of a certain historical table, and so on), and then the next historical table is added in the same order or the reverse order.
[0250] In another embodiment, the addition order can be from the beginning to the end of each historical table. In another embodiment, the addition order can be from the end to the beginning of each historical table. In another embodiment, the addition order can be from a predefined position in each historical table to the end of the table. In another embodiment, the addition order can be from a predefined position in each historical table to the beginning of the table. In another embodiment, the addition order of historical candidates can be in a staggered manner within each historical table (e.g., the first candidate added is from the beginning of all historical tables, the second candidate added is from the end of all historical tables, and so on).
[0251] In one embodiment, multiple cross-component model historical tables are used, but not all historical tables will be used to create the candidate list. Only the historical tables whose regions are close to the region where the current block is located can be used to create the candidate list.
[0252] In one embodiment, if historical candidates are used, the range for selecting non - adjacent candidates can be narrowed by using the smaller distance between each non - adjacent candidate position and its neighbors. In another embodiment, if historical candidates are used, the number of non - adjacent candidates can be reduced by measuring the distance from the upper - left position of the current block to the candidate position, and then excluding candidates with a distance greater than a predefined threshold. In another embodiment, if historical candidates are used, the number of non - adjacent candidates can be reduced by skipping candidates that are not located in the same region. In another embodiment, if historical candidates are used, the number of non - adjacent candidates can be reduced by skipping candidates that are not located in neighboring regions. The range of neighboring regions is predefined and can be M times N regions, where M and N can be any value greater than 0. In another embodiment, if historical candidates are used, the range for selecting non - adjacent candidates can be narrowed by skipping the second search pattern.
[0253] Models generated based on other inherited models
[0254] In another embodiment, a single cross - component model can be generated from multiple cross - component models. For example, if candidates are encoded using multiple cross - component models (e.g., MMLM, or CCCM using multiple models), a single cross - component model can be generated by selecting the first or second cross - component model in the multi - cross - component model.
[0255] Candidate list construction
[0256] In one embodiment, the candidate list is constructed by adding candidates in a predefined order until the maximum number of candidates is reached. The added candidates can include all or part of the above - mentioned candidates, but are not limited to the above - mentioned candidates. For example, the candidate list can include spatially - adjacent candidates, temporally - adjacent candidates, historical candidates, non - adjacent - adjacent candidates, single - model candidates generated based on other inherited models, or combined models (such as the inheritance of multiple cross - component models mentioned in the later section). For example, the candidate list can include the same candidates as the previous example, but the candidates are added to the list in a different order.
[0257] In another embodiment, if all predefined adjacent and historical candidates have been added but the maximum number of candidates has not been reached, some default candidates are added to the candidate list until the maximum number of candidates is reached.
[0258] In a sub - embodiment, the default candidates include, but are not limited to, the candidates described below. The final scaling parameter α comes from the set {0, 1 / 8, - 1 / 8, + 2 / 8, - 2 / 8, + 3 / 8, - 3 / 8, + 4 / 8, - 4 / 8}, and the offset parameter β = 1 / (1 << bit_depth) or is derived based on neighboring luma and chroma samples. For example, if the average of neighboring luma and chroma samples is lumaAvg and chromaAvg, then β is derived by β = chromaAvg - α·lumaAvg. The average of neighboring luma samples (lumaAvg) can be calculated from all selected luma samples, the luma DC mode value of the current luma CB, or the average of the maximum and minimum luma samples (e.g., lumaAvg = or ). Similarly, the average of neighboring chroma samples (chromaAvg) can be calculated from all selected chroma samples, the chroma DC mode value of the current chroma CB, or the average of the maximum and minimum chroma samples (e.g., or ).
[0259] In another sub - embodiment, the default candidate includes but is not limited to the candidates described below. The default candidate is α·G + β, where G is the luma sample gradient rather than the down - sampled luma sample L. Sixteen GLM filters described in the section titled "Gradient Linear Model (GLM)" are applied. The final scaling parameter α is from the set {0, 1 / 8, - 1 / 8, + 2 / 8, - 2 / 8, + 3 / 8, - 3 / 8, + 4 / 8, - 4 / 8}. The offset parameter β = 1 / (1 << bit_depth) or is derived based on neighboring luma and chroma samples.
[0260] In another embodiment, the default candidate can be an early candidate with refined incremental scaling parameters. For example, if the scaling parameter of the early candidate is α, the scaling parameter of the default candidate is (α + Δα), where Δα can be 1 / 8, - 1 / 8, + 2 / 8, - 2 / 8, + 3 / 8, - 3 / 8, + 4 / 8, - 4 / 8. The offset parameter of the default candidate will be derived from (α + Δα) and the average of neighboring luma and chroma samples of the current block.
[0261] Removing or modifying similar neighbouring model parameters
[0262] When inheriting cross-component model parameters from other blocks, the similarity between the inherited model and existing models in the candidate list or those model candidates derived from the neighboring reconstruction samples of the current block can be further checked (e.g., CCLM, MMLM, or CCCM models derived from the neighboring reconstruction samples of the current block). If the model of the candidate parameters is similar to the existing model, the model will not be included in the candidate list. In one embodiment, the similarity of (α×lumaAvg + β) or α between existing candidates can be compared to decide whether to include the candidate model. For example, if (α×lumaAvg + β) or α of the candidate is the same as one of the existing candidates, the candidate model will not be included. Another example is that if the difference between (α×lumaAvg + β) or α of the candidate and the existing candidate is less than a threshold, the candidate model will not be included. Additionally, the threshold can be adaptively adjusted based on encoding and decoding information (e.g., the size or region of the current block). Another example is that when comparing similarities, if both the candidate and the existing model use CCCM, the value of (c0C + c1N + c2S + c3E + c4W + c5P + c6B) can be checked to decide whether to include the candidate model. In another embodiment, if the candidate position points to the same CU as the existing candidate, the model of the candidate parameters will not be included. In another embodiment, if the model of the candidate is similar to the existing candidate model, the inherited model parameters can be adjusted to make the inherited model different from the existing candidate model. For example, if the inherited scaling parameter is similar to the existing candidate model, a predefined offset (e.g., 1 >> S or -(1 >> S), where S is the shift parameter) can be added to make the inherited parameter different from the existing candidate model.
[0263] Reordering the candidates in the list
[0264] The candidates in the list can be reordered to reduce the syntax overhead when signaling the selected candidate index. The reordering rule can depend on the encoding and decoding information of neighboring blocks or the model error. For example, if the neighboring upper or left block is encoded by MMLM, the MMLM candidates in the list can be moved to the beginning of the current list. Similarly, if the neighboring upper or left block is encoded by a single model LM or CCCM, the single model LM or CCCM candidates in the list can be moved to the beginning of the current list. Similarly, if the neighboring upper or left block uses GLM, the GLM-related candidates in the list can be moved to the beginning of the current list.
[0265] In another embodiment, the reordering rule is based on comparing the model error by applying the candidate model to the neighboring template of the current block and then comparing it with the reconstruction samples of the neighboring template. For example, as Figure 22 shown, the size of the upper neighboring template 2220 of the current block is wa ×h a , the size of the left adjacent template 2230 of the current block 2210 is w b ×h b . Assume there are K models in the current candidate list, and α k and β k are the final scaling and offset parameters after inheriting candidate k. The model error of candidate k corresponding to the upper adjacent template is:
[0266]
[0267] where, and are the luminance reconstruction samples (e.g., after the downsampling process or after applying the GLM mode) and chrominance reconstruction samples at position (i, j) in the upper template, 0 ≤ i < w a and 0 ≤ j < h a .
[0268] Similarly, the model error of candidate k through the left adjacent template is:
[0269]
[0270] where and are the reconstructed luminance samples (e.g., after applying the downsampling process or GLM mode) and reconstructed chrominance samples at the left template position (m, n), 0 ≤ m < w b and 0 ≤ n < h b . Then the model error of candidate k is: After calculating the model errors of all candidates, a model error list E = {e 0 , e 1 , e 2 , …, e k , …, e K} can be obtained. Then, the candidate indices in the inherited candidate list can be rearranged by sorting the model error list in ascending order. In another embodiment, if candidate k uses CCCM prediction, and are defined as:
[0271]
[0272] where c0 k , c1 k , c2 k , c3 k , c4 k , c5 k and c6 kis the final filtering coefficient after inheriting candidate k. P and B are the non - linear term and the bias term. In another embodiment, if the above - mentioned neighboring template is not available, then Similarly, if the left neighboring template is not available, then If neither template is available, the candidate index re - ordering method using the model error is not applied. In another embodiment, not all positions within the above - mentioned and left neighboring templates are used to calculate the model error. Some positions within the above - mentioned and left neighboring templates can be selected to calculate the model error. For example, a first starting position and a first subsampling interval can be defined depending on the width of the current block to partially select positions within the above - mentioned neighboring template. Similarly, a second starting position and a second subsampling interval can be defined depending on the height of the current block to partially select positions within the left neighboring template. Another example, h a or h b can be a constant value (e.g., h a or h b can be 1, 2, 3, 4, 5, or 6). Another example, h a or h b can depend on the block size. If the current block size is greater than or equal to a threshold, h a or h b equals the first value. Otherwise, h a or h b equals the second value. In another embodiment, after re - ordering candidates based on the template cost, the redundancy of the candidates can be further checked. If the difference in the template cost between a candidate and the previous candidate in its list is less than a threshold, the candidate is considered redundant. If a candidate is considered redundant, it can be removed from the list or it can be moved to the end of the list. Inherit from candidates in the candidate list of neighbors.
[0273] The candidates in the current inheritance candidate list can come from neighboring blocks. For example, it can inherit the top k candidates in the candidate list of the neighboring block's inheritance. As Figure 23 shown, the current block can inherit the candidate list of the upper neighboring block (Candidate list of Block A) the first two candidates in ) and the first two candidates in the candidate list of the left neighboring block (Candidate list of BlockL). For one embodiment, after adding neighboring spatial candidates and non - neighboring spatial candidates, if the current inherited candidate list is not full, the candidates in the neighboring block candidate list are included in the current inherited candidate list. For another embodiment, when including the candidates in the neighboring block candidate list, the candidates in the left neighboring block candidate list are included before the candidates in the upper neighboring block candidate list. For yet another embodiment, when including the candidates in the neighboring block candidate list, the candidates in the upper neighboring block candidate list are included before the candidates in the left neighboring block candidate list.
[0274] Signaling the inheriting candidate index in the list (Inheriting candidates from the candidates inthe candidate list of neighbours)
[0275] An on / off flag can be signaled to indicate whether the current block inherits the cross - component model parameters of the neighboring block. This flag can be signaled per CU / CB, per PU, per TU / TB, per color component, or per chrominance color component. High - level syntax can be signaled in the SPS, PPS (Picture Parameter Set), PH (Picture Header), or SH (Slice Header) to indicate whether the proposed method is allowed for the current sequence, picture, or slice.
[0276] If the current block inherits the cross - component model parameters of the neighboring block, the inheriting candidate index is signaled. This index can be signaled (e.g., using truncated unary code, Exp - Golomb code, or fixed - length code signaling) and shared between the current Cb and Cr blocks. For another example, this index can be signaled per color component. For example, one inheriting index is signaled for the Cb component and another inheriting index is signaled for the Cr component. For another example, it can use the chrominance intra - prediction syntax (e.g., IntraPredModeC[xCb][yCb]) to store the inheriting index.
[0277] If the current block inherits the cross-component model parameters of the neighboring block, the current chrominance intra prediction mode (e.g., IntraPredModeC[xCb][yCb] defined in the VVC standard) is temporarily set to the cross-component mode (e.g., CCLM_LT) during the bitstream syntax parsing phase. Subsequently, during the prediction phase or the reconstruction phase, the candidate list is derived, and then the inherited candidate model is determined by inheriting the candidate index. After obtaining the inherited model, the coding and decoding information of the current block is updated according to the inherited candidate model. The coding and decoding information of the current block includes but is not limited to the prediction mode (e.g., CCLM_LT or MMLM_LT), the relevant sub-mode flag (e.g., CCCM mode flag), the prediction mode (e.g., GLM mode index), and the current model parameters. Then, the prediction of the current block is generated according to the updated coding and decoding information.
[0278] Inheriting multiple cross-component models (Signalling the inherit candidate indexin the list)
[0279] The final prediction of the current block can be a combination of multiple cross-component models, or a fusion of the selected cross-component model and the prediction by non-cross-component coding and decoding tools (e.g., intra angular prediction mode, intra plane / DC mode, or inter prediction mode). In one embodiment, if the current candidate list size is N, k candidates (where k ≤ N) can be selected from a total of N candidates. Then, by applying the cross-component models of the selected k candidates using the corresponding luminance reconstruction samples, k predictions are generated respectively. The final prediction of the current block is the combined result of these k predictions. For example, if two candidate predictions (denoted as p cand1 and p cand2 ) are combined, the final prediction of the current block at the (x,y) position is p final (x,y) = (1 - α) × p cand1 (x,y) + α × p cand2 (x,y), where α is a weighting factor. In addition, the weighting factor α can be predefined or implicitly derived by the neighboring template cost. For example, by using the template cost defined in the section titled "Inheriting non-adjacent spatial neighboring models", the corresponding template costs of the two candidates are e cand1 and e cand2 , then α is e cand1 / (e cand1 +e cand2 ). In another embodiment, if two candidate models are combined, the selected model comes from the first two candidates in the list. In yet another embodiment, if i candidate models are combined, the selected models come from the first i candidates in the list.
[0280] In another embodiment, if the current candidate list size is N, k candidates can be selected from a total of N candidates (where k ≤ N). These k cross-component models can be merged into a final cross-component model by weighted averaging of the corresponding model parameters. For example, if a cross-component model has M parameters, the j-th parameter of the final cross-component model is the weighted average of the j-th parameters of the selected k candidates, where j is 1...M. Then, the final prediction is generated by applying the final cross-component model to the corresponding luminance reconstruction samples. For example, if two candidate models are and the final cross-component model is where α is a weighting factor that can be predefined or implicitly derived from the neighboring template cost, and is the x-th model parameter of the y-th candidate. For example, by using the template cost defined in the section titled "Inheriting Non-Adjacent Spatial Neighboring Models", the corresponding template costs of the two candidates are e cand1 and e cand2 , then α is e cand1 / (e cand1 +e cand2 ). As another example, one of the two candidate models comes from a spatially adjacent neighboring candidate, and the other comes from a non-adjacent spatial candidate or a historical candidate. If the spatially adjacent neighboring candidate is not available, both candidate models come from non-adjacent spatial candidates or historical candidates. In another embodiment, if two candidate models are merged, the selected models come from the first two candidates in the list. In another embodiment, if i candidate models are merged, the selected models come from the first i candidates in the list.
[0281] In another embodiment, two cross-component models are merged into a final model by weighted averaging of the corresponding model parameters, where one of the two cross-component models comes from the aforementioned spatial neighboring candidate and the other comes from the left spatial neighboring candidate. The aforementioned spatial neighboring candidate is a neighboring candidate whose vertical position is less than or equal to the top block boundary position of the current block. The left spatial neighboring candidate is a neighboring candidate whose horizontal position is less than or equal to the left block boundary position of the current block. The weighting factor α is determined according to the horizontal and vertical spatial positions within the current block. For example, if two candidate predictions (denoted as p above and p left ) are merged, the final prediction at the (x, y) position of the current block is p final (x,y) = (1 - α) × p above (x,y) + α × p left(x, y), where α = u / (x + u). In another embodiment, the above spatial neighboring candidate is the first candidate in the list whose vertical position is less than or equal to the top block boundary position of the current block. The left spatial neighboring candidate is the first candidate in the list whose horizontal position is less than or equal to the left block boundary position of the current block.
[0282] In another embodiment, a cross-component model candidate can be combined with the prediction of a non-cross-component encoding / decoding tool. For example, a cross-component model candidate is selected from the list, and its prediction is denoted as p ccm . Another prediction can come from chroma DM, chroma DIMD, or the intra angle mode, denoted as p non-ccm . The final prediction at the (x, y) position of the current block is p final (x, y) = (1 - α) × p ccm (x, y) + α × p non-ccm (x, y), where α is a weighting factor that can be predefined or implicitly derived from the neighboring template cost. In the same example, the prediction of the non-cross-component encoding / decoding tool can be predefined or signaled. The prediction of the non-cross-component encoding / decoding tool is chroma DM or chroma DIMD. In another example, the prediction of the non-cross-component encoding / decoding tool is signaled, but the index of the cross-component model candidate is predefined or determined by the encoding / decoding mode of neighboring blocks. In the same example, if at least one neighboring spatial block is encoded using the CCCM mode, the first candidate with CCCM model parameters is selected. If at least one neighboring spatial block is encoded using the GLM mode, the first candidate with GLM mode parameters is selected. Similarly, if at least one neighboring spatial block is encoded using the MMLM mode, the first candidate with MMLM parameters is selected.
[0283] In another implementation, a cross-component model candidate can be combined with the prediction of the current cross-component model. For example, a cross-component model candidate is selected from the list, and its prediction is denoted as p ccm . Another prediction can come from the cross-component prediction mode of the current neighboring reconstructed samples and is denoted as p curr-ccm . The final prediction at the (x, y) position of the current block is p final (x, y) = (1 - α) × p ccm (x, y) + α × p curr-ccm(x, y), where α is a weight factor that can be predefined or implicitly derived based on neighboring template costs. For still the same example, the prediction of the current cross-component model can be predefined or signaled. The predictions of non-cross-component encoding / decoding tools are CCCM_LT, LM_LT (i.e., a single-model LM uses top and left neighboring samples to derive the model), or MMLM_LT (i.e., a multi-model LM uses top and left neighboring samples to derive the model). In one implementation, the selected cross-component model candidate is the first candidate in the list.
[0284] In another implementation, multiple cross-component models can be combined into a final cross-component model. For example, one model can be selected from one candidate and a second model from another candidate as a multi-model mode. The selected candidates can be CCLM / MMLM / GLM / CCCM encoding candidates. The multi-model classification threshold can be the average of the offset parameters of the two selection modes (e.g., the offset / β in CCLM, or c6×B or c6 in CCCM). In one implementation, if two candidate models are combined, the selected models are the first two candidates in the list. In another implementation, the classification threshold is set to the average of the neighboring luma and chroma samples of the current block.
[0285] Refining the inherited candidate positions
[0286] In one implementation, the final inherited model of the current block comes from the cross-component model indicating the candidate position and has an incremental position. For example, if the currently selected candidate position is then an incremental position can be further signaled, to indicate the position of the final inherited model. That is, the final inherited model of the current block comes from the cross-component model at In one implementation, the signaled incremental position can only have a horizontal incremental position or a vertical incremental position, i.e., or Furthermore, the signaled incremental position can be shared among multiple color components or signaled for each color component. For example, the signaled incremental position is shared by the current Cb and Cr blocks, or the signaled incremental position is only used for the current Cb block or the current Cr block. Additionally, the signaled or may have a sign bit to indicate a positive incremental position or a negative incremental position. When indicating the magnitude of or , it can be signaled through a look-up table index. For example, the look-up table is {1, 2, 4, 8, 16, …}, if If it is equal to 8, then the signaling table index is 3 (the first table index is 0).
[0287] In one embodiment, when selecting a candidate from a candidate list, further search for models at neighboring positions of the selected candidate. The final inherited model can come from neighboring positions of the selected candidate. Search for positions of a predefined search pattern within the area surrounding the selected candidate. In one embodiment, the neighboring positions searched are different from the selected candidate in the horizontal or vertical direction, i.e., the incremental position is or In another embodiment, the neighboring positions searched are different from the selected candidate diagonally, i.e., the incremental position is where Note that the incremental position can be positive or negative.
[0288] In another embodiment, further search for models at neighboring positions of the candidate only when the selected candidate is a non-adjacent candidate. Search for positions of a predefined search pattern within the area surrounding the selected candidate. For example, assume that the distance between non-adjacent candidates is the width and height of the current coding / decoding block. After selecting a non-adjacent candidate, further search for positions where both the horizontal distance and the vertical distance are less than the width and height of the current coding / decoding block, i.e., within the range of ±width, within the range of ±height. In one embodiment, the neighboring positions searched are different from the selected candidate in the horizontal or vertical direction, i.e., the position difference is or In another embodiment, the neighboring positions searched are different from the selected candidate diagonally, i.e., the position difference is where
[0289] Inheriting from shared cross-component models
[0290] In one embodiment, the current picture is divided into multiple non-overlapping regions, each region having a size of M×N. Derive a shared cross-component model for each region respectively. The neighboring available luminance / chrominance reconstruction samples of the current region are used to derive the shared cross-component model of the current region. Then, for blocks within the current region, it can be determined whether to inherit the shared cross-component model or derive a cross-component model through the neighboring available luminance / chrominance reconstruction samples of the block. In one embodiment, M×N can be a predefined value (e.g., 32x32 for chrominance format), a signaling value (e.g., signaled at the sequence / picture / slice / tile level), a derived value (e.g., depending on the CTU size), or the maximum allowed transform block size.
[0291] In another embodiment, each region may have multiple shared cross-component models. For example, multiple shared cross-component models may be derived using various neighborhood templates (e.g., top and left neighborhood samples, only top neighborhood samples, only left neighborhood samples). Additionally, the shared cross-component model for the current region may be inherited from a previously used cross-component model. For example, the shared model may be inherited from models in adjacent spatial neighbors, non-adjacent spatial neighbors, temporal neighbors, or the historical list.
[0292] When signaling, a first flag may be used to determine whether the current cross-component model is inherited from a shared cross-component model. If the current cross-component model is inherited from a shared cross-component model, a second syntax indicates the inheritance index of the shared cross-component model (e.g., signaled using a truncated unary code, Exp-Golomb code, or fixed-length code).
[0293] As described above, cross-component prediction with inherited model parameters may be implemented at the encoder side or the decoder side. For example, any proposed cross-component prediction method may be implemented in the intra / inter-frame codec module in the decoder (e.g., Intra Pred.150 / MC 152 in Figure 1B or the intra / inter-frame codec module in the encoder (e.g., IntraPred.110 / Inter Pred.112 in Figure 1A ). Any proposed cross-component prediction method with inherited model parameters may also be implemented as a circuit, coupled to the intra / inter-frame codec module of the decoder or encoder. However, the decoder or encoder may also use additional processing units to perform the required cross-component prediction processing. Although the intra-prediction units (e.g., units 110 / 112 in Figure 1A and units 150 / 152 in Figure 1B are shown as separate processing units, they may correspond to executable software or firmware code stored on a medium, such as a hard disk or flash memory, for a CPU (Central Processing Unit) or a programmable device (e.g., a DSP (Digital Signal Processor) or an FPGA (Field Programmable Gate Array)).
[0294] Figure 24Shows a flowchart of an exemplary video coding and decoding system that incorporates a shared cross-component model with inheritance of a history table using a predefined insertion order, according to one embodiment of the present invention. The steps shown in the flowchart can be implemented as program code executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart can also be implemented based on hardware, such as one or more electronic devices or processors arranged to execute the steps in the flowchart. According to the method, input data related to a current block is received in step 2410, including a first color block and a second color block, where the input data includes pixel data to be encoded at the encoder side or encoded data related to the current block to be decoded at the decoder side. A prediction candidate list is determined in step 2420, including one or more cross-component prediction candidates inherited from a cross-component model history table. In step 2430, a target model parameter set related to a target inheritance prediction model is determined based on an inheritance model parameter set related to the target inheritance prediction model selected from the prediction candidate list. In step 2440, the second color block is encoded or decoded using prediction data, where the prediction data includes a cross-color prediction generated by applying the target inheritance prediction model and the target model parameter set to a reconstructed first color block.
[0295] Figure 25 Shows a flowchart of an exemplary video coding and decoding system that incorporates a shared cross-component model with inheritance of a history table using a specific reset point, according to one embodiment of the present invention. According to the method, input data related to a current block is received in step 2510, including a first color block and a second color block, where the input data includes pixel data to be encoded at the encoder side or encoded data related to the current block to be decoded at the decoder side. A prediction candidate list is determined in step 2520, including one or more cross-component prediction candidates inherited from a cross-component model history table, where the cross-component model history table is reset at a specific point related to an image region including non-CTUs. In step 2530, a target model parameter set related to a target inheritance prediction model is determined based on an inheritance model parameter set related to the target inheritance prediction model selected from the prediction candidate list. In step 2540, the second color block is encoded or decoded using prediction data, where the prediction data includes a cross-color prediction generated by applying the target inheritance prediction model and the target model parameter set to a reconstructed first color block.
[0296] Figure 26A flowchart of an exemplary video coding and decoding system is shown, which combines the use of multiple historical tables to inherit a shared cross-component model according to an embodiment of the present invention. According to the method, input data related to a current block is received in step 2610, including a first color block and a second color block, where the input data includes pixel data to be encoded at the encoder side or encoded data related to the current block to be decoded at the decoder side. In step 2620, a prediction candidate list is determined, including one or more cross-component prediction candidates inherited from multiple cross-component model historical tables. In step 2630, a target model parameter set related to a target inherited prediction model is determined based on an inherited model parameter set related to the target inherited prediction model selected from the prediction candidate list. In step 2640, the second color block is encoded or decoded using prediction data, where the prediction data includes cross-color prediction generated by applying the target inherited prediction model and the target model parameter set to the reconstructed first color block.
[0297] The flowchart shown is intended to illustrate an example of video coding and decoding according to the present invention. A person skilled in the art can modify each step, rearrange the steps, split the steps or combine the steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics are used to illustrate examples of practicing the present invention. A person skilled in the art can practice the present invention by replacing these syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0298] The foregoing description is intended to enable a person of ordinary skill in the art to practice the present invention according to a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the present invention is not intended to be limited to the specific embodiments shown and described, but should be accorded the widest scope in accordance with the principles and novel features disclosed herein. In the foregoing detailed description, various specific details are shown to provide a thorough understanding of the present invention. However, those skilled in the art will understand that the present invention can be practiced.
[0299] Embodiments of the present invention as described above can be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuits integrated into a video compression chip, or program codes integrated into video compression software to perform the processing described herein. An embodiment of the present invention can also be program codes to be executed on a digital signal processor (DSP) to perform the processing described herein. The invention may also relate to multiple functions executed by a computer processor, a digital signal processor, a microprocessor, or a field programmable gate array (FPGA). These processors can be configured to perform specific tasks by executing machine-readable software codes or firmware codes that define the methods embodied in the present invention. The software codes or firmware codes can be developed in different programming languages and different formats or styles. The software codes can also be compiled for different target platforms. However, different software code formats, styles, and languages, as well as other methods of configuring codes to perform tasks according to the present invention, do not deviate from the spirit and scope of the present invention.
[0300] The invention can be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are illustrative only and not restrictive in all respects. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes within the meaning and equivalent scope of the claims should be included within their scope.
Claims
1. A method for encoding and decoding a color image using an encoding and decoding tool, comprising one or more cross-component model-related modes, the method comprising: Receiving input data related to a current block, including a first color block and a second color block, wherein the input data includes pixel data to be encoded at an encoder side or encoded data related to the current block to be decoded at a decoder side; Determining a prediction candidate list including one or more cross-component prediction candidates inherited from a cross-component model history table; Deriving a target model parameter set related to the target inherited prediction model based on an inherited model parameter set related to the target inherited prediction model selected from the prediction candidate list; and Encoding or decoding the second color block using prediction data, the prediction data including a cross-color prediction generated by applying the target inherited prediction model and the target model parameter set to a reconstructed first color block.
2. The method according to claim 1, wherein the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list according to a predefined order.
3. The method according to claim 2, wherein the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the beginning to the end of the cross-component model history table.
4. The method according to claim 2, wherein the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the end to the beginning of the cross-component model history table.
5. The method according to claim 2, wherein the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from a predefined position in the cross-component model history table to the end or the beginning of the cross-component model history table.
6. The method according to claim 2, wherein the one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the cross-component model history table in an interleaved manner.
7. An apparatus for video encoding and decoding, the apparatus comprising one or more electronic devices or processors arranged to: Receiving input data related to a current block, including a first color block and a second color block, wherein the input data includes pixel data to be encoded at an encoder side or encoded data related to the current block to be decoded at a decoder side; Determining a prediction candidate list including one or more cross-component prediction candidates inherited from a cross-component model history table; Deriving a target model parameter set related to the target inherited prediction model based on an inherited model parameter set related to the target inherited prediction model selected from the prediction candidate list; and Encoding or decoding the second color block using prediction data, the prediction data including a cross-color prediction generated by applying the target inherited prediction model and the target model parameter set to a reconstructed first color block.
8. A method for encoding and decoding a color image using an encoding and decoding tool, comprising one or more cross-component model-related modes, the method comprising: Receive input data related to a current block, including a first color block and a second color block, where the input data includes pixel data to be encoded at an encoder side or encoded data related to the current block to be decoded at a decoder side; Determine a prediction candidate list including one or more cross-component prediction candidates inherited from a cross-component model history table, where the cross-component model history table is reset at a specific point related to an image region including non-CTUs; Derive a target model parameter set related to the target inheritance prediction model based on an inheritance model parameter set related to the target inheritance prediction model selected from the prediction candidate list; and Encode or decode the second color block using prediction data, where the prediction data includes cross-color prediction generated by applying the target inheritance prediction model and the target model parameter set to a reconstructed first color block.
9. The method according to claim 8, wherein the image region corresponds to a current picture, slice or tile.
10. The method according to claim 8, wherein the image region corresponds to every M CTU rows or every N CTUs, where M and N are positive integers.
11. The method according to claim 8, wherein the specific point related to the image region corresponds to the start or the end of the image region.
12. An apparatus for video coding and decoding, the apparatus including one or more electronic devices or processors arranged to: Receive input data related to a current block, including a first color block and a second color block, where the input data includes pixel data to be encoded at an encoder side or encoded data related to the current block to be decoded at a decoder side; Determine a prediction candidate list including one or more cross-component prediction candidates inherited from a cross-component model history table, where the cross-component model history table is reset at a specific point related to an image region including non-CTUs; Derive a target model parameter set related to the target inheritance prediction model based on an inheritance model parameter set related to the target inheritance prediction model selected from the prediction candidate list; and Encode or decode the second color block using prediction data, where the prediction data includes cross-color prediction generated by applying the target inheritance prediction model and the target model parameter set to a reconstructed first color block.
13. A method for coding and decoding a color image using coding and decoding tools, including one or more cross-component model related modes, the method comprising: Receive input data related to a current block, including a first color block and a second color block, where the input data includes pixel data to be encoded at an encoder side or encoded data related to the current block to be decoded at a decoder side; Determine a prediction candidate list including one or more cross-component prediction candidates inherited from a plurality of cross-component model history tables; Derive a target model parameter set related to the target inheritance prediction model based on an inheritance model parameter set related to the target inheritance prediction model selected from the prediction candidate list; and Encode or decode the second color block using prediction data, the prediction data including cross-color prediction generated by applying the target inheritance prediction model and the target model parameter set to a reconstructed first color block.
14. The method according to claim 13, wherein each image is divided into a plurality of regions, and a cross-component model history table is maintained for each of the plurality of regions.
15. The method according to claim 14, wherein the sizes of the plurality of regions are predefined.
16. The method according to claim 15, wherein the sizes of the plurality of regions correspond to X times Y CTUs, where X and Y are positive integers.
17. The method according to claim 13, wherein each image is divided into N regions, and the plurality of cross-component model history tables correspond to N history tables, where N is an integer greater than 1.
18. The method according to claim 13, wherein cross-component model history table 0 is used to store all previous cross-component models.
19. The method according to claim 18, wherein cross-component model history table 0 is always updated during the encoding or decoding process.
20. The method according to claim 18, wherein cross-component model history table 0 and one additional history table among the plurality of cross-component model history tables are updated during the encoding or decoding process.
21. The method according to claim 20, wherein the additional history table is determined according to the current position of the current block.
22. The method according to claim 13, wherein at least two cross-component model history tables are updated at different frequencies.
23. The method according to claim 13, wherein the plurality of cross-component model history tables are used to store different types of cross-component models.
24. The method according to claim 23, wherein the different types of cross-component models correspond to single model and multi-model, gradient model and non-gradient model, or simple linear model and complex model.
25. The method according to claim 23, wherein the different types of cross-component models correspond to different reconstructed luminance intensities or different reconstructed chrominance intensities.
26. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the beginning to the end of a cross-component model history table, and then inserted from the next cross-component model history table in the same order or in the reverse order.
27. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the end to the beginning of a cross-component model history table, and then inserted from the next cross-component model history table in the same order or in the reverse order.
28. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are from a predefined position to the end or the beginning of a cross-component model history table, and then inserted from the next cross-component model history table in the same order or in the reverse order.
29. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from one cross-component model history table in a staggered manner and then inserted from the next cross-component model history table in the same order or in the reverse order.
30. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the beginning to the end of each of the plurality of cross-component model history tables.
31. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are inserted into the prediction candidate list from the end to the beginning of each of the plurality of cross-component model history tables.
32. The method according to claim 13, wherein one or more inherited cross-component prediction candidates are inserted into the prediction candidate list at the end or at the beginning of each of the plurality of cross-component model history tables from a predefined position.
33. The method according to claim 13, wherein only a subset of the plurality of cross-component model history tables, whose corresponding regions are close to the current region containing the current block, is used to create the prediction candidate list.
34. The method according to claim 13, wherein when one or more cross-component prediction candidates inherited from the plurality of cross-component model history tables are used to create the prediction candidate list, the range of selecting non-adjacent candidates is narrowed.
35. The method according to claim 34, wherein the range of selecting non-adjacent candidates is narrowed by measuring the distance from the upper left position of the current block to the target candidate position, and then cases where the distance between the target candidate and the distance is greater than a predefined threshold are excluded.
36. The method according to claim 34, wherein the non-adjacent candidates are not located in the same region as the current block, thereby skipping insertion into the prediction candidate list.
37. An apparatus for video coding and decoding, the apparatus comprising one or more electronic devices or processors arranged to: receive input data related to a current block, including a first color block and a second color block, wherein the input data includes pixel data to be encoded at an encoder side or encoded data related to the current block to be decoded at a decoder side; determine a prediction candidate list including one or more cross-component prediction candidates inherited from a plurality of cross-component model history tables; derive a target model parameter set related to a target inherited prediction model based on an inherited model parameter set related to the target inherited prediction model selected from the prediction candidate list; encode or decode the second color block using prediction data, including cross-color prediction generated by applying the target inherited prediction model and the target model parameter set to the reconstructed first color block.