Multi-hypothesis mixing method and device for cross-component model merging mode in video coding and decoding

By employing a multi-hypothesis hybrid approach and combining various cross-component models, the prediction process for chroma blocks is optimized, solving the encoding and decoding efficiency and accuracy issues of existing cross-component models in VVC and achieving more efficient video encoding and decoding results.

CN120982098APending Publication Date: 2025-11-18MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480027001.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-20
Filing Date
2024-02-20
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing video coding and decoding technologies struggle to effectively utilize cross-component models to improve the efficiency and accuracy of chroma prediction when dealing with multi-functional video coding and decoding standards such as VVC.

Method used

A multi-hypothesis hybrid approach is adopted, which generates the final predictor by mixing the basic predictor and one or more cross-component prediction candidates and using predefined weights or signaling information. The prediction process of chroma blocks is optimized by combining various cross-component models such as CCLM, MMLM, CCCM and GLM.

Benefits of technology

It improves the prediction accuracy and encoding/decoding efficiency of chroma blocks, reduces redundancy and computational complexity in the encoding process, and enhances the overall performance of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982098A_ABST
    Figure CN120982098A_ABST
Patent Text Reader

Abstract

A multi-hypothesis hybrid video coding and decoding method and apparatus for cross component model merge mode. According to the method, input data associated with a current block is received, the current block comprises a first color block and a second color block, and the input data comprises pixel data to be encoded at an encoder end or encoded data to be decoded at a decoder end and associated with the current block. A base predictor for a second color block is determined, where the base predictor corresponds to an intra predictor or an inter predictor. One or more hypotheses of a second color block are determined, wherein each of the one or more hypotheses corresponds to one cross-component prediction candidate. A final predictor is derived by mixing the base predictor and the one or more hypotheses. And encoding or decoding the second color block using prediction data including the final predictor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing This invention is a non-provisional application, claiming priority to U.S. Provisional Patent Application No. 63 / 485,937, filed on February 20, 2023. The aforementioned U.S. Provisional Patent Application is incorporated herein by reference in its entirety. [Technical Field] This invention relates to video encoding and decoding systems. Specifically, this invention relates to multi-hypothesis blending for cross-component model merge mode. [Background Technology] Versatile Video Coding (VVC) is the latest international video coding standard developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) Joint Video Experts Team (JVET). This standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology – Codec representation of immersive media – Part 3: Versatile Video Coding, published in February 2021. VVC was developed based on its predecessor, High Efficiency Video Coding (HEVC), by adding more coding and decoding tools to improve coding and decoding efficiency and handle various types of video sources, including three-dimensional (3D) video signals.

[0004] Figure 1AAn exemplary adaptive inter-frame / intra-frame video coding system incorporating loop processing is illustrated. For intra-frame prediction 110, prediction data is derived from previously encoded video data in the current frame. For inter-frame prediction 112, motion estimation (ME) is performed at the encoder, and motion compensation (MC) is performed based on the result of ME to provide prediction data derived from other frames and motion data. Switch 114 selects either intra-frame prediction 110 or inter-frame prediction 112, and the selected prediction data is provided to adder 116 to form a prediction error, also known as a residual. The prediction error is then transformed (T) 118 and subsequently quantized (Q) 120. The transformed and quantized residual is then encoded by entropy encoder 122 to be included in the video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packaged with additional information, such as motion and encoding / decoding modes associated with intra-frame and inter-frame prediction, and other information related to loop filters applied to the underlying image regions. Additional information related to intra-frame prediction 110, inter-frame prediction 112, and loop filter 130 is provided to entropy encoder 122, such as... Figure 1A As shown. When using inter-frame prediction mode, one or more reference images must also be reconstructed at the encoder end. Therefore, the transformed and quantized residuals are processed by inverse quantization (IQ) 124 and inverse transformation (IT) 126 to recover the residuals. Then, in reconstruction (REC) 128, the residuals are added back to the prediction data 136 to reconstruct the video data. The reconstructed video data can be stored in the reference image buffer 134 and used for prediction of other frames.

[0005] like Figure 1AAs shown, the input video data undergoes a series of processing steps in the encoding system. Due to these processing steps, the reconstructed video data from REC 128 may be subject to various impairments. Therefore, a loop filter 130 is typically applied to the reconstructed video data before it is stored in the reference image buffer 134 to improve video quality. For example, a deblocking filter (DF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF) may be used. Loop filter information may need to be incorporated into the bitstream so that the decoder can correctly recover the required information. Therefore, loop filter information is also provided to the entropy encoder 122 for incorporation into the bitstream. Figure 1A In the process, the loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference image buffer 134. Figure 1A The system described herein is intended to demonstrate an exemplary architecture of a typical video encoder. It may correspond to a High Efficiency Video Codec (HEVC) system, VP8, VP9, ​​H.264, or VVC.

[0006] like Figure 1B As shown, the decoder can use the same or partially the same functional modules as the encoder, except for transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses entropy decoder 140 instead of entropy encoder 122 to decode the video bitstream into quantized transform coefficients and the required encoding / decoding information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). Intra-frame prediction 150 at the decoder end does not require mode search. Instead, the decoder only needs to generate intra-frame predictions based on the intra-frame prediction information received from entropy decoder 140. Furthermore, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from entropy decoder 140, without needing to perform motion estimation.

[0007] According to VVC, the input image is divided into non-overlapping square regions called Coding Tree Units (CTUs), similar to HEVC. Each CTU can be further divided into one or more smaller Coding Units (CUs). The resulting CU partitions can be square or rectangular in shape. Furthermore, VVC divides the CTUs into Prediction Units (PUs) as units for applying prediction processes, such as inter-frame prediction and intra-frame prediction.

[0008] The VVC standard incorporates several new encoding and decoding tools to further improve encoding and decoding efficiency compared to the HEVC standard. Some of the new tools related to this invention are described below.

[0009] Intra-mode codec with 67 intra-prediction modes To capture arbitrary edge orientations presented in natural video, the number of intra-directional modes in VVC has been expanded from 33 in HEVC to 65. New directional modes not included in HEVC are indicated by dashed arrows in Figure 2, while planar and DC modes remain unchanged. These denser intra-directional prediction modes are applicable to all block sizes and for both luma and chroma intra-prediction.

[0010] In VVC, for non-square blocks, several traditional angle intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes.

[0011] In HEVC, each intra-codec block is a square, with each side having a length that is a power of 2. Therefore, generating intra-predictors using DC mode requires no division. In VVC, blocks can be rectangular, which typically requires division for each block. To avoid division in DC prediction, for non-square blocks, only the longer side is used to calculate the average.

[0012] Wide-angle intra-frame prediction for non-square blocks Traditional angular intra-prediction directions are defined as clockwise from 45 degrees to -135 degrees. In VVC, for non-square blocks, several traditional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes use the original mode index for communication, which is remapped to the wide-angle mode index after parsing. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding / decoding method also remains unchanged.

[0013] To support these predicted directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as shown in Figures 3A and 3B.

[0014] The number of replacement modes in wide-angle directional mode depends on the aspect ratio of the block. Table 1 lists the replaced intra-prediction modes.

[0015] Table 1 – Intra-prediction modes replaced by wide-angle mode

[0016] VVC supports 4:2:2 and 4:4:4 chroma formats, as well as 4:2:0. The 4:2:2 chroma format's derived mode (DM) export table was originally ported from HEVC, with its number of entries expanded from 35 to 67 to match the expansion of intra-prediction modes. Because the HEVC specification does not support prediction angles below −135° and above 45°, the luma intra-prediction modes (ranging from 2 to 5) are mapped to 2. Therefore, the 4:2:2 chroma format's chroma DM export table is updated by replacing the values ​​of some entries in the mapping table to more accurately convert the prediction angle of the chroma blocks.

[0017] Cross-Component Linear Model (CCLM) Prediction To reduce redundancy of cross components, a cross-component linear model (CCLM) prediction mode is used in VVC, where chroma samples are predicted based on reconstructed luminance samples from the same CU using a linear model, as shown below: (1) in This represents the predicted chromaticity samples in the CU. This represents the downsampled reconstructed luminance samples from the same CU.

[0018] CCLM parameters ( and This is derived from a maximum of four neighboring chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block size is W×H, then W' and H' are set to... -W' = W, H' = H, when applying CCLM_LT mode; -W' = W + H, when applying CCLM_T mode; -H' = H + W, when applying CCLM_L mode.

[0019] The upper nearest neighbor positions are denoted as S[0,–1]…S[W'–1,–1], and the left nearest neighbor positions are denoted as S[–1,0]…S[–1,H'–1]. Then, four samples are selected as... -S[W' / 4,–1],S[3 W' / 4,–1],S[–1,H' / 4],S[–1,3 H' / 4], when LM mode is applied and the samples above and to the left of the neighboring sample are available; -S[W' / 8,–1],S[3 W' / 8,–1],S[5 W' / 8,–1],S[7 W' / 8,–1], when applying LM-A mode or when only the above-mentioned neighboring samples are available; -S[–1,H' / 8],S[–1,3 H' / 8],S[–1,5 H' / 8],S[–1,7 H' / 8], when applying LM-L mode or when only the left neighboring sample is available.

[0020] Four neighboring brightness samples at the selected location were downsampled and compared four times to find the two larger values: x 0 A and x 1 A and two smaller values: x 0 B and x 1 B Their corresponding chromaticity sample values ​​are respectively represented as y 0 A , y 1 A , y 0 B and y 1 B .Then x A , x B , y A and y B Exported as: x A =(x 0 A + x 1 A +1)>>1; x B =(x 0 B + x 1 B +1)>>1; y A =(y 0 A + y 1 A +1)>>1; yB =(y 0 B + y 1 B +1)>>1; (1) Finally, the linear model parameters are obtained according to the following equation. and .

[0021] (2) (3) Figure 4 This shows an example of the positions of the left and top samples, as well as the current block sample, involved in CCLM_LT mode. Figure 4 Showing Chroma block 410, corresponding The relative sample positions of luminance block 420 and its neighboring samples (displayed as solid circles).

[0022] Calculation parameters The division operation is implemented using a lookup table. To reduce the memory required to store the table, the difference ( diff Values ​​(difference between maximum and minimum values) and parameters It is expressed using an exponent. For example, the difference is approximated using four significant digits and an exponent. Therefore, the table for 1 / diff is reduced to 16 elements, corresponding to 16 significant digit values, as shown below: DivTable[] ={ 0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}, (4) This will help reduce the complexity of the calculations and the memory size required to store the tables.

[0023] In addition to the templates mentioned above and the left-hand template being used together to calculate linear model coefficients, they can also be used alternately in two other CCLM modes (referred to as CCLM_T and CCLM_L modes).

[0024] In CCLM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In CCLM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples.

[0025] In CCLM_LT mode, the left and top templates are used to calculate the coefficients of the linear model.

[0026] To match the chroma sample positions in a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type-0" and "type-2" respectively.

[0027] (6) (7) Note that when the upper reference line is at the CTU boundary, only one luminance line (the universal line buffer in intra-frame prediction) is used to generate downsampled luminance samples.

[0028] This parameter calculation is performed as part of the decoding process, not just the encoder search operation. Therefore, the α and β values ​​are not passed to the decoder using syntax.

[0029] For chroma intra-mode encoding and decoding, a total of eight intra-modes are allowed. These modes include five traditional intra-modes and three cross-component linear model modes (CCLM_LT, CCLM_T, and CCLM_L). The chroma mode signaling and derivation process are shown in Table 2. Chroma mode encoding and decoding directly depends on the intra-prediction mode of the corresponding luma block. Due to the independent block partitioning structure for luma and chroma components enabled in the I slice, one chroma block may correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra-prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.

[0030] Table 2. Deriving Chroma Prediction Mode from Luminance Mode when CCLM is Enabled

[0031] Regardless of the value of sps_cclm_enabled_flag, a single binary table is used, as shown in Table 3.

[0032] Table 3 – Unified Binarization Table for Chromaticity Prediction Modes

[0033] In Table 3, the first bit (bin) indicates whether it is normal mode (0) or CCLM mode (1). If it is LM mode, the next bit indicates CCLM_LT (0) or not. If it is not CCLM_LT, the next 1-bit indicates whether it is CCLM_L (0) or CCLM_T (1). In this case, when sps_cclm_enabled_flag is 0, the first bit of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy encoding and decoding. In other words, the first bit is inferred to be 0 and therefore not encoded. This single binarization table is used for the cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two bits in Table 4 are context-encoded using their own context model, while the remaining bits are side-encoded.

[0034] Furthermore, to reduce luma-chroma latency in dual-tree systems, when the 64x64 luma codec tree node is not segmented (and ISP is not used for the 64x64 CU) or QT segmented, the chroma CU in the 32x32 / 32x16 chroma codec tree node can use CCLM in the following manner: If the 32x32 chroma node is not segmented or is segmented into QT segments, all chroma CUs in the 32x32 node can use CCLM.

[0035] If a 32x32 chroma node is split into horizontal BTs, and the 32x16 child nodes are not split or are split using vertical BTs, all chroma CUs in the 32x16 chroma node can use CCLMs.

[0036] Under all other luma and chroma codec tree splitting conditions, chroma CUs are not allowed to use CCLM.

[0037] Multiple Model CCLM (MMLM) In JEM (J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Video Exploration Team (JVET), July 2017), a multi-model CCLM mode (MMLM) was proposed to predict chroma samples of the entire CU from luminance samples using two models. In MMLM, the neighboring luminance samples and neighboring chroma samples of the current block are divided into two groups, each used as a training set to derive a linear model (i.e., derive specific α and β for a specific group). Furthermore, samples of the current luminance block are also classified according to the classification rules of neighboring luminance samples. Three MMLM model modes (MMLM_LT, MMLM_T, and MMLM_L) allow selection of neighboring samples from the left and top, top only, and left only, respectively.

[0038] Figure 5 This shows an example of classifying neighboring samples into two groups. Threshold ( Threshold The value is calculated as the average of the neighboring reconstructed brightness samples. Neighboring sample Rec′ L [x,y]<= Threshold It was classified as group 1; while the neighboring sample Rec′ L [x,y]> Threshold It was classified as group 2.

[0039] (8) Therefore, MMLM uses two models based on the sample level of neighboring samples.

[0040] Slope adjustment of CCLM CCLM uses a two-parameter model to map luminance values ​​to chrominance values, such as... Figure 6A As shown. The slope parameter "a" and the bias parameter "b" define the following mapping: chromaVal = a lumaVal + b, The slope parameter "u" is adjusted by signaling to update the model in the following form, such as... Figure 6B As shown: chromaVal = a' lumaVal + b', in a' = a + u, b' = b - u y r .

[0041] Through this selection, the mapping function revolves around a value with brightness y. r The point is tilted or rotated. The average value of the reference brightness sample used for model creation is used as y. r This allows for meaningful modifications to the model. Figure 6A and 6B The process is explained.

[0042] Implementation of CCLM slope adjustment The slope adjustment parameter is provided as an integer between -4 and 4 and is signaled in the bitstream. The unit of the slope adjustment parameter is (1 / 8) of the chroma sample value per luminance sample value (for 10-bit content).

[0043] The adjustment applies to CCLM models that use reference samples above and to the left of the block (e.g., "LM_CHROMA_IDX" and "MMLM_CHROMA_IDX"), but not to "one-sided" modes. This choice is based on a trade-off between encoding / decoding efficiency and complexity. "LM_CHROMA_IDX" and "MMLM_CHROMA_IDX" refer to CCLM_LT and MMLM_LT in this invention. "One-sided" modes refer to CCLM_L, CCLM_T, MMLM_L, and MMLM_T in this invention.

[0044] When applying slope adjustment to a multi-mode CCLM model, two models can be adjusted, so at most two slope updates are signaled for a single chroma block.

[0045] CCLM slope adjustment encoder method The proposed encoder method performs a search based on the Sum of Absolute Transformed Differences (SATD) to find the optimal value for the slope update of Cr, and performs a similar SATD-based search to find the optimal value for the slope update of Cb. If either result is a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based update of Cr, SATD-based update of Cb) is included in the rate-distortion (RD) checklist of TU.

[0046] Convolutional Cross-Component Model (CCCM) - Single Model and Multiple Models In CCCM, a convolutional model is applied to improve chromaticity prediction performance. The convolutional model has a 7-tap filter, consisting of 5 taps plus a sign-shape spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of a center (C) luminance sample co-located with the chromaticity sample to be predicted, and its top / north (N), bottom / south (S), left / west (W), and right / east (E) neighbors, as shown below. Figure 7 As shown.

[0047] The non-linear term (denoted as P) is represented as the square of the center brightness sample C, scaled to the range of sample values ​​for the content: P = (C C + midVal )>>bitDepth.

[0048] For example, for 10-bit content, the non-linear term is calculated as follows: P = (C) C + 512)>>10.

[0049] The offset term (denoted as B) represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).

[0050] The output of the filter is calculated as the filter coefficients c. i Convolve the input values ​​and crop them to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B.

[0051] Filter coefficients c i It is calculated by minimizing the MSE between the predicted and reconstructed chromaticity samples in the reference region. Figure 8 An example of a reference region is shown, consisting of six rows of chroma samples above and to the left of the PU. The reference region extends to the right by one PU width and below the PU boundary by one PU height. The region is adjusted to contain only available samples. The extension around this region (referred to as "padding") is used to support the "side samples" of the plus-shaped spatial filter in Figure 7, and padding is applied when the region is unavailable.

[0052] MSE minimization is performed by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and chrominance output. The autocorrelation matrix is ​​decomposed using LDL, and the final filter coefficients are calculated by back substitution. This process largely follows the calculation of ALF filter coefficients in ECM, but LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. In newer ECMs, CCCM uses a Gaussian elimination-based method for MSE minimization.

[0053] In addition, similar to CCLM, CCCM can be configured to use either a single-model or multi-model variant. The multi-model variant uses two models: one for samples above the average luminance reference value and the other for the remaining samples (following the spirit of CCLM design). For PUs with at least 128 reference samples, the multi-model CCCM mode can be selected.

[0054] Gradient Linear Model (GLM) For the YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chromaticity samples from the gradient of luminance samples. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.

[0055] Compared to CCLM, two-parameter GLM uses the gradient of luminance samples, rather than downsampled luminance values, to derive a linear model. Specifically, when applying two-parameter GLM, the input to the CCLM process is the downsampled luminance samples... The gradient of the brightness sample The rest of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remains unchanged.

[0056] .

[0057] In the three-parameter GLM, chromaticity samples can be predicted based on the luminance sample gradient and downsampled luminance values ​​with different parameters. The model parameters of the three-parameter GLM are derived from 6 rows and 6 columns of adjacent samples using the MSE minimization method based on LDL decomposition used in CCCM. .

[0058] In terms of communication, when the current CU is in CCLM mode, a flag is sent to indicate whether both Cb and Cr components are in GLM mode; if GLM is in GLM mode, another flag is sent to indicate which of the two GLM modes is selected, and a syntax element is sent to select one of the four gradient filters for gradient calculation.

[0059] As shown in Figure 9, GLM enables four gradient filters (910-940).

[0060] Template-based Intra Mode Derivation (TIMD) Template-based intra-prediction (TIMD) modes implicitly derive the intra-prediction mode of the CU from the neighboring templates of the encoder and decoder, rather than sending the exact intra-prediction mode bits to the decoder. As shown in Figure 10, the template prediction samples (1012 and 1014) of the current block 1010 are generated using the template reference samples (1020 and 1022) for each candidate mode. The cost is calculated as the SATD between the template prediction sample and the reconstructed sample. The intra-prediction mode with the lowest cost is selected as the TIMD mode and used for intra-prediction of the CU. The candidate modes can be the 67 intra-prediction modes in the VVC or extended to 131 intra-prediction modes. Typically, the MPM can provide some clues indicating the orientation information of the CU. Therefore, in order to reduce the intra-prediction mode search space and take advantage of the characteristics of the CU, the intra-prediction mode is implicitly derived from the MPM list.

[0061] For each intra-prediction mode in the MPM, the SATD between the template prediction sample and the reconstructed sample is calculated. The top two intra-prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are fused with weights after applying PDPC processing, and this weighted intra-prediction is used to encode and decode the current CU. Position-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.

[0062] The costs of the two selected modes are compared with a threshold. In the test, cost factor 2 is applied as follows: costMode2<2 costMode1.

[0063] If this condition is met, fusion is applied; otherwise, only Mode 1 is used. The weights of the modes are calculated based on their SATD costs as follows: weight1 = costMode2 / (costMode1+ costMode2), weight2 = 1 - weight1.

[0064] This invention discloses a method and apparatus for improving the performance of cross-component prediction encoding and decoding. [Summary of the Invention] A method and apparatus for video encoding and decoding using an encoding / decoding tool that includes one or more cross-component model correlation modes are disclosed. According to the method, input data associated with a current block is received, the current block including a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or encoded data associated with the current block to be decoded at the decoder end. A basic predictor for the second color block is determined, wherein the basic predictor corresponds to an intra-frame predictor or an inter-frame predictor. One or more hypotheses for the second color block are determined, wherein each of the one or more hypotheses corresponds to a cross-component prediction candidate. A final predictor is derived by mixing the basic predictor and the one or more hypotheses. The second color block is encoded or decoded using prediction data containing the final predictor.

[0066] In one embodiment, a final predictor is derived by mixing a base predictor and the one or more hypotheses using predefined weights. In another embodiment, one or more flags are signaled or parsed to indicate whether the one or more hypotheses are added to derive the final predictor. In yet another embodiment, when the one or more flags indicate that the one or more hypotheses should be added to derive the final predictor, one or more syntaxes related to the one or more hypotheses and / or related to the predefined weights used to mix the base predictor and the one or more hypotheses are signaled or parsed.

[0067] In one embodiment, a cross-component prediction candidate is selected from a cross-component merging candidate list. In one embodiment, the candidate index associated with the cross-component prediction candidate selected from the cross-component merging candidate list is explicitly signaled or parsed, or implicitly derived. In one embodiment, the first candidate in the cross-component merging candidate list is selected as the cross-component prediction candidate. In another embodiment, the candidate in the cross-component merging candidate list with the minimum template matching cost or boundary matching cost is selected as the cross-component prediction candidate. In yet another embodiment, member candidates in the cross-component merging candidate list are classified into one or more categories, and the selection of the one or more hypotheses depends on the one or more categories.

[0068] In one embodiment, one or more prediction patterns selected for generating the one or more hypotheses belong to a different category than one or more previously selected patterns. In one embodiment, the one or more categories include a linear model category and a convolutional model category. In another embodiment, the one or more categories include a single-model category and a multi-model category.

[0069] In one embodiment, only one or more member candidates belonging to one or more specific categories in the cross-component merging candidate list are allowed to generate the one or more hypotheses. In another embodiment, the one or more specific categories correspond to cross-component convolution model patterns. In one embodiment, if the base predictor corresponds to a cross-component pattern, then the one or more specific categories do not include spatially inherited candidate patterns.

[0070] In one embodiment, the final predictor is derived by mixing the base predictor with the one or more hypotheses using a weighted summation that includes weighting parameters. In one embodiment, the weighting parameters are selected from a set of weighting parameters. In one embodiment, the weighting parameters are indicated by a signaled or parsed syntax. In another embodiment, the weighting parameters are implicitly determined. For example, the weighting parameters can be implicitly determined based on pattern information from one or more neighboring blocks, cross-component prediction patterns used to generate the one or more hypotheses, or template matching cost or boundary matching cost. In one embodiment, the weighting parameters are derived by minimizing the MSE between the predicted second color sample of the final predictor and the reconstructed second color sample in one or more reference regions.

[0071] In one embodiment, the base predictor is derived based on cross-component prediction hypotheses, each of which, along with the base predictor, is based on an encoding pattern. In another embodiment, the base predictor is derived from a cross-component prediction candidate selected from a cross-component merging candidate list. [Attached Image Description] Figure 1A An exemplary adaptive inter-frame / intra-frame video coding system incorporating loop processing is shown.

[0073] Figure 1B It shows Figure 1A The corresponding decoder of the encoder.

[0074] Figure 2 This illustrates the intra-frame prediction mode used in the VVC video codec standard.

[0075] Figures 3A and 3B show examples of wide-angle intra-frame prediction for blocks with a width greater than their height (Figure 3A) and blocks with a height greater than their width (Figure 3B), respectively.

[0076] Figure 4 This shows an example of the top-left sample and sample position of the current block in CCLM_LT mode.

[0077] Figure 5 An example of classifying neighboring samples into two groups is shown.

[0078] Figure 6AAn example of the CCLM model is shown.

[0079] Figure 6B An example of the effect of the slope adjustment parameter "u" used for model updates is shown.

[0080] Figure 7 An example of the spatial portion of a convolution filter is shown.

[0081] Figure 8 An example of a padded reference region used to derive filter coefficients is shown.

[0082] Figure 9 Four gradient patterns are shown for the gradient linear model (GLM).

[0083] Figure 10 shows an example of a template-based intra-frame mode derivation (TIMD) mode, where TIMD implicitly derives the intra-frame prediction mode of the CU using neighboring templates at the encoder and decoder.

[0084] Figure 11 An example of CCM information propagation is shown, where the dashed blocks are encoded in cross component modes (e.g., CCLM, MMLM, GLM, CCCM).

[0085] Figure 12A An example of the parameters of the inheritance space proximity model is shown. Figure 12B An example of inheritance time proximity model parameters is shown.

[0086] Figure 13A -B illustrates two search patterns for inheriting the non-adjacent spatial proximity model.

[0087] Figures 14A-B show historical records based on regions with the same starting geometric position as the current region (Figure 14A) or historical records based on regions containing the center geometric position of the current region (Figure 14B). Figure 14B Example of constructing the current region's historical record table.

[0088] Figure 15 An example of a neighboring template used to calculate model error is shown.

[0089] Figure 16 shows an example of candidate reordering based on template matching cost, where reconstructed samples on the current block template are used as gold data.

[0090] Figure 17 shows an example of pattern classification based on filter shape, where the convolutional model can have 3 filter shape options.

[0091] Figure 18 shows a flowchart of an exemplary video codec system using a multi-hypothesis mixing method for cross-component model merging according to an embodiment of the present invention.

Detailed Implementation Methods

[0093] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the invention can be practiced without one or more specific details, or using other methods, components, etc. In other instances, well-known structures or operations have not been shown or described in detail to avoid obscuring various aspects of the invention. Embodiments of the invention will be best understood by referring to the accompanying drawings, in which similar parts are designated with similar numerals throughout. The following description is merely illustrative, simply illustrating certain selected apparatus and method embodiments consistent with the invention claimed herein.

[0094] To improve the prediction accuracy or encoding / decoding performance of cross-component prediction, several schemes related to inherited cross-component models have been disclosed.

[0095] Inheriting neighbor model parameters to refine cross component model parameters The final scaling parameters of the current block are inherited from neighboring blocks and further refined via dA (e.g., the derivation or signaling of dA can be similar to or identical to the method described in the aforementioned "Guiding Parameter Set for Refining Cross Component Model Parameters"). Once the final scaling parameters are determined, the offset parameters (e.g., in CCLM) are... This is derived based on the inherited scaling parameters and the average of the luminance and chrominance samples from the current block. For example, if the final scaling parameters are inherited from selected neighboring blocks, and the inherited scaling parameters are... So the final scaling parameter is ( + dA). In another embodiment, the final scaling parameter is inherited from the history list and further refined by dA. For example, the history list records the j most recent final scaling parameter entries for previous CCLM codec blocks. The final scaling parameter is then selected from a chosen entry in the history list. Inheritance, and the final scaling parameter is ( + dA). In another embodiment, the final scaling parameter is inherited from the history list or neighboring blocks, but only the MSB (most significant bit) portion of the inherited scaling parameter is taken, and the LSB (least significant bit) of the final scaling parameter comes from dA. In yet another embodiment, the final scaling parameter is inherited from the history list or neighboring blocks, but is no longer refined by dA.

[0096] In another embodiment, after inheriting the model parameters, the offset parameters can be further refined in dB. For example, if the final offset parameters are inherited from selected neighboring blocks, and the inherited offset parameters are... So the final offset parameter is ( + dB). In another embodiment, the final offset parameter is inherited from the history list and further refined by dB. For example, the history list records the j most recent final scaling parameter entries for previous CCLM codec blocks. The final scaling parameter is then selected from a chosen entry in the history list. Inheritance, and the final scaling parameter is ( + dB). In another embodiment, the final offset parameter is inherited from the history list or neighboring blocks, but not further refined by dB.

[0097] In another embodiment, if the inherited neighboring blocks use CCCM encoding / decoding, then the inherited filter coefficients ( Offset parameters (e.g., in CCCM) or The filter coefficients can be re-derived based on the inherited parameters and the average values ​​of the luminance and chrominance samples at the corresponding neighboring locations of the current block. In another embodiment, only a portion of the filter coefficients are inherited (e.g., only n out of 7 filter coefficients are inherited, where...). The remaining filter coefficients are further re-derived using neighboring luminance and chrominance samples from the current block.

[0098] In another embodiment, if the inherited candidate applies the GLM gradient mode to its brightness reconstruction sample, then the current block should also inherit the candidate's GLM gradient mode and apply it to the current brightness reconstruction sample.

[0099] In another embodiment, if the inherited neighboring blocks use multiple cross-component model encoding / decoding (e.g., MMLM or CCCM with multiple models), the classification threshold is also inherited to classify the neighboring samples of the current block into multiple groups, and the inherited multiple cross-component model parameters are further assigned to each group. In another embodiment, the classification threshold is the average of the neighboring reconstructed luminance samples, and the inherited multiple cross-component model parameters are further assigned to each group. Similarly, once the final scaling parameters for each group are determined, the offset parameters for each group are re-derived based on the inherited scaling parameters and the average of the neighboring luminance and chrominance samples for each group of the current block. For example, if a multi-model CCCM is used, once the final coefficient parameters for each group (e.g., in CCCM) are determined... arrive ,Apart from Then the offset parameter for each group (e.g., in CCCM) or Based on the inherited coefficient parameters and the neighboring luminance and chromaticity samples of each group in the current block, the values ​​are re-exported.

[0100] In another embodiment, the inherited model parameters may depend on the color components. For example, the Cb and Cr components may inherit model parameters or model derivation methods from the same or different candidate models. As another example, only one color component inherits model parameters, while the other color component derives its model parameters based on the inherited model derivation method (e.g., if the inheritance candidate is encoded or decoded by MMLM or CCCM, the current block also derives its model parameters based on MMLM or CCCM using the current neighbor reconstructed samples). Yet another example, only one color component inherits model parameters, while the other color component derives its model parameters using the current neighbor reconstructed samples.

[0101] For example, if the Cb and Cr components can be derived from different candidate inheritance model parameters or model derivation methods, the inheritance model of Cr can depend on the inheritance model of Cb. For example, possible cases include, but are not limited to: (1) if the inheritance model of Cb is CCCM, then the inheritance model of Cr should be CCCM; (2) if the inheritance model of Cb is CCLM, then the inheritance model of Cr should be CCLM; (3) if the inheritance model of Cb is MMLM, then the inheritance model of Cr should be MMLM; (4) if the inheritance model of Cb is CCLM, then the inheritance model of Cr should be CCLM or MMLM; (5) if the inheritance model of Cb is MMLM, then the inheritance model of Cr should be CCLM or MMLM; (6) if the inheritance model of Cb is GLM, then the inheritance model of Cr should be GLM.

[0102] In another embodiment, after decoding a block, the Cross-Component Model (CCM) information of the current block is exported and stored so that neighboring blocks can be reconstructed subsequently using inherited neighbor model parameters. The CCM information mentioned in this disclosure includes, but is not limited to, prediction modes (e.g., CCLM, MMLM, CCCM), GLM pattern indexes, model parameters, or classification thresholds. For example, even if the current block uses inter-frame predictive coding, the cross-component model parameters of the current block can be derived using the current luma and chroma reconstruction or prediction samples. Subsequently, if another block is predicted using inherited neighbor model parameters, it can inherit the model parameters from the current block. As another example, if the current block uses cross-component predictive coding, the cross-component model parameters of the current block are re-derived using the current luma and chroma reconstruction or prediction samples. As yet another example, the stored cross-component model can be CCCM, CCLM_LT (i.e., a single-model LM derived using upper and left neighbor samples), or MMLM_LT (a multi-model LM derived using upper and left neighbor samples). For example, even if the current block is encoded using non-cross-component intra-predictive coding (e.g., DC, planar, intra-angle mode, MIP, or ISP), the cross-component model parameters for the current block are derived using the current luma and chroma reconstructed or predicted samples. Alternatively, even if the current block is encoded using cross-component predictive coding, the cross-component model parameters for the current block are re-derived using the current luma and chroma reconstructed or predicted samples. The re-derived model parameters are then combined with the original cross-component model used to reconstruct the current block. To combine with the original cross-component model, the model combination methods mentioned in the section titled "Inheriting Multiple Cross-Component Models" can be used. For example, suppose the original cross-component model parameters are... The re-exported cross-component model parameters are The final cross-component model is as follows: ,in It is a weighting factor that can be predefined or implicitly derived based on the cost of neighboring templates.

[0103] In another embodiment, when inheriting a cross-component model from a neighbor merge candidate encoded and decoded by cross-component modes (e.g., CCLM and CCCM), a flag can be sent to indicate / select whether to use a re-derived model. If the flag is 0, the cross-component model used to encode the neighbor merge candidate is inherited. If the flag is 1, the cross-component model re-derived based on the luminance and chrominance reconstruction or prediction samples of the neighbor merge candidate is inherited.

[0104] For example, when the current slice is a non-intra-frame slice (e.g., a P-slice or a B-slice), the cross-component model of the current block is exported and stored so that adjacent blocks can be reconstructed later using inherited neighbor model parameters. In another embodiment, when the current block is inter-frame encoded / decoded, the CCM information of the current inter-frame encoded / decoded block is obtained by copying the CCM information from its reference block, which is located in a reference image containing the CCM information and is positioned by the motion information of the current inter-frame encoded / decoded block. For example, as shown in FIG11, if block B in P / B image 1120 is inter-frame encoded / decoded, the CCM information of block B is obtained by copying the CCM information from its reference block A in I image 1110. It should be noted that the current block can also copy CCM information from intra-frame encoded / decoded blocks in the P / B image. For example, as... Figure 11 As shown, if block D in P / B image 1130 is inter-frame encoded, the CCM information of block D is obtained by copying the CCM information from its reference block E, which is intra-frame encoded in P / B image 1120. In another embodiment, if a reference block in a reference image is also inter-frame encoded, the CCM information of the reference block is obtained by copying the CCM information from another reference block in another reference image. For example, as... Figure 11 As shown, in the current P / B image 1130, the current block C uses inter-frame coding, and its reference block B also uses inter-frame coding. Since the CCM information of block B is obtained by copying the CCM information of block A, the CCM information of block A will also be propagated to the current block C. In another embodiment, when the current block uses bidirectional predictive inter-frame coding, if one of its reference blocks is intra-frame coding and contains CCM information, the CCM information of the current block is obtained by copying the CCM information of its intra-frame coding reference block in the reference image. For example, suppose block F uses bidirectional predictive inter-frame coding and has reference blocks G and H. Block G is intra-frame coding and contains CCM information. The CCM information of block F is obtained by copying the CCM information of block G, which is coded in CCM mode. In yet another embodiment, when the current block uses bidirectional prediction for inter-frame coding, the CCM information of the current block is a combination of the CCM models of its reference blocks (as mentioned in the section titled "Inheriting Multiple Cross Component Models").

[0105] When deriving a cross-component model for the current block using the current luminance and chrominance reconstructed or predicted samples, in one embodiment, if the error of the currently derived model is greater than a threshold, the currently derived model is discarded and not stored. For example, the current luminance reconstructed sample can be input into the model, the distortion between the model output and the current chrominance reconstructed sample can be calculated, and then the calculated distortion can be normalized using the current block size or number of samples used to calculate the distortion. If the normalized distortion is greater than or equal to the threshold, the currently derived model is discarded and not stored.

[0106] Whether to derive a cross-component model for the current block can depend on the size or area of ​​the current block. For example, for small blocks (e.g., block width / height less than or equal to a threshold, or block area less than or equal to a threshold), deriving a cross-component model is not allowed. Similarly, for large blocks (e.g., block width / height greater than or equal to a threshold, or block area greater than or equal to a threshold), deriving a cross-component model is not allowed.

[0107] Inherit CCM information In one embodiment, the cross-component model (CCM) information of the inherited cross-component model can be stored along with the inherited model parameters. As described above in this disclosure, CCM information includes, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM), model indices indicating which model shape is used in the convolutional model, classification thresholds for multiple models, downsampling filter flags, downsampling filter indices, the number of neighboring lines used to derive the model, template types used to derive the model, post-filtering flags, or model parameters.

[0108] In one embodiment, a CCLM model can be inherited. In addition to storing model parameters, the prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCLM model.

[0109] In another embodiment, a CCLM model with nonlinear terms can be inherited. In addition to storing model parameters, the CCM information can also store prediction patterns indicating that the inherited model is a CCLM model with nonlinear terms.

[0110] In one embodiment, a CCCM model can be inherited. In addition to storing model parameters, the CCM information can also store prediction patterns to indicate that the inherited model is a CCCM model. Luminance and chromaticity offsets used to adjust the CCCM model inputs can also be stored in the CCM information.

[0111] In another embodiment, CCCM models with different convolutional filter shapes can be inherited. In addition to model parameters and prediction modes, the CCM information can also store CCCM mode indexes indicating which convolutional filter shape the inherited CCCM model uses. For example, a CCCM model with different convolutional filter shapes can only contain horizontal spatial terms. Another example is that a CCCM model with different convolutional filter shapes can only contain vertical spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes can only contain diagonal spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes can only contain anti-diagonal spatial terms. Finally, a CCCM model with different convolutional filter shapes can contain X-shaped spatial terms.

[0112] In another embodiment, a CCCM model using unsampled samples can be inherited. In addition to storing model parameters, prediction patterns can also be stored in the CCM information to indicate that the inherited model is a CCCM model using unsampled samples.

[0113] In another embodiment, a CCCM model with multiple downsampling filters can be inherited. In addition to storing model parameters, the prediction mode can also be stored in the CCM information to indicate that the inherited model is a CCCM model with multiple downsampling filters, and the model index can also be stored in the CCM information to indicate which variant of the CCCM model with multiple downsampling filters was inherited.

[0114] In another embodiment, a hybrid CCCM model consisting of various terms (e.g., spatial, gradient, positional, nonlinear, and bias terms) can be inherited. The gradient term can be computed in either a downsampled or non-downsampled domain. The positional term can be computed relative to the top-left corner coordinates of the current block or image. In addition to storing model parameters, prediction patterns can be stored in the CCM information to indicate that the inherited model is a hybrid CCCM model consisting of various terms. If multiple types of hybrid CCCM models exist, model indexes can also be stored in the CCM information to indicate which type of hybrid CCCM model was inherited. For example, the gradient and location-based cross-component model (GL-CCCM) proposed in JVET-AB0119 (Ramin G. Youvalari et al., “Non-EE2: Gradient and location based convolutional cross-component model (GL-CCCM) for intra prediction”, ITU-T SG16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Exploration Group (JVET), 28th meeting, Mainz, Germany, October 20-28, 2022, document: JVET-AB0119) is a hybrid CCCM model consisting of a spatial term located at the center, two gradient terms representing the horizontal and vertical directions respectively, two position terms X and Y representing the relative horizontal and vertical positions respectively, a nonlinear term, and a bias term. In addition to storing model parameters, prediction patterns can also be stored in the CCM information to indicate that the inherited model is a GL-CCCM model.

[0115] In one embodiment, a GLM model can be inherited. In addition to storing model parameters, the CCM information can also store a prediction mode indicating that the inherited model is a GLM model, and can further store a downsampling filter index indicating which gradient downsampling filter the inherited GLM model uses.

[0116] In another embodiment, a GLM model with a luminance term can be inherited. In addition to storing model parameters, the CCM information can also store a prediction mode indicating that the inherited model is a GLM model with a luminance term, and the CCM information can also store a downsampling filter index indicating which gradient downsampling filter is used by the inherited GLM model with a luminance term.

[0117] In one embodiment, any type of cross-component multi-model can be inherited. In addition to storing model parameters and prediction modes, the CCM information can also store a multi-model on / off flag to indicate whether the inherited CCM model is a multi-model. If the multi-model on / off flag is true, the multi-model classification threshold is also stored in the CCM information.

[0118] In one embodiment, the CCM information may contain information indicating how to derive the inherited model. For example, the CCM information may include the number of adjacent lines used to derive the cross component model and / or the template type used to derive the model. For example, a set of templates may be used to derive the CCCM model. This set of templates includes templates with different positions, sizes, and shapes. The CCM information may store indices of the templates on which the derived inherited CCCM model is based. For example, the inherited CCCM model may be derived based on a top-only template, a left-only template, or both a left-and-top template. As another example, the inherited CCCM model may be derived based on a 6-line template or a 2-line template.

[0119] In one embodiment, a post-filter flag can be stored in the CCM information. This information describes how the inherited model is used within its parent block. If the post-filter flag is enabled, it indicates that a filter has been applied to the predictions of the block to which the inherited model belongs.

[0120] Inheritance spatial proximity model parameters In another embodiment, the inherited model parameters can come from a neighboring block. Figure 12A An example of inheritance space proximity model parameters is shown. Models from blocks at predefined locations are added to the candidate list in a predefined order. For example, the predefined order could be B. 0, A 0, B 1, A1 and B2, or A 0, B 0, B 1, A1 and B2.

[0121] In another embodiment, the predefined positions include those positions immediately above (W>>1) or ((W>>1) - 1) (if W is greater than or equal to TH) (i.e., the position directly above the center of the top line of the current block. Assuming the position of the current block is (x, y), the predefined positions can be (x + (W>>1), y - 1) or (x + (W>>1) - 1, y - 1)), and those positions immediately to the left (H>>1) or ((H>>1) - 1) (if H is greater than or equal to TH) (i.e., the position immediately to the left row of the current block. Assuming the position of the current block is (x, y), the predefined positions can be (x - 1, y + (H>>1)) or (x - 1, y + (H>>1) - 1)), where W and H are the width and height of the current block, and TH is a threshold that can be 2, 4, 8, 16, 32, or 64.

[0122] In another embodiment, the maximum number of models inherited from spatial proximity is less than the number of predefined locations. For example, if the predefined locations are as follows: Figure 12A As shown, there are 5 predefined positions. If the predefined order is B... 0, A 0, B 1, If A1 and B2 are selected, and the maximum number of models inherited from spatial proximity is 4, then a model from B2 will only be added to the candidate list if one of the preceding blocks is unavailable or not encoded / decoded in the cross component model.

[0123] Inheritance time proximity model parameters In another embodiment, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images. For example, such as Figure 12B As shown, the current block is located at (x, y), and the block size is [size missing]. The inherited model parameters can come from blocks in previously encoded slices / images at positions (x', y'), (x', y'+ h / 2), (x' + w / 2, y'), (x' + w / 2, y' + h / 2), (x' + w, y'), (x', y' + h), or (x' + w, y' + h), where x' = x + Δx and y' = y + Δy. In one embodiment, if the prediction mode of the current block is intra-frame, Δx and Δy are set to 0. If the prediction mode of the current block is inter-frame, Δx and Δy are set to the horizontal and vertical motion vectors of the current block. In another embodiment, if the current block is inter-frame bidirectional prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 0. In yet another embodiment, if the current block is inter-frame bidirectional prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 1.

[0124] In another embodiment, if the current block is an inter-frame bidirectional prediction, the inherited model parameters can come from blocks in previously encoded / decoded slices / images in the reference list. For example, if the horizontal and vertical motion vectors in reference image list 0 are... and Then the motion vector can be scaled to other reference images in reference lists 0 and 1. If the motion vector is scaled to the first reference image in reference list 0... Reference image is The model can be derived from the first one in reference list 0. Referencing the block in the image, and setting Δx and Δy to... Another example is if the horizontal and vertical motion vectors in reference image list 0 are... and And the motion vector is scaled to the first one in reference list 1. Reference image is The model can be derived from the first one in reference list 1. Referencing the block in the image, and setting Δx and Δy to... .

[0125] Inheriting the non-adjacent spatial proximity model In another embodiment, the inherited model parameters can come from spatially neighboring blocks, especially non-adjacent neighboring blocks that are not immediately adjacent to the current block. Models from blocks at predefined locations are added to the candidate list in a predefined order. In another embodiment, the distance between locations closer to the current coded block is less than the distance between locations farther from the current block.

[0126] In another embodiment, the maximum number of models inherited from non-adjacent spatial neighbors is less than the number of predefined locations. For example, if the predefined locations are as follows: Figure 13A -B shows two patterns. Figure 13A Model 1310 and Figure 13B (Sample 1320 in the text). The maximum number of models that can be added to the candidate list if inherited from non-adjacent spatial neighbors is... Then only if the number of available models obtained from search sample 1 is less than Only then are models from positions in search pattern 2 added to the subsequent list.

[0127] Inheriting model parameters from history table In one embodiment, the inherited model parameters can come from a cross-component model history table. Cross-component models in the history table can be added to the candidate list in a predefined order. In one embodiment, the order of adding historical candidates can be from the beginning to the end of the table. In another embodiment, the order of adding historical candidates can be from a predefined position to the end of the table. In yet another embodiment, the order of adding historical candidates can be from the end to the beginning of the table. In yet another embodiment, the order of adding historical candidates can be staggered (e.g., the first candidate added is from the beginning of the table, the second candidate added is from the end of the table, and so on).

[0128] In one embodiment, a single cross-component model history table can be maintained to store previous cross-component models, and the cross-component model history table can be reset at the current image, current slice, current tile, every M CTU rows, or at the start of every N CTU, where N and M can be any values ​​greater than 0. In another embodiment, the cross-component model history table can be reset at the current image, current slice, current tile, current CTU row, or at the end of the current CTU.

[0129] In another embodiment, an image can be divided into multiple regions, and a history table is maintained for each region. History table 0 and an additional history table are updated during the encoding / decoding process. The additional history table can be determined by the current location. For example, if the current CU is located in the second region, the additional history table to be updated is history table 2.

[0130] In another embodiment, multiple history tables are used for different update frequencies. For example, the first history table is updated every CU, the second history table is updated every two CUs, the third history table is updated every four CUs, and so on.

[0131] In another embodiment, multiple history tables are used to store different types of cross-component models. For example, a first history table stores a single model, and a second history table stores multiple models. Another example is that a first history table stores gradient models, and a second history table stores non-gradient models. Yet another example is that a first history table stores simple linear models (e.g., y = ax + b, a model with fewer parameters), and a second history table stores complex models (e.g., CCCM, a model with more parameters).

[0132] In another embodiment, multiple history tables are used for different reconstructed luminance intensities. For example, if the average value of the reconstructed luminance samples in the current block is greater than a predefined threshold, the cross-component model is stored in the first history table; otherwise, the cross-component model is stored in the second history table. In another embodiment, multiple history tables are used for different reconstructed chrominance intensities. For example, if the average value of adjacent reconstructed chrominance samples in the current block is greater than a predefined threshold, the cross-component model is stored in the first history table; otherwise, the cross-component model is stored in the second history table.

[0133] In one embodiment, when adding historical candidates to a candidate list from multiple historical tables, the addition order may be from the beginning to the end of a table, and then adding the next historical table in the same or reverse order. In another embodiment, the addition order may be from the end to the beginning of a table, and then adding the next historical table in the same or reverse order. In yet another embodiment, the addition order may be from a predefined position in a table to the end of a table, and then adding the next historical table in the same or reverse order. In yet another embodiment, the addition order may be from a predefined position in a table to the beginning of a table, and then adding the next historical table in the same or reverse order. In yet another embodiment, the order of adding historical candidates may be staggered in a historical table (e.g., the first candidate added comes from the beginning of a historical table, the second candidate added comes from the end of a historical table, and so on), and then adding the next historical table in the same or reverse order.

[0134] In another embodiment, the addition order may be from the beginning to the end of each historical table. In another embodiment, the addition order may be from the end to the beginning of each historical table. In another embodiment, the addition order may be from a predefined position in each historical table to the end of each historical table. In another embodiment, the addition order may be from a predefined position in each historical table to the beginning of each historical table. In yet another embodiment, the order of adding historical candidates may be staggered within each specific historical table (e.g., the first candidate is added from the beginning of all historical tables, the second candidate is added from the end of all historical tables, and so on).

[0135] In one embodiment, multiple cross-component model history tables are used, but not all history tables are used to create the candidate list. Only history tables whose regions are close to the current block region can be used to create the candidate list.

[0136] In one embodiment, if historical candidates are used, the range of non-adjacent candidates can be reduced by using smaller distances between non-adjacent candidate positions. In another embodiment, if historical candidates are used, the number of non-adjacent candidates can be reduced by measuring the distance from the top-left corner of the current block to the candidate position, and then excluding candidates whose distance is greater than a predefined threshold. In another embodiment, if historical candidates are used, the number of non-adjacent candidates can be reduced by skipping candidates that are not in the same region. In another embodiment, if historical candidates are used, the number of non-adjacent candidates can be reduced by skipping candidates that are not located in a neighboring region. The range of the neighboring region is predefined and can be a region of size M multiplied by N, where M and N can be any values ​​greater than 0. In another embodiment, if historical candidates are used, the range of non-adjacent candidates can be reduced by skipping a second search pattern.

[0137] In another embodiment, an image can be divided into multiple regions, and at least one history table is maintained in each region. For a region of the current image, it can use or combine the history tables of one or more regions from previously encoded / decoded images as the initial history table. For example, if an image is divided into N regions, a history table can be implicitly or explicitly selected from one of the N regions of a previously encoded image as the initial history table. The index of one of the N regions can be signaled or implicitly derived from the corresponding region of the previously encoded / decoded image. Figure 14A As shown in -B, the current image 1420 is a P / B codec image, and the previous image 1410 is an intra-frame codec image. Each image is divided into four regions, as shown by four rectangular boxes. According to an embodiment of the present invention, the corresponding region in the previous codec image can be as follows: Figure 14A The region 1412 shown has the same starting geometric position as the current region 1422, or includes the region 1412 containing the center geometric position of the current region 1422 as shown in Figure 22B. For example, it can combine multiple history tables from previously encoded / decoded regions / images to construct a history table for the current region.

[0138] Candidate list construction In one embodiment, a candidate list is constructed by adding candidates in a predefined order until a maximum number of candidates is reached. The added candidates may include, but are not limited to, all of the candidates described above. For example, the candidate list may include spatially neighboring candidates, temporally neighboring candidates, historical candidates, non-neighboring neighboring candidates, and single-model candidates generated based on other inherited models or combined models (as mentioned in later sections: inheriting multiple cross-component models). In another example, the candidate list may include the same candidates as in the previous example, but the candidates are added to the list in a different order.

[0139] In another embodiment, if all predefined neighboring and historical candidates are added but the maximum number of candidates is not reached, some default candidates are added to the candidate list until the maximum number of candidates is reached.

[0140] In one sub-implementation, the default candidates include, but are not limited to, those described below. Final scaling parameters. From collection And offset parameter Alternatively, it can be derived based on neighboring luminance and chrominance samples. For example, if the average values ​​of neighboring luminance and chrominance samples are lumaAvg and chromaAvg, then... pass Export. Average value of neighboring brightness samples ( This can be achieved by using all selected luminance samples, the current luminance CB's luminance DC mode value, or the average of the maximum and minimum luminance samples (e.g., ,or ) calculation. Similarly, the average value of neighboring chromaticity samples (i.e., This can be done by using all selected chromaticity samples, the chromaticity DC mode value of the current chromaticity CB, or the average of the maximum and minimum chromaticity samples (e.g., ,or )calculate.

[0141] In another sub-implementation, the default candidate includes, but is not limited to, the candidates described below. The default candidate is... ,in It is the gradient of the brightness sample, not the downsampled brightness sample. Sixteen GLM filters, described in the chapter entitled "Gradient Linear Model (GLM)," were applied. The final scaling parameters... From collection Offset parameter Alternatively, it can be derived based on nearby luminance and chrominance samples.

[0142] In another embodiment, the default candidate can be based on an earlier candidate in the candidate list with incremental scaling parameter refinement. For example, if the scaling parameter of the earlier candidate is , then the scaling parameter of the default candidate is , where can be from the set { . The offset parameter of the default candidate will be derived by and the average of the neighboring luminance and chrominance samples of the current block.

[0143] In another embodiment, the default candidate can be a shortcut indicating a cross-component mode (i.e., using the current neighboring luminance / chrominance reconstruction samples to derive the cross-component model), rather than inheriting parameters from the neighbors. For example, the default candidate can be CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T, single-model CCCM, multi-model CCCM, or a cross-component model with specified GLM-type samples.

[0144] In another embodiment, the default candidate can be a cross-component mode (i.e., using the current neighboring luminance / chrominance reconstruction samples to derive the cross-component model), rather than inheriting parameters from the neighbors, and also has a scaling parameter update ( ). Then the scaling parameter of the default candidate is . For example, the default candidate can be CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, or MMLM_T. For another example, can be from the set . The offset parameter of the default candidate will be derived by and the average of the neighboring luminance and chrominance samples of the current block. For another example, for each color component, can be different.

[0145] In another embodiment, the default candidate can be an earlier candidate with some selected model parameters. For example, assume the earlier candidate has m parameters, and it can select k out of the m parameters from the earlier candidate as the default candidate, where 0 < k < m and m > 1.

[0146] In another embodiment, the default candidate can be the first model of an earlier MMLM candidate (i.e., the model used when the sample value is less than or equal to the classification threshold). In another embodiment, the default candidate can be the second model of an earlier MMLM candidate (i.e., the model used when the sample value is greater than or equal to the classification threshold). In yet another embodiment, the default candidate can be a combination of the two models of an earlier MMLM candidate. For example, if the models of the earlier MMLM candidate are and The default candidate model parameters can be... ,in It is a weighting factor that can be predefined or implicitly derived based on the cost of neighboring templates. It is the x-th parameter of the y-th model.

[0147] In another embodiment, default candidates can be derived from reconstructed samples in non-adjacent neighborhoods. Let the current block position be (x, y) and the block size be w×h. If a reconstructed sample in the MxN region (x+dx, y+dy) is "available," then the reconstructed luma and chroma samples in that region can be used to derive default candidates. For example, MxN can be 8x8. Or, for example, MxN can be 16x8. Or, for example, MxN can be 16x16. Or, for example, MxN can be w×h. "Available" can mean that a reconstructed sample within the current block is available, or that a reconstructed sample within k rows of neighboring samples is available. k can be defined by the neighborhood search region of the IBC or the neighborhood buffer of other intra-frame coding / decoding tools (e.g., multi-reference line intra-prediction, CCLM, or CCCM).

[0148] In another embodiment, let the current block position be (x, y) and the block size be w×h. If a reconstructed sample is available in this region, then the sample located at (x, y) can be used. mid +dx, y mid The reconstructed samples in the MxN region of (x + dy) derive the default candidate, where (x mid , y mid ) = (x + w / 2, y + h / 2).

[0149] In another embodiment, the default candidate derived from the reconstructed samples in non-adjacent regions can be any type of cross-component model or certain specific types of cross-component models. For example, the derived model can be CCLM, MMLM, CCCM, CCCM multi-model, or other cross-component models. For another example, the derived model is a CCCM model. For another example, the derived model is a CCLM model. For another example, the derived model is a CCCM or CCCM multi-model.

[0150] In another embodiment, assume two sets of values. and Defined as: , .

[0151] and All values ​​in (dx, dy) are positive. .

[0152] In another embodiment, the current block position is (x, y), and the block size is w × h. Let... and Given two fixed positive numbers (dx, dy), it can be: or .

[0153] When constructing the candidate list, candidates are inserted into the list according to a predefined order. For example, the predefined order could be spatially neighboring candidates, temporal candidates, spatially non-neighboring candidates, historical candidates, and then the default candidate. In one embodiment, if a cross-component model is derived for a non-LM codec block (e.g., as mentioned in the section entitled "Inheriting Neighbor Model Parameters to Refine Cross-Component Model Parameters"), the candidate models for the non-LM codec block are included in the list after the candidate models for the LM codec block. In another embodiment, if a cross-component model is derived for a non-LM codec block, the candidate models for the non-LM codec block are included in the list before the default candidate is included. In yet another embodiment, if a cross-component model is derived for a non-LM codec block, the candidate models for the non-LM codec block are included in the list with lower priority than the candidate models for the LM codec block.

[0154] When constructing the candidate list, only candidates with a specific prediction mode can be added to the list. For example, it can restrict that only candidates derived through CCLM or MMLM modes are allowed to be added to the list. For another example, it can restrict that only candidates derived through a single-model mode (e.g., CCLM or CCLM with a single model) are allowed to be added to the list. For yet another example, it can restrict that only candidates derived through a multi-model mode (e.g., MMLM or CCCM with a multi-model) are allowed to be added to the list. For yet another example, it can restrict that only candidates derived through a GLM mode are allowed to be added to the list. For yet another example, it can restrict that only candidates derived through a specific mode (e.g., CCLM, MMLM, CCCM, CCCM with a multi-model, or GLM) are allowed to be added to the list. In one embodiment, if only candidates with a specific prediction mode can be added to the list, during prediction mode signaling, the prediction mode can be signaled first, followed by a signal indicating whether the proposed cross-component merging mode should be used. If the proposed cross-component merging mode is used, a candidate index signal is sent.

[0155] In chroma intra-frame fusion mode, intra-frame predictions from non-CCLM codecs and CCLM codecs are fused together to obtain the final intra-frame prediction. In one embodiment, when inheriting cross-component model parameters from the block / position of the chroma intra-frame fusion mode codec, the model parameters used to obtain the intra-frame prediction for CCLM codecs are inherited and further refined. In another embodiment, the fusion weights, the codec mode of the intra-frame prediction for non-CCLM codecs, and the model parameters used to obtain the intra-frame prediction for CCLM codecs are inherited and further refined. In yet another embodiment, the codec mode of the intra-frame prediction for non-CCLM codecs is implicitly derived (e.g., derived as DM or planar mode), and the fusion weights and the model parameters used to obtain the intra-frame prediction for CCLM codecs are inherited and further refined. In another embodiment, if the non-CCLM codec intra-prediction of the block / position encoded / decoded by the chroma intra-fusion mode can be implicitly derived (e.g., the non-CCLM codec intra-prediction is DM or planar mode), the fusion weights and model parameters used to obtain the CCLM codec intra-prediction can be inherited and further refined.

[0156] Candidates in the reordered list Candidates in the list can be reordered to reduce the syntax overhead when selecting candidate indices for signaling. Reordering rules can rely on encoding / decoding information or model errors of neighboring blocks. For example, if the neighboring upper or left block is encoded using MMLM, MMLM candidates in the list can be moved to the beginning of the current list. Similarly, if the neighboring upper or left block is encoded using Single Model LM or CCCM, Single Model LM or CCCM candidates in the list can be moved to the beginning of the current list. Likewise, if the neighboring upper or left block uses GLM, GLM-related candidates in the list can be moved to the beginning of the current list.

[0157] In another embodiment, the reordering rule is based on model error. A candidate model is applied to the neighboring templates of the current block, and then the error is compared with reconstructed samples from the neighboring templates. For example, as... Figure 15 As shown, the size of the template 1520 adjacent to the current block is... The size of the left adjacent template 1530 of the current block 1510 is Suppose there are K models in the current candidate list. and These are the final scaling and offset parameters after inheriting from candidate k. The model error of candidate k corresponding to the upper neighboring template is: , in, and These are the luminance reconstruction sample (e.g., after downsampling or applying GLM mode) and the chrominance reconstruction sample at position (i, j) in the template above, where ,and .

[0158] Similarly, the model error between candidate k and its left neighboring template is: , in, and These are the luminance reconstruction sample (e.g., after downsampling or GLM mode) and chrominance reconstruction sample at position (m, n) in the left template, where... as well as .

[0159] The model error of candidate k is then: .

[0160] After calculating the errors of all candidate models, a list of model errors can be obtained. Then, the candidate indices in the inherited candidate list can be reordered by sorting the model error list in ascending order.

[0161] In another embodiment, if candidate k is predicted using CCCM, and Defined as: , , in and These are the final filter coefficients obtained after inheriting candidate k. P and B are the nonlinear term and bias term, respectively.

[0162] In another embodiment, if the aforementioned adjacent template is unavailable, then Similarly, if the adjacent template on the left is unavailable, then If neither template is available, the candidate index reordering method using model error should not be applied.

[0163] In another embodiment, not all locations within the upper and left adjacent templates are used to calculate the model error. A subset of locations within the upper and left adjacent templates can be selected for model error calculation. For example, a first starting position and a first subsampling interval can be defined, depending on the width of the current block, to partially select locations within the upper adjacent template. Similarly, a second starting position and a second subsampling interval can be defined, depending on the height of the current block, to partially select locations within the left adjacent template.

[0164] In another example, or It can be a constant value (e.g., or It can be 1, 2, 3, 4, 5, or 6). Another example, or This can depend on the block size. If the current block size is greater than or equal to the threshold, or It equals the first value. Otherwise, or It equals the second value.

[0165] In another embodiment, different types of candidates are reordered separately before being added to the final candidate list. For each type of candidate, candidates are added to a predefined size. The candidates are placed in the main candidate list. The candidates in the main candidate list are then reordered. Then the top candidates in the main candidate list are... The candidate with the lowest cost is added to the final candidate list, where In another embodiment, candidates are categorized into different types based on their source, including, but not limited to, spatial proximity models, temporal proximity models, non-adjacent spatial proximity models, and historical candidates. In another embodiment, candidates are categorized into different types based on the cross-component model pattern. For example, the type could be CCLM, MMLM, CCCM, and CCCM multi-model. Another example is that the type could be GLM-non-active or GLM-active.

[0166] In another embodiment, after reordering the candidates based on template cost, the redundancy of the candidates can be further checked. If the template cost difference between a candidate and the previous candidate in the list is less than a certain threshold, the candidate is considered redundant. If a candidate is considered redundant, it can be removed from the list or moved to the end of the list.

[0167] Inheritance candidate index in the signaling list An on / off flag can be signaled to indicate whether the current block inherits cross-component model parameters from neighboring blocks. This flag can be signaled for each CU / CB, each PU, each TU / TB, each color component, or each chromaticity color component. High-level syntax can be signaled in the SPS, PPS (Image Parameter Set), PH (Image Header), or SH (Slice Header) to indicate whether the proposed method is permitted for the current sequence, image, or slice.

[0168] The signaling (also known as communication) specifies the maximum allowed number of candidates to indicate the maximum size of the candidate list for merging. This number can be signaled in CU / CB, PU, ​​TU / TB, color component, or chroma color component. High-level syntax can be used in SPS, PPS, PH, or SH to indicate whether the proposed method is suitable for the current sequence, image, or slice. The maximum allowed number of candidates for the proposed method (i.e., CCP merging mode) can be shared with the maximum allowed number of candidates for inter-frame merging mode.

[0169] If the current block inherits cross-component model parameters from neighboring blocks, then the signaling inheritable candidate index is used. This index can be signaled (e.g., using truncated unary codes, Exp-Golomb codes, or fixed-length code signaling) and shared between the current Cb and Cr blocks. As another example, this index can be signaled for each color component. For example, one inheritable candidate index can be signaled for the Cb component, and another inheritable candidate index can be signaled for the Cr component. As yet another example, the chroma intra-frame prediction syntax (e.g., IntraPredModeC[xCb][yCb]) can be used to store the inheritable candidate indices.

[0170] If the current block inherits cross-component model parameters from neighboring blocks, the current chroma intra-prediction mode (e.g., IntraPredModeC[xCb][yCb] as defined in the VVC standard) is temporarily set to the cross-component mode (e.g., CCLM_LA) during the bitstream syntax parsing phase. Subsequently, during the prediction or reconstruction phase, a candidate list is derived, and the inherited candidate model is determined by the inherited candidate index. After obtaining the inherited model, the codec information of the current block is updated according to the inherited candidate model. The codec information of the current block includes, but is not limited to, the prediction mode (e.g., CCLM_LA or MMLM_LA), the associated submode flag (e.g., CCCM mode flag), the prediction pattern (e.g., GLM pattern index), and the current model parameters. Then, the prediction for the current block is generated based on the updated codec information.

[0171] Inheriting multiple cross-component models The final prediction for the current block can be a combination of predictions from multiple cross-component models, or a fusion of predictions from the selected cross-component model and non-cross-component codec tools (e.g., intra-angle prediction mode, intra-plane / DC mode, or inter-frame prediction mode). In one embodiment, if the current candidate list size is N, k candidates can be selected from a total of N candidates (where k ≤ N). Then, k predictions are generated using the cross-component models of the selected k candidates and the corresponding brightness reconstruction samples. The final prediction for the current block is the result of combining these k predictions. For example, if two candidate predictions (denoted as...) and If the blocks are combined, the final prediction for the current block at position (x, y) is: ,in It is a weighting factor. Furthermore, the weighting factor... The template cost can be predefined or implicitly derived from the neighboring template cost (i.e., model error). For example, using the template costs defined in the section titled "Reordering Candidates in the List," the corresponding template costs for two candidates are... and ,but for In another embodiment, if two candidate models are combined, the selected model is chosen from the first two candidates in the list. In yet another embodiment, if i candidate models are combined, the selected model is chosen from the first i candidates in the list.

[0172] In another embodiment, if the current candidate list is N in size, k candidates can be selected from a total of N candidates (where k ≤ N). These k cross-component models can be combined into a final cross-component model by weighted averaging of their respective model parameters. For example, if a cross-component model has M parameters, the j-th parameter of the final cross-component model is the weighted average of the j-th parameters of the selected k candidates, where j is from 1 to M. The final prediction is then generated by applying the final cross-component model to the corresponding brightness reconstruction samples. For example, if the two candidate models are... and The final cross-component model is ,in It is a weighting factor that can be predefined or implicitly derived through the cost of neighboring templates. It is the x-th model parameter of the y-th candidate. For example, using the template cost defined in the section titled "Reordering Candidates in the List," the corresponding template costs for the two candidates are: and ,So yes For example, two candidate models may be selected, one from a spatially adjacent candidate and the other from a non-adjacent spatial candidate or a historical candidate. If a spatially adjacent candidate is unavailable, both candidate models may be selected from non-adjacent spatial candidates or historical candidates. In another embodiment, if two candidate models are merged, the selected model is selected from the first two candidates in the list. In yet another embodiment, if i candidate models are merged, the selected model is selected from the first i candidates in the list, where i is an integer greater than 1.

[0173] In another embodiment, the two cross-component models are merged into a final model by weighted averaging of the corresponding model parameters. One of the two cross-component models comes from the upper spatial neighbor candidate, and the other comes from the left spatial neighbor candidate. The upper spatial neighbor candidate is a neighbor candidate whose vertical position is less than or equal to the top boundary position of the current block. The left spatial neighbor candidate is a neighbor candidate whose horizontal position is less than or equal to the left boundary position of the current block. Weighting factor Determined based on the horizontal and vertical spatial location within the current block. For example, if two candidate predictions (represented as...) and The blocks are merged, and the final prediction for the current block at position (x, y) is... ,in In another embodiment, the top spatial neighbor candidate is the first candidate in the list whose vertical position is less than or equal to the top boundary position of the current block. The left spatial neighbor candidate is the first candidate in the list whose horizontal position is less than or equal to the left boundary position of the current block.

[0174] In another embodiment, cross-component model candidates can be combined with predictions from non-cross-component codec tools. For example, a cross-component model candidate can be selected from a candidate list, and its prediction is represented as... Another prediction can come from chroma DM, chroma DIMD, or intra-frame angle mode, and is represented as... The final prediction for the current block at position (x, y) is: ,in This is a weighting factor that can be derived implicitly based on a predefined value or through the neighbor template cost. For the same example, a predefined or signaled non-cross component codec tool can be used. The non-cross component codec tool is either chroma DM or chroma DIMD. For another example, a signaled non-cross component codec tool is used, but the index of the cross component model candidate is predefined or determined by the codec modes of neighboring blocks. For the same example, if at least one neighboring spatial block uses CCCM mode coding, the first candidate with CCCM model parameters is selected. If at least one neighboring spatial block uses GLM mode coding, the first candidate with GLM mode parameters is selected. Similarly, if at least one neighboring spatial block uses MMLM mode coding, the first candidate with MMLM parameters is selected.

[0175] In another embodiment, cross-component model candidates can be combined with the predictions of the current cross-component model. For example, a cross-component model candidate can be selected from the list, and its prediction is denoted as... Another prediction can come from the cross-component prediction pattern derived by the model based on the current neighboring reconstructed samples, and is denoted as... The final prediction for the current block at position (x, y) is: ,in It is a weighting factor that can be predefined or implicitly derived based on the cost of neighboring templates.

[0176] In another embodiment, multiple cross-component models can be combined into a single final cross-component model. For example, a model can be selected from one candidate and a second model from another candidate to form a multi-model pattern. The selected candidates can be CCLM / MMLM / GLM / CCCM codec candidates. The multi-model classification threshold can be an offset parameter between the two selected patterns (e.g., offset in CCLM / ...). or in CCCM or The average value of the light and chromaticity samples of the current block. In one embodiment, if two candidate models are combined, the selected model is the top two candidates in the list. In another embodiment, the classification threshold is set to the average value of the light and chromaticity samples of the current block.

[0177] Multi-hypothesis mixing of cross component model merging methods Let P be the current prediction for the current block; based on predefined weights, the additional hypothesis H is mixed with the current prediction P to form the final prediction P for the current block. final As shown in the following formula: , Prediction P can be generated based on any intra / inter-frame codec tool. Additional hypotheses are predictions of the current block generated based on the CC merge mode. The CC merge mode means that the cross-component model is inherited from spatial, historical, or temporal neighboring blocks / locations. A flag is sent to indicate whether additional hypotheses to be added with existing predictions are to be mixed. If the flag indicates that additional hypotheses should be added, the syntax related to how the additional hypotheses H and / or weights are communicated at the encoder end or resolved at the decoder end.

[0178] In one embodiment, multiple additional hypotheses can be added. The assumption is that N hypotheses H1, H2, ..., H... N Adding it to the existing prediction P0 results in the final prediction P. final It can be calculated as follows, where These are some predefined weights.

[0179] , , …, .

[0180] After signaling the syntax related to adding an additional hypothesis (including flags and syntax related to how to generate the additional hypothesis), another flag can be signaled to indicate whether to add another additional hypothesis. If the flag indicates to add another hypothesis, the syntax related to how to generate the additional hypothesis H and / or weights is signaled at the encoder, or parsed at the decoder.

[0181] In one embodiment, a pattern is selected from the cross-component merging candidate list to generate additional hypotheses. The inherited candidate index of the selected pattern can be indicated explicitly or implicitly. For example, the candidate index can be explicitly notified using the methods described in the section entitled "Messaging Inherited Candidate Indices in a List". The candidates in the list can be reordered using the methods described in the section entitled "Reordering Candidates in a List". For another example, only the first k candidates (k>=2) in the list can be selected, and the candidate index can be explicitly notified using the methods described in the section entitled "Messaging Inherited Candidate Indices in a List". The candidates in the list can be reordered using the methods described in the section entitled "Reordering Candidates in a List". For another example, after excluding candidates from the list whose previously added hypotheses have already been selected, the first candidate in the list can be implicitly selected. The candidate objects in the list can be reordered using the methods described in the section entitled "Reordering Candidates in a List".

[0182] In another embodiment, a pattern is selected from the list of cross-component merging candidates to generate additional hypotheses. The inherited candidate index of this pattern can be implicitly determined. After excluding candidates whose hypotheses have already been selected from the list, the candidate with the lowest template matching cost or boundary matching cost can be implicitly selected. The template matching cost and boundary matching cost are explained below: 1. Template matching cost settings: - Step 1: As shown in Figure 16, the reconstructed sample of the current block template is used as the gold data. The figure shows the current luma block 1610 and the current chroma block 1620, as well as the corresponding luma template 1612 and chroma template 1622.

[0183] - Step 2: For each inherited candidate pattern in the candidate list, generate the predicted sample value in the current chromaticity block chromaticity template based on the inherited model parameters and the luminance samples co-located in the luminance template.

[0184] - Step 3: For each inherited candidate pattern in the candidate list, calculate the distortion between the predicted sample within the chromaticity template and the golden data.

[0185] 2. Boundary Matching Cost Setting: The boundary matching cost of the inherited candidate pattern refers to a measure of discontinuity between the predicted sample within the current block generated by the inherited candidate pattern and the adjacent reconstructed samples within one or more neighboring blocks. Discontinuity measures include the top boundary matching cost and / or the left boundary matching cost. Top boundary matching refers to the comparison between the predicted sample at the top of the current block and the reconstructed sample of the block above the current block. Left boundary matching refers to the comparison between the predicted sample of the block to the left of the current block and the reconstructed sample of the block to the left of the current block.

[0186] In one embodiment, candidates in the cross-component merging candidate list can be categorized. The pattern selected to generate the additional hypothesis cannot belong to the same category as the pattern previously selected to generate the previously added hypothesis.

[0187] In one sub-implementation, patterns are categorized based on whether they are linear or convolutional models. For example, CCLM and MMLM belong to the same group, while CCCM belongs to another group.

[0188] In another sub-implementation, patterns are categorized based on whether they are single-model or multi-model patterns. For example, CCLM and CCCM belong to the same group, while MMLM and MMCCCM (multi-model CCCM) belong to another group.

[0189] In another sub-implementation, the patterns are classified based on the sample locations used to derive the cross-component model. For example, CCLM_LT and MMLM_LT belong to the same group, while CCLM_L and MMLM_L belong to another group.

[0190] In another sub-implementation, patterns are categorized based on the filter shape they use. For example, as shown in Figure 17, a convolutional model may have three filter shape choices (1710, 1720, and 1730). Convolutional models using filter shape 0 belong to the first group. Convolutional models using filter shape 1 and filter shape 2 are located in the second and third groups, respectively.

[0191] In another sub-implementation, the patterns are classified based on whether gradient brightness samples are used in the model. For example, GLM and CCLM are in different groups.

[0192] In another sub-implementation, the modes are classified according to the gradient filters used. For example, two-parameter GLMs and three-parameter GLMs using the first GLM filter are located in the first group. Two-parameter GLMs and three-parameter GLMs using the second, third, and fourth GLM filters are located in the second, third, and fourth groups, respectively. In another sub-implementation, the patterns are classified based on whether the co-located reconstructed luminance sample values ​​used in the model have been downsampled. In yet another sub-implementation, the patterns are classified based on the downsampling filter used to generate the downsampled reconstructed luminance samples.

[0193] In another sub-implementation, patterns are classified based on whether the model contains nonlinear terms.

[0194] In one embodiment, candidates in the cross-component merging candidate list can be categorized. The patterns can be categorized based on the classification criteria described above. The patterns selected for generating additional hypotheses can only come from certain categories. For example, the selected pattern must be a convolutional model pattern. As another example, the selected pattern must use gradient terms.

[0195] In one sub-implementation, patterns within the same category can be further subdivided. The selected pattern must not belong to the same subcategory as the pattern previously selected to generate a previously added hypothesis. Patterns can be subdivided based on the classification criteria described above. For example, if the CCCM category has subcategories 1) CCCM using filter shape 0, 2) CCCM using filter shape 1, and 3) CCCM using filter shape 2, and a hypothesis has already been generated by the CCCM pattern using filter shape 0, then the pattern selected to generate another hypothesis must come from group 2 or group 3 (using filter shape 1 or filter shape 2).

[0196] In one embodiment, if the existing predictor P is a cross-component pattern, the successor candidate selected for generating the additional hypothesis cannot be a spatial successor candidate.

[0197] In one embodiment, the weighting parameter α is selected from a set of weighting parameters. The signal indicates the selected α. For example, the set of weighting parameters may include α = 1 / 4 or -1 / 8. The signal indicates the index = 0 or 1.

[0198] In another embodiment, the weighting parameter α is selected from a set of weighting parameters. The weighting parameter α is implicitly determined based on the mode information of neighboring blocks. For example, if more neighboring blocks are in cross-component prediction mode, a larger α is selected. As another example, if more neighboring blocks are in inter-frame mode, and the existing prediction is in inter-frame mode, a smaller α is selected.

[0199] In another embodiment, the weighting parameter α is selected from a set of weighting parameters. The weighting parameter α is implicitly determined based on the type of the inherited pattern from which the hypothesis is generated. For example, if the inherited pattern is a CCCM pattern, a larger α is selected. As another example, if the inherited pattern is a CCLM pattern, a smaller α is selected.

[0200] In another embodiment, the weighting parameter α is selected from a set of weighting parameters. The weighting parameter α is implicitly determined based on template matching cost or boundary matching cost. A parameter associated with minimizing template matching cost or boundary matching cost is selected. The methods for calculating template matching cost and boundary matching cost have been described above. For example, if the set of weighting parameters contains... And assuming it is H, then for each parameter in the weighted parameter set, the predicted sample value within the chroma template of the current chroma block is calculated by... To generate it. Template cost is the distortion between the golden data (i.e., the reconstructed chroma sample on the current block template) and the predicted sample within the chroma template.

[0201] In another embodiment, the weighting parameter α is implicitly derived based on either the template matching cost (TIMD cost) or the boundary matching cost. The methods for calculating the template matching cost and boundary matching cost have been described above. The template matching cost and boundary matching cost can be determined at the decoder without message passing. For example, if the predicted cost is greater than the assumed cost, α uses a larger weight. If the predicted cost is less than the assumed cost, α uses a smaller weight. As another example, α is determined according to the following formula: .

[0202] In another embodiment, the weight parameter α can be implicitly derived using a method similar to that used to derive CCCM parameters, as described in the section titled "Convolutional Cross-Component Model (CCCM) - Single Model and Multiple Models". Let , where β is the offset value. The weight parameter α can be obtained by minimizing the MSE between the predicted chroma sample P_final and the reconstructed chroma sample in the reference region, as described in the "Convolutional Cross Component Model (CCCM) - Single Model and Multiple Model" section.

[0203] In one embodiment, the final prediction for the current block can be a mixture of N hypotheses using one or more predefined weights, where N is greater than or equal to 2. Each hypothesis is a prediction of the current block generated based on a mode. This mode can be any intra-frame or inter-frame mode. At least one hypothesis is generated based on a CC merging mode. The CC merging mode means that the cross-component model is inherited from spatial, historical, or temporal neighboring blocks / locations. A flag is sent at the block level to indicate whether the merging mode is applied. For example, this flag can be sent at the CU level and / or PU level and / or CTU level. As another example, the flag can be signaled at the CB level, PB level, CTB level, TU / TB level, any predefined region level, or a combination thereof.

[0204] In one embodiment, the final prediction for the current block can be a mixture of two hypotheses using predefined weights. The mode for generating the hypotheses can be any intra-frame or inter-frame mode. At least one hypothesis is generated based on a CC merging mode. A flag is sent at the block level to indicate whether the merging mode is applied.

[0205] In one embodiment, the final prediction for the current block can be a mixture of two hypotheses using predefined weights. One mode for generating hypotheses is a cross-component pattern derived from a neighbor template, and the other is a CC merging pattern. The CC merging pattern is selected from a CC merging candidate list. In one sub-implementation, the selected pattern cannot be a spatially inherited candidate pattern.

[0206] In one embodiment, the final prediction for the current block can be a mixture of two hypotheses. Both hypotheses, H1 and H2, are generated based on the CC merge mode. A flag is sent at the block level to indicate whether the mixed mode is applied.

[0207] In one embodiment, when the flag indicates that Blending Mode is disabled, the original syntax for signaling / parsing the intra-prediction mode of the current block is followed. When the flag indicates that Blending Mode is enabled, the following syntax needs to be sent at the encoder, indicating the mode and / or weights for generating H1 and H2, and parsed and decided at the decoder.

[0208] In one embodiment, both patterns for generating H1 and H2 are selected from a cross-component merging candidate list. The inherited candidate index of the selected pattern can be explicitly or implicitly indicated. For example, the candidates in the list can be reordered according to the method described in the section entitled "Reordering Candidates in the List". The candidate index can be explicitly signaled according to the method described in the section entitled "Signaling Inherited Candidate Indices in the List". As another example, only the first k candidates (k>= 3) can be selected from the list, and the candidate index can be explicitly signaled according to the method described in the section entitled "Signaling Inherited Candidate Indices in the List". As another example, the first two candidates can be implicitly selected. As yet another example, the index can be explicitly signaled to indicate the candidate index of the first pattern, and the candidate index of the second pattern is equal to the signaled index + k, where k can be 1, 2, 3, 4, or 5.

[0209] In one embodiment, both patterns that generate H1 and H2 are selected from a cross-component merging candidate list. The inheritance candidate index of the selected pattern is implicitly determined. The two candidates with the lowest template matching cost or boundary matching cost in the list can be implicitly selected.

[0210] In one embodiment, both patterns that generate H1 and H2 are selected from a cross-component merging candidate list. The candidates in the cross-component merging candidate list are categorized. The patterns can be categorized according to the classification criteria described above. The two selected patterns must come from different categories.

[0211] In another embodiment, the two patterns come from certain categories. For example, both patterns are CCCM patterns. In a sub-implementation, patterns within the same category can be further subdivided. The two selected patterns must not belong to the same subcategory. Patterns can be subclassified according to the classification criteria described above.

[0212] In one embodiment, weights for mixing H1 and H2 are selected from a set of weighted parameters. An index can be explicitly communicated to indicate the selected weights. For example, the weight set ( The weights can also be (3, 5), (5, 3), (-2, 10), (10, -2), (4, 4), or any subset thereof. In another embodiment, the weights can be implicitly determined based on pattern information from neighboring blocks. For example, if H1 is a linear model (e.g., CCLM and MMLM) and H2 is a convolutional model (e.g., CCCM), and more neighboring blocks are convolutional model patterns, then H2 has a larger mixed weight. In another embodiment, the weights can be implicitly determined based on the type of inheritance pattern that generates the hypothesis. For example, if the inheritance pattern is a CCCM pattern, then a larger weight is selected. As another example, if the inheritance pattern is a CCLM pattern, then a smaller weight is selected. In another embodiment, the weights can be implicitly determined based on template matching cost or boundary matching cost. The weights associated with the minimum template matching cost or boundary matching cost are selected. The methods for calculating template matching cost and boundary matching cost have been described in the preceding paragraphs. For example, if template matching cost is used, then for each weight in the weighted parameter set, by calculating... This is used to generate predicted sample values ​​within the chroma template of the current chroma block. Template cost is the distortion between the golden data (i.e., the chroma samples reconstructed on the current block template) and the predicted samples within the chroma template.

[0213] In another embodiment, the weighting parameter ( The α is implicitly derived based on template matching cost (TIMD cost) or boundary matching cost. The calculation methods for template matching cost and boundary matching cost have been described previously. Template matching cost and boundary matching cost can be determined at the decoder without communication. For example, if the pattern generating H1 has a high TIMD cost, H1 uses a smaller weight during the mixing process. If the pattern generating H2 has a high TIMD cost, H2 uses a smaller weight during the mixing process. If the prediction cost is less than the assumed cost, a smaller weight is chosen for α. For another example... , .

[0214] In another embodiment, the weighting parameter ( It can be implicitly derived using a method similar to that used to derive CCCM parameters, as described in the "Convolutional Cross Component Model (CCCM)" section. Let... , where β is the offset value. Weight parameters , And β can be minimized by the predicted chromaticity samples The MSE between the reconstructed chromaticity samples in the reference region is derived, as described in the "Convolutional Cross Component Model (CCCM)" section.

[0215] In one embodiment, when combining two cross-component models, a final cross-component model can be generated by combining / weighting the parameters of the two models. If a term appears only in one model, the parameter associated with that term in the other model will be treated as zero. For example, if a 7-tap CCCM model is to be combined with a 3-parameter GLM model, the 7-tap CCCM model is: The 3-parameter GLM model is As described in the sections on "Convolutional Cross-Component Model (CCCM)" and "Gradient Linear Model (GLM)," respectively. Let the combined weights be ( Since the C term in the 7-tap CCCM model is equivalent to the L term in the 3-parameter GLM model, and the 7-tap CCCM model... And the 3-parameter GLM model All of these are bias terms, therefore the final combined model is as follows, which is essentially an 8-tap model: .

[0216] In another embodiment, when generating the final cross-component model from two cross-component models, some parameters are obtained by combining / weighting the parameters of the two cross-component models, while the remaining parameters can be re-derived. For example, if we want to combine a 7-tap CCCM model with a 3-parameter GLM model to generate an 8-tap model, let the 7-tap CCCM model be... The 3-parameter GLM model is As described in the sections on "Convolutional Cross-Component Model (CCCM)" and "Gradient Linear Model (GLM)," the final model combination is as follows: , in{ The calculation method is as described above, and (The deviation term) is based on { And the re-derived values ​​of neighboring brightness and chromaticity reconstructed samples. For example, this can be achieved by minimizing the predicted chromaticity samples. The mean square error (MSE) between the reconstructed chromaticity samples in the reference region is used to derive the... As described in the section “Convolutional Cross Component Model (CCCM)”.

[0217] The aforementioned multi-hypothesis mixing method for cross-component model merging can be implemented at the encoder or decoder end. For example, any proposed multi-hypothesis mixing method for cross-component prediction can be implemented in the decoder's intra / inter-frame codec module (e.g., intra-prediction 150 / MC 152 in Figure 1B) or in the encoder's intra / inter-frame codec module (e.g., intra-prediction 110 / inter-prediction 112 in Figure 1A). Any proposed multi-hypothesis mixing method can also be implemented as a circuit coupled to the decoder or encoder's intra / inter-frame codec module. However, the decoder or encoder can also use additional processing units to implement the required cross-component prediction processing. Although the intra-prediction units (e.g., units 110 / 112 in Figure 1A and units 150 / 152 in Figure 1B) are shown as separate processing units, they may correspond to executable software or firmware code stored on a medium (e.g., hard disk or flash memory) for a CPU (Central Processing Unit) or a programmable device (e.g., a DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array)).

[0218] Figure 18A flowchart illustrating an exemplary video encoding / decoding system using a multi-hypothesis mixing method for cross-component model merging according to an embodiment of the present invention is shown. The steps shown in the flowchart can be implemented as program code executable on one or more processors (e.g., one or more CPUs) on the encoder side. The steps shown in the flowchart can also be implemented in hardware, such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, in step 1810, input data associated with a current block, comprising a first color block and a second color block, is received, wherein the input data includes pixel data to be encoded on the encoder side or encoded data associated with the current block to be decoded on the decoder side. In step 1820, a basic predictor for the second color block is determined, wherein the basic predictor corresponds to an intra-frame predictor or an inter-frame predictor. In step 1830, one or more hypotheses for the second color block are determined, wherein each of the one or more hypotheses corresponds to a cross-component prediction candidate. In step 1840, a final predictor is derived by mixing the basic predictor with the one or more hypotheses. In step 1850, the second color block is encoded or decoded using prediction data containing the final predictor.

[0219] The flowchart shown is intended to illustrate an example of video encoding and decoding according to the present invention. Those skilled in the art can modify each step, rearrange the steps, split the steps, or combine the steps to implement the invention without departing from its spirit. Specific syntax and semantics are used in this disclosure to illustrate examples of implementing the invention. Those skilled in the art can implement the invention by using equivalent syntactic and semantic substitutions without departing from its spirit.

[0220] The foregoing description is intended to enable those skilled in the art to practice the invention within the context of specific applications and requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the invention is not intended to be limited to the specific embodiments shown and described, but should be given the broadest scope in accordance with the principles and novel features disclosed herein. Various specific details have been shown in the foregoing detailed description to provide a thorough understanding of the invention. However, those skilled in the art will understand that the invention can be practiced.

[0221] The embodiments of the invention described above can be implemented in various hardware, software code, or combinations thereof. For example, one embodiment of the invention may be program code integrated into a video compression chip or into video compression software to perform the processes described herein. Another embodiment of the invention may be program code to be executed on a digital signal processor (DSP) to perform the processes described herein. The invention may also relate to multiple functions performed by a computer processor, digital signal processor, microprocessor, or field-programmable gate array (FPGA). These processors can be configured to execute machine-readable software code or firmware code to perform specific methods embodied in the invention. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, styles, and languages ​​of the software code, as well as other configuration codes to perform methods consistent with the tasks of the invention, do not depart from the spirit and scope of the invention.

[0222] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The examples described are for illustrative purposes only and not for limitation. Therefore, the scope of the invention is indicated by the appended claims rather than the foregoing description. All variations within the meaning and equivalence of the claims should be included within their scope.

Claims

1. An encoding / decoding method for encoding and decoding a color image using an encoding / decoding tool that includes one or more cross-component model correlation modes, the method comprising: Receive input data related to a current block including a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or encoded data related to the current block to be decoded at the decoder end; Determine the base predictor of the second color patch, wherein the base predictor corresponds to an intra-frame predictor or an inter-frame predictor; Determine one or more hypotheses for the second color patch, wherein each of the one or more hypotheses corresponds to a cross component prediction candidate; The final predictor is derived by mixing the base predictor and one or more of the hypotheses; and The second color block is encoded or decoded using the prediction data including the final predictor.

2. The method of claim 1, wherein the final predictor is derived by using a predefined weighted mixture of the base predictor and the one or more hypotheses.

3. The method of claim 1, wherein one or more flags are communicated or parsed to indicate whether the one or more assumptions are added to derive the final predictor.

4. The method of claim 3, wherein whether the target flag of the one or more flags is signaled or parsed depends on the flag of the one or more flags previously signaled or parsed.

5. The method of claim 3, wherein when the one or more flags indicate that the one or more hypotheses are added to derive the final predictor, one or more syntaxes related to the one or more hypotheses and / or related to a predefined weighting for mixing the base predictor and the one or more hypotheses are signaled or parsed.

6. The method of claim 1, wherein the cross component prediction candidate is selected from the cross component merging candidate list.

7. The method of claim 6, wherein the candidate index associated with the cross component prediction candidate selected from the cross component merging candidate list is explicitly signaled or parsed, or implicitly deduced.

8. The method of claim 6, wherein the first candidate in the cross component merging candidate list is selected as the cross component prediction candidate.

9. The method of claim 6, wherein the candidate with the lowest template matching cost or boundary matching cost in the cross component merging candidate list is selected as the cross component prediction candidate.

10. The method of claim 6, wherein the member candidates in the cross component merging candidate list are classified into one or more categories, and the selection of the one or more hypotheses depends on the one or more categories.

11. The method of claim 10, wherein, The one or more prediction models selected to generate the one or more hypotheses belong to a different category than one or more previously selected models.

12. The method of claim 10, wherein the one or more categories include a linear model category and a convolutional model category.

13. The method of claim 10, wherein the one or more categories include single-model categories and multi-model categories.

14. The method of claim 10, wherein only one or more member candidates belonging to one or more specific categories in the cross component merging candidate list are allowed to generate the one or more hypotheses.

15. The method of claim 14, wherein the one or more specific categories correspond to cross-component convolution model patterns.

16. The method of claim 14, wherein if the base predictor corresponds to a cross component pattern, the one or more specific category exclusion spatial inheritance candidate patterns are excluded.

17. The method of claim 1, wherein the final predictor is derived by using a weighted mixture of the base predictor and the one or more hypotheses, including weighting parameters.

18. The method of claim 17, wherein the weighting parameter is implicitly determined based on pattern information from one or more neighboring blocks, cross-component prediction pattern for generating the one or more hypotheses, template matching cost, or boundary matching cost.

19. The method of claim 17, wherein the weighting parameter is selected from a set of weighting parameters.

20. The method of claim 17, wherein the weighting parameter is indicated by the syntax being communicated or parsed.

21. The method of claim 17, wherein the weighting parameters are implicitly determined.

22. The method of claim 17, wherein the weighting parameters are derived by minimizing the MSE between the predicted second color sample of the final predictor and the reconstructed second color sample in one or more reference regions.

23. The method of claim 1, wherein the base predictor is derived based on cross-component prediction hypotheses, and each of the base predictor and the one or more hypotheses is based on an encoding / decoding mode.

24. The method of claim 23, wherein the base predictor is derived based on a cross component prediction candidate selected from the cross component merging candidate list.

25. A video encoding / decoding apparatus, the apparatus comprising one or more electronic devices or processors configured to: Receive input data related to a current block including a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or encoded data related to the current block to be decoded at the decoder end; Determine the base predictor of the second color patch, wherein the base predictor corresponds to an intra-frame predictor or an inter-frame predictor; Determine one or more hypotheses for the second color patch, wherein each of the one or more hypotheses corresponds to a cross component prediction candidate; The final predictor is derived by mixing the base predictor and one or more of the hypotheses; and The second color block is encoded or decoded using the prediction data including the final predictor.