Temporal candidate method and apparatus for cross-component model merge mode in video coding system

By deriving motion vectors from neighboring blocks and inheriting model parameters from previously encoded and decoded images, the encoding and decoding performance of cross-component prediction is improved, addressing the inefficiency problem in existing technologies and enhancing the efficiency and quality of video encoding and decoding.

CN120982087APending Publication Date: 2025-11-18MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480026986.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-18
Filing Date
2024-04-18
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing video coding and decoding technologies suffer from inefficiency in cross-component prediction, especially when dealing with multi-functional video coding and decoding standards such as VVC, where it is difficult to effectively utilize information from neighboring blocks to determine time candidates.

Method used

By deriving motion vectors from neighboring blocks to determine temporal candidates, cross-component prediction data is generated using cross-component model parameters, and model parameters are inherited from previously encoded and decoded images, thus improving the encoding and decoding performance of cross-component prediction.

Benefits of technology

It improves the efficiency and quality of video encoding and decoding by effectively utilizing information from neighboring blocks and previous images, thereby enhancing the accuracy of cross-component prediction and encoding/decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982087A_ABST
    Figure CN120982087A_ABST
Patent Text Reader

Abstract

A method and apparatus derives motion vectors from neighboring blocks to determine temporal candidates and inherits a cross-component model from two or more previously coded pictures for cross-component prediction. According to one method, a target motion vector is determined based on neighboring blocks, and temporal candidates are determined based on the target motion vector to be included in a candidate list. According to another method, a plurality of co-located pictures are determined for a current block and a temporal candidate is determined based on two or more co-located pictures.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related citations This invention is a non-provisional application and claims priority with U.S. Provisional Patent Application No. 63 / 496,704, filed April 18, 2023. This U.S. Provisional Patent Application is incorporated herein by reference in its entirety. Technical Field

[0002] This invention relates to video encoding and decoding systems. In particular, it relates to deriving motion vectors from neighboring blocks to determine temporal candidates or inheriting cross-component models from two or more previously encoded and decoded images. Background Technology

[0003] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Group (JVET), comprised of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). This standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology – Coding representation of immersive media – Part 3: Versatile Video Coding, published in February 2021. VVC was developed based on its predecessor, High Efficiency Video Coding (HEVC), by adding more coding and decoding tools to improve coding and decoding efficiency and handle various types of video sources, including three-dimensional (3D) video signals.

[0004] Figure 1AAn exemplary adaptive intra / inter-frame video coding system incorporating loop processing is described. For intra-frame prediction 110, prediction data is derived from previously encoded video data in the current frame. For inter-frame prediction 112, motion estimation (ME) is performed at the encoder, and motion compensation (MC) is performed based on the result of ME to provide prediction data from other frames and motion data. Switch 114 selects either intra-frame prediction 110 or inter-frame prediction 112 and provides the selected prediction data to adder 116 to form a prediction error, also known as a residual. The prediction error is then processed by transform (T) 118, followed by quantization (Q) 120. The residuals from the transform and quantization are then encoded by entropy encoder 122 and included in the video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packaged with side information, such as motion and encoding / decoding modes associated with intra-frame and inter-frame prediction, and parameters associated with loop filters applied to the base image regions. Side information related to intra-frame prediction 110, inter-frame prediction 112, and loop filter 130, such as Figure 1A The data is provided to the entropy encoder 122. When using inter-frame prediction mode, a reference image must also be reconstructed at the encoder. Therefore, the residuals from the transform and quantization are processed by inverse quantization (IQ) 124 and inverse transformation (IT) 126 to recover the residuals. Then, the residuals are added to the prediction data 136 at reconstruction (REC) 128 to reconstruct the video data. The reconstructed video data can be stored in the reference image buffer 134 and used for prediction of other frames.

[0005] like Figure 1A As shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subject to various degradations due to these processes. Therefore, a loop filter 130 is typically applied to the reconstructed video data before it is stored in the reference picture buffer 134 to improve video quality. For example, a deblocking filter (DF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF) can be used. Loop filter information may need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. Figure 1AIn the process, the loop filter 130 is applied to the reconstructed video, and then the reconstructed samples are stored in the reference image buffer 134. Figure 1A The system described herein is intended to illustrate an exemplary architecture of a typical video encoder. It may correspond to a High Efficiency Video Codec (HEVC) system, VP8, VP9, ​​H.264, or VVC.

[0006] like Figure 1B As shown, the decoder can use functional modules similar to or partially identical to the encoder, except for transform 118 and quantization 120, since the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses entropy decoder 140 to decode the video bitstream into quantized transform coefficients and the required encoding / decoding information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). Intra-frame prediction 150 at the decoder end does not require mode search. Instead, the decoder only needs to generate intra-frame predictions based on the intra-frame prediction information received from entropy decoder 140. Furthermore, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from entropy decoder 140, without motion estimation.

[0007] According to VVC, the input image is divided into non-overlapping block regions called codec tree units (CTUs), similar to HEVC. Each CTU can be further divided into one or more smaller codec units (CUs). The resulting CU partitions can be square or rectangular. VVC also divides CTUs into prediction units (PUs), which serve as units for applying prediction processes such as inter-frame prediction and intra-frame prediction.

[0008] The VVC standard includes various new encoding and decoding tools to further improve encoding and decoding efficiency compared to the HEVC standard. Some of the new tools related to this invention are described below.

[0009] Cross-Component Linear Model (CCLM) Prediction To reduce cross-component redundancy, VVC uses a cross-component linear model (CCLM) prediction mode. This mode uses a linear model to predict chromaticity samples based on reconstructed luminance samples from the same CU, as shown below: (1) in This represents the predicted chromaticity samples in the CU. This represents the downsampled reconstructed luminance samples from the same CU.

[0010] CCLM parameters ( and The chroma block size is derived based on at most four neighboring chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block size is W×H, then W' and H' are set as follows: When W' = W, H' = H is applied in LM_LA mode; W' = W + H when applying LM_A mode; H' = H + W when applying LM_L mode.

[0011] The upper nearest neighbor is denoted as S[0, -1]…S[W' - 1, -1] and the left nearest neighbor is denoted as S[-1, 0]…S[-1, H' - 1]. Then, four samples are selected as... When applying the LM_LA mode and both the upper and left neighboring samples are available, select S[W' / 4, -1], S[3 W' / 4, -1], S[ -1, H' / 4], S[ -1, 3 H' / 4]; When applying LM_A mode or when only the upper neighboring sample is available, select S[W' / 8, -1], S[3]. W' / 8, -1], S[ 5 W' / 8, -1], S[ 7 W' / 8, -1]; When applying LM_L mode or when only the left neighbor sample is available, select S[-1, H' / 8], S[-1, 3]. H' / 8], S[-1, 5 H' / 8], S[ -1, 7 H' / 8].

[0012] Four neighboring brightness samples at the selected location were downsampled and compared four times to find the two larger values: x 0 A and x 1 A And two smaller values: x 0 B and x 1 B Their corresponding chromaticity sample values ​​are represented as y 0 A , y1 A , y 0 B and y 1 B .Then x A , x B , y A and y B Exported as: x A = (x 0 A + x 1 A + 1)>>1; x B = (x 0 B + x 1 B + 1)>>1; y A = (y 0 A + y 1 A + 1)>>1; y B = (y 0 B + y 1 B + 1)>>1. (2) Finally, the linear model parameters are obtained according to the following procedure. and .

[0013] (3) (4) Figure 2 This shows examples of the positions of the left and top samples involved in the LM_LA pattern, as well as the current block sample. Figure 2 Showing Chroma block 210, corresponding The relative sample positions of luminance block 220 and its neighboring samples (displayed as solid circles).

[0014] In addition to the templates mentioned above and the left-hand template being used together to calculate the coefficients of the linear model, they can also be used alternately in two other LM modes, called LM_A and LM_L modes.

[0015] In LM_A mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H+W) samples.

[0016] In LM_LA mode, the linear model coefficients are calculated using the left and top templates.

[0017] In this disclosure, the terms {LM_LA, LM_L, LM_A} and {CCLM_LT, CCLM_L, CCLM_T} are used interchangeably.

[0018] To match the chroma sample positions in a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filters is specified by the Sequence Parameter Set (SPS) level flags. These two downsampling filters are as follows, corresponding to "Type-0" and "Type-2" content, respectively.

[0019] (5) (6) Note that when the upper reference line is at the CTU boundary, only one luminance line (the universal line buffer in intra-frame prediction) is used to create the downsampled luminance sample.

[0020] This parameter calculation is performed as part of the decoding process, not just the encoder's search operation. Therefore, there is no syntax for passing the α and β values ​​to the decoder.

[0021] Multi-model CCLM (MMLM) In JEM (J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, and J. Boyce, Algorithm description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Video Exploration Team (JVET), July 2017), a multi-model CCLM mode (MMLM) was proposed to predict chromaticity samples from luminance samples across the entire CU using two models. In MMLM, neighboring luminance and chromaticity samples of the current block are classified into two groups, each used as a training set to derive a linear model (i.e., derive specific α and β for a specific group). Furthermore, samples from the current luminance block are also classified based on the same rules to classify neighboring luminance samples. Three MMLM model modes (MMLM_LA, MMLM_T, and MMLM_L) are allowed to select neighboring samples from the left and top, top only, and left only, respectively.

[0022] In this disclosure, the terms MMLM_LA and MMLM_LT are used interchangeably.

[0023] Figure 3 This example shows how neighboring samples are classified into two groups. The threshold is calculated as the average of the neighboring reconstructed brightness samples. Neighboring samples with a Rec′L [x,y] <= threshold are classified into group 1; while neighboring samples with a Rec′L [x,y] > threshold are classified into group 2.

[0024] (7) Convolutional Cross-Component Models (CCCM) - Single Model and Multiple Models In CCCM, a convolutional model is applied to improve chromaticity prediction performance. The convolutional model has a 7-tap filter, consisting of a 5-tap spatial component with a sign shape, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter includes the center (C) luminance sample co-located with the chromaticity sample to be predicted and its top / north (N), bottom / south (S), left / west (W), and right / east (E) neighbors, such as... Figure 4 As shown.

[0025] The non-linear term (denoted as P) is represented as the square of the center brightness sample C, scaled to the range of sample values ​​for the content: P = (C C + midVal )>>bitDepth.

[0026] For example, for 10-bit content, the non-linear term is calculated as follows: P = (C C + 512)>>10.

[0027] The offset term (denoted as B) represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).

[0028] The output of the filter is calculated as the filter coefficients c. i Convolve the input values ​​and crop them to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B Filter coefficients c i It is calculated by minimizing the MSE between the predicted and reconstructed chromaticity samples in the reference region. Figure 5 This illustrates an example of a reference region consisting of six rows of chroma samples above and to the left of the PU. The reference region extends to the right by one PU width and downwards by one PU height. The region is adjusted to include only available samples. Region expansion (represented as "filling") needs to be supported. Figure 5 The plus sign shape space filter in the "side sample" is used to fill in the unavailable areas.

[0029] MSE minimization is performed by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and the chrominance output. The autocorrelation matrix is ​​decomposed using LDL, and the final filter coefficients are calculated using back-substitution. This process roughly follows the calculation of ALF filter coefficients in ECM, but LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.

[0030] Furthermore, similar to CCLM, CCCM offers options for single-model or multi-model variants. The multi-model variant uses two models: one for samples above the average luminance reference value, and the other for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode is selectable for PUs with at least 128 reference samples available.

[0031] Gradient Linear Model (GLM) For the YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chromaticity samples from the gradient of luminance samples. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.

[0032] Compared to CCLM, GLM uses luminance sample gradients to derive a linear model, rather than downsampled luminance values. Specifically, when applying GLM, the input to the CCLM process is the downsampled luminance sample... The gradient of the brightness sample Replacement. Other parts of CCLM (such as parameter derivation and linear transformation of predicted samples) remain unchanged.

[0033] .

[0034] In a three-parameter GLM, different parameters can be used to predict chromaticity samples based on the luminance sample gradient and the downsampled luminance value. The model parameters of a three-parameter GLM are derived from neighboring samples in 6 rows and columns through LDL decomposition based on the MSE minimization method, as used in CCCM.

[0035] .

[0036] In terms of signal transmission, when CCLM mode is enabled at the current CU, a flag is transmitted to indicate whether GLM is enabled for the Cb and Cr components; if GLM is enabled, another flag is transmitted to indicate which GLM mode is selected, and a syntax element is further transmitted to select the four gradient filters used for gradient calculation. Figure 6 One of them (610-640).

[0037] Spatial candidate derivation In VVC, the derivation of spatial merge candidates is the same as in HEVC, except that the positions of the first two merge candidates are swapped. Figure 7 At the positions depicted, a maximum of four merge candidates (B0, A0, B1, and A1) for the current CU 710 are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is only considered if one or more neighboring CUs of positions B0, A0, B1, and A1 are unavailable (e.g., belonging to another slice or tile) or if intra-frame coding is used. After adding a candidate for position A1, the addition of the remaining candidates undergoes a redundancy check to ensure that candidates with the same motion information are excluded from the list, thereby improving encoding and decoding efficiency.

[0038] Time Candidate Derivation In this step, only one candidate is added to the list. Specifically, when deriving this time-merging candidate for the current CU 810, it is based on, as... Figure 8 The scaling motion vector is derived from the co-located CU 820 of the shown co-located reference image. The list of reference images and reference indices used to derive the co-located CU are explicitly specified in the slice header. The scaling motion vector 830 of the time-merging candidate is shown below. Figure 8As shown by the dashed lines, the distances tb and td are obtained by scaling the motion vector 840 of the co-located CU using the POC (Picture Order Count) distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located image and the reference image. The reference image index of the temporal merging candidate is set to zero.

[0039] The position of the time candidate is Figure 9 Choose between candidate C0 and C1. If the CU at position C0 is unavailable, is intra-coded, or is outside the current CTU line, then position C1 is used. Otherwise, position C0 is used when merging candidates at derivation time.

[0040] Non-neighbor space candidates For example, in JVET-L0399 (Yu Han et al., “CE4.4.6: Improvement on Merge / Skip mode”, Joint Video Exploration Team (JVET) of ITU-TSG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 12th Meeting: Macau, China, October 3-12, 2018, Document: JVET-L0399), non-proximity spatial merge candidates are inserted into the regular merge candidate list after the TMVP (Time MVP). The pattern of spatial merge candidates is as follows: Figure 10 As shown. The distance between non-nearest neighbor space candidates and the current decode block is based on the width and height of the current decode block. Line buffer constraints do not apply.

[0041] In this invention, the method and apparatus derive cross-component prediction models for intra- and non-intra-coded blocks based on two or more co-located images. Furthermore, the method and apparatus derive motion vectors from neighboring blocks to determine temporal candidates. Summary of the Invention

[0042] This invention discloses a method and apparatus for deriving motion vectors from neighboring blocks to determine temporal candidates. According to the method, input data associated with a current block is received, the current block including a first color block and a second color block, wherein the input data includes pixel data to be encoded at an encoder or data associated with the current block to be decoded at a decoder, and wherein the current block is located in an intra-frame or non-intra-frame slice or image. A target motion vector is determined based on one or more neighboring blocks of the current block. A temporal candidate is determined based on the target motion vector, wherein a target cross-component model is generated for the second color block using one or more cross-component model parameters associated with the temporal candidate. A candidate list containing the temporal candidate is determined. The second color block is encoded or decoded using the candidate list, wherein when the temporal candidate is selected, cross-component prediction data is generated for the second color block according to the target cross-component model.

[0043] In one embodiment, the one or more neighboring blocks are located at one or more predefined locations. In one embodiment, the selected target motion vector corresponds to an L0 or L1 motion vector. In one embodiment, if the target neighboring block of the one or more neighboring blocks is not an intra-frame block, then the target neighboring block of the one or more neighboring blocks will not be selected for the current block to determine the time candidate.

[0044] In one embodiment, one or more neighboring blocks of the current block correspond to a predefined list of locations, and a target neighboring block is selected from the predefined list of locations according to a predefined check order. In one embodiment, the target neighboring block corresponds to a first location in the predefined list of locations, the first location being an intra-frame block. In another embodiment, the target neighboring block corresponds to a first location in the predefined list of locations, wherein the first location has a reference image as a co-located image of the current block.

[0045] In one embodiment, the target cross-component model is inherited from a previously encoded image, and the previously encoded image is selected from one or more reference images in a reference list. In one embodiment, the target reference image selected from the one or more reference images in the reference list is signaled in the image or slice header.

[0046] In one embodiment, a target reference image is implicitly selected from a plurality of reference images in one or more reference lists according to a set of predefined rules. In one embodiment, the set of predefined rules includes selecting a candidate reference image with the smallest picture order count (POC) difference or the smallest quantization parameter (QP) difference, or a candidate reference image with the smallest QP among the plurality of reference images, as the target reference image. In another embodiment, the set of predefined rules includes selecting a candidate reference image corresponding to the most recently encoded I-image as the target reference image.

[0047] The present invention also discloses a method and apparatus for inheriting a cross-component model from two or more previously encoded / decoded images. According to the method, input data associated with a current block is received, the current block comprising a first color block and a second color block, wherein the input data includes pixel data to be encoded at an encoder end or data to be decoded at a decoder end associated with the current block, and wherein the current block is located in an intra-frame non-intra-frame slice or image. A plurality of co-location images are determined for the current block, wherein the plurality of co-location images correspond to previously encoded / decoded images. Temporal candidates are determined based on two or more co-location images, wherein a target cross-component model is derived by inheriting one or more cross-component model parameters from the two or more co-location images. A candidate list containing the temporal candidates is determined. The second color block is encoded or decoded using the candidate list, wherein when the temporal candidate is selected, cross-component prediction data is generated for the second color block according to the target cross-component model.

[0048] In one embodiment, the total number of co-located images is signaled or parsed in the bitstream. In another embodiment, the total number of co-located images is signaled or parsed in the PH (image header), SH (slice header), PPS (image parameter set), or SPS (sequence parameter set) of the bitstream.

[0049] In one embodiment, multiple co-located pictures correspond to a picture set consisting of N previously encoded pictures. In one embodiment, a target index is signaled or parsed from the bitstream to indicate one of two or more co-located pictures selected from the picture set. In one embodiment, the index value of a member picture is smaller if the difference in picture order count (POC) between a member picture in the picture set and the current picture is small. In one embodiment, the index value of a member picture is smaller if the difference in picture-level quantization parameter (QP) between a member picture in the picture set and the current picture is small. In another embodiment, the index value of a member picture is smaller if the picture-level QP of a member picture in the picture set is small. In yet another embodiment, the index value of a member picture is smaller if the picture-level QP of a member picture in the picture set is large.

[0050] In one embodiment, two or more co-located images are selected from an image set according to a set of predefined rules. In one embodiment, the set of predefined rules corresponds to selecting one or more candidate images from the image set that have a smaller QP, a larger QP, a smaller QP difference from the current image, a smaller POC difference from the current image, and / or a combination thereof. Attached Figure Description

[0051] Figure 1A An exemplary adaptive inter-frame / intra-frame video encoding and decoding system incorporating loop processing is described.

[0052] Figure 1B Explanation Figure 1A The corresponding decoder of the encoder.

[0053] Figure 2 This shows an example of the positions of the left and top samples involved in the LM_LA mode, as well as the current block sample.

[0054] Figure 3 An example is shown where neighboring samples are classified into two groups.

[0055] Figure 4 An example illustrating the spatial portion of a convolutional filter is provided.

[0056] Figure 5 An example of a padded reference region is shown for deriving filter coefficients.

[0057] Figure 6 This section describes four gradient modes used in the Gradient Linear Model (GLM).

[0058] Figure 7 This describes the neighboring blocks used to derive VVC space merge candidates.

[0059] Figure 8 An example of temporal candidate export is illustrated, in which scaled motion vectors are derived based on the Picture Order Count (POC) distance.

[0060] Figure 9 This indicates the position of the time candidate selected between candidate C0 and C1.

[0061] Figure 10 An exemplary pattern for non-adjacent spatial merging candidates is illustrated.

[0062] Figure 11 An example illustrating the inheritance time proximity model parameters is provided.

[0063] Figure 12 An example of a time candidate position according to an embodiment of the present invention is illustrated.

[0064] Figure 13A -B specifies two search patterns used to inherit the proximity model of non-adjacent spaces.

[0065] Figure 14 A flowchart illustrating an exemplary video encoding / decoding system is provided, which derives motion vectors from neighboring blocks to determine time candidates, according to an embodiment of the present invention.

[0066] Figure 15A flowchart illustrating an exemplary video encoding / decoding system that inherits a cross-component model from two or more previously encoded / decoded images, according to an embodiment of the invention. Detailed Implementation

[0067] It will be readily understood that the components of the present invention, as described and illustrated in the figures, can be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of embodiments of the systems and methods of the present invention, as illustrated in the figures, is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the invention. References to “one embodiment,” “an embodiment,” or similar language throughout this specification mean that a particular feature, structure, or characteristic described in the description associated with that embodiment may be included in at least one embodiment of the invention. Therefore, the phrase “in one embodiment” or “in one embodiment” appearing in various places in this specification does not necessarily refer to the same embodiment.

[0068] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize that the invention can be practiced without including one or more specific details, or other methods, components, etc., can be used. In other instances, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the invention. Illustrative embodiments of the invention will be best understood by reference to the accompanying drawings, wherein like parts are indicated by like numerals throughout the drawings. The following description is merely by way of example and simply illustrates certain selected embodiments of apparatus and methods consistent with the claims of the invention.

[0069] To improve the encoding and decoding performance of cross-component prediction, various schemes related to inheriting cross-component models are disclosed.

[0070] Inherited neighbor model parameters When applying cross-component predictive coding / decoding tools to the current block to generate a predicted signal, the model parameters contained in the cross-component model (CCM) information (see the section titled "Inheriting CCM Information" for more details) can be inherited from neighboring blocks. Further details about neighboring blocks are described in the sections titled "Inheriting Spatial Neighbor Model Parameters," "Inheriting Temporal Neighbor Model Parameters," "Inheriting Non-Neighboring Spatial Neighbor Models," and "Inheriting Model Parameters from History Tables."

[0071] The final scaling parameters for the current block are inherited from neighboring blocks and / or further refined by dA. Once the final scaling parameters are determined, the offset parameters (e.g., In CCLM, the scaling parameters are derived based on inherited scaling parameters and / or the average of the luminance and chrominance samples from the current block. For example, if the final scaling parameters are inherited from selected neighboring blocks, and the inherited scaling parameters are... The final scaling parameter is ( + dA). In another embodiment, the final scaling parameter is inherited from the history list and / or further refined by dA. For example, the history list records the j most recent final scaling parameter entries from previous CCLM encoded blocks. The final scaling parameter is then inherited from selected entries in the history list. The final scaling parameter is ( +dA). In another embodiment, the final scaling parameter is inherited from the history list or neighboring blocks, but is not further refined by dA.

[0072] In another embodiment, after inheriting the model parameters, the offset parameters can be further refined by dB. For example, if the final offset parameters are inherited from selected neighboring blocks, and the inherited offset parameters are... Then the final offset parameter is ( +dB). In another embodiment, the final offset parameter is inherited from the history list and further refined by dB. For example, the history list records the j most recent final offset parameter entries from previous CCLM encoded blocks. The final offset parameter is then inherited from selected entries in the history list. The final offset parameter is ( + dB). In another embodiment, the final offset parameter is inherited from the history list or neighboring blocks, but is not further refined by dB.

[0073] In another embodiment, if the inherited neighboring block is coded using CCCM, then the inherited filter coefficients ( Offset parameters (e.g., or In CCCM, it can be re-derived based on inherited parameters and the average value of luminance and chrominance samples from the corresponding neighboring positions of the current block.

[0074] In another embodiment, if the inherited candidate applies the GLM gradient mode to its brightness reconstruction sample, the current block should also inherit the candidate's GLM gradient mode and apply it to the current brightness reconstruction sample.

[0075] In another embodiment, if the inherited neighboring blocks are encoded with multiple cross-component models (e.g., MMLM or CCCM with multiple models), the classification threshold is also inherited to classify the neighboring samples of the current block into multiple groups, and the inherited multiple cross-component model parameters are further assigned to each group.

[0076] Inheriting CCM information In one embodiment, inherited cross-component model (CCM) information may be stored along with inherited model parameters. As described earlier in this disclosure, CCM information includes, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM), model indices indicating which model shape to use in the convolutional model, classification thresholds for multiple models, downsampling filter flags, downsampling filter indices, the number of neighboring lines used to derive the model, the template type used to derive the model, post-filter flags, or model parameters.

[0077] In one embodiment, a CCLM model can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCLM model.

[0078] In another embodiment, a CCLM model with nonlinear terms can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCLM model with nonlinear terms.

[0079] In one embodiment, a CCCM model can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCCM model. Luminance and chromaticity offsets used to adjust the CCCM model inputs can also be stored in the CCM information.

[0080] In another embodiment, CCCM models with different convolutional filter shapes can be inherited. In addition to model parameters and prediction modes, a CCCM mode index can be stored in the CCM information to indicate which convolutional filter shape the inherited CCCM model uses. For example, a CCCM model with different convolutional filter shapes may contain only horizontal spatial terms. As another example, a CCCM model with different convolutional filter shapes may contain only vertical spatial terms. As another example, a CCCM model with different convolutional filter shapes may contain only diagonal spatial terms. As another example, a CCCM model with different convolutional filter shapes may contain only anti-diagonal spatial terms. As another example, a CCCM model with different convolutional filter shapes may contain X-shaped spatial terms.

[0081] In another embodiment, a CCCM model using non-downsampled samples can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCCM model using non-downsampled samples.

[0082] In another embodiment, a CCCM model with multiple downsampling filters can be inherited. In addition to storing model parameters, a prediction mode can be stored in the CCM information to indicate that the inherited model is a CCCM model with multiple downsampling filters, and a model index can also be stored in the CCM information to indicate which variant of the CCCM model with multiple downsampling filters is being inherited.

[0083] In another embodiment, a hybrid CCCM model consisting of various terms (e.g., spatial, gradient, positional, nonlinear, and bias terms) can be inherited. The gradient term can be computed in either a downsampled or non-downsampled domain. The positional term can be computed relative to the top-left coordinates of the current block or image. In addition to storing model parameters, a prediction pattern can be stored in the CCM information to indicate that the inherited model is a hybrid CCCM model consisting of various terms. If multiple types of hybrid CCCM models exist, a model index can also be stored in the CCM information to indicate which type of hybrid CCCM model is being inherited. For example, the gradient and location-based convolutional cross-component model (GL-CCCM) proposed in JVET-AB0119 (Ramin G. Youvalari et al., “Non-EE2: Gradient and location based convolutional cross-component model (GL-CCCM) for intra prediction”, Jointvideo Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 28th meeting, Mainz, Germany, October 20-28, 2022, document: JVET-AB0119) is a hybrid CCCM model that includes a spatial term at the center location, two gradient terms in the horizontal and vertical directions, two X and Y position terms relative to the horizontal and vertical locations, a nonlinear term, and a bias term. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a GL-CCCM model.

[0084] In one embodiment, a GLM model can be inherited. In addition to storing model parameters, a prediction pattern can be stored in the CCM information to indicate that the inherited model is a GLM model, and a downsampling filter index can also be stored in the CCM information to indicate which gradient downsampling filter the inherited GLM model uses.

[0085] In another embodiment, a GLM model with a luminance term can be inherited. In addition to storing model parameters, a prediction mode can be stored in the CCM information to indicate that the inherited model is a GLM model with a luminance term, and a downsampling filter index can also be stored in the CCM information to indicate which gradient downsampling filter is used by the inherited GLM model with a luminance term.

[0086] In one embodiment, any type of cross-component multi-model can be inherited. In addition to storing model parameters and prediction patterns, a multi-model on / off flag can be stored in the CCM information to indicate whether the inherited CCM model is multi-model. If the multi-model on / off flag is true, the multi-model classification threshold is also stored in the CCM information.

[0087] In one embodiment, CCM information may include information indicating how the inherited model is derived. For example, CCM information may include the number of adjacent lines used to derive the cross-component model and / or the template type used to derive the model. For example, a set of templates may be used to derive the CCCM model. This set of templates includes templates with different positions, sizes, and shapes. CCM information may store the template index on which the inherited CCCM model is based. For example, the inherited CCCM model may be derived based on a top-only template, a left-only template, or both a left and top template. As another example, the inherited CCCM model may be derived based on a 6-line template or a 2-line template.

[0088] In one embodiment, a post-filter flag can be stored in the CCM information. This information describes how the inherited model is used in the block and where the inherited model comes from. If the post-filter flag is enabled, this indicates that predictions are applied to the block from which the inherited model comes.

[0089] Inheritance spatial proximity model parameters According to another embodiment, inherited model parameters can come from an adjacent block. Models from blocks at predefined locations are added to the candidate list in a predefined order. For example, the predefined location could be... Figure 7 The positions shown can be in the predefined order of B0, A0, B1, A1 and B2, or A0, B0, B1, A1 and B2. This block can be a chroma block or a single-tree luminance block.

[0090] In one embodiment, the predefined location and predefined order can be the same as the spatial candidates for the inter-frame merge mode.

[0091] According to another embodiment, assuming the current block's position, width, and height are (x, y), W, and H, respectively, the predefined positions include the position immediately above (W>>1) or ((W>>1) - 1) (i.e., the position immediately above the current block, such as (x + W>>1, y-1) or (x + (W+1)>>1, y-1)), if W is greater than or equal to TH, and the position immediately to the left (H>>1) or ((H>>1) - 1) (i.e., the position immediately to the left of the current block, such as (x-1, y+H>>1) or (x-1, y+(H+1)>>1)), if H is greater than or equal to TH, where W and H are the width and height of the current block, and TH is a threshold that can be 2, 4, 8, 16, 32, or 64.

[0092] According to another embodiment, the maximum number of inheritance models from spatial neighbors is less than the number of predefined locations. For example, if the predefined locations are as follows: Figure 7 As shown, there are 5 predefined positions. If the predefined order is B0, A0, B1, A1 and B2, and the maximum number of inherited models from spatial neighbors is 4, then a model from B2 will only be added to the candidate list if one of the preceding blocks is unavailable or not encoded in the cross-component model.

[0093] Inheritance time proximity model parameters In one embodiment, if the current slice / image is an intra-non-intra slice / image, the inherited model parameters can be derived from blocks in previously encoded slices / images.

[0094] According to another embodiment, if the current slice / image is an intra-frame or non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images. For example, such as Figure 11 As shown, the current block is located at (x, y), and the block size is... The inherited model parameters can come from blocks at positions (x', y'), (x', y' + h / 2), (x' + w / 2, y'), (x' + w / 2, y' + h / 2), (x' + w, y'), (x', y' + h), or (x' + w, y' + h) in previously encoded slices / images, where x' = x + Δx and y' = y + Δy. In one embodiment, if the prediction mode of the current block is intra-frame, Δx and Δy are set to 0. If the prediction mode of the current block is inter-frame, Δx and Δy are set to the horizontal and vertical motion vectors of the current block. In another embodiment, if the current block is inter-frame bidirectional prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 0. In yet another embodiment, if the current block is inter-frame bidirectional prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 1.

[0095] According to another embodiment, if the current block is an inter-frame bidirectional prediction, the inherited model parameters can come from blocks in previously encoded slices / images in the reference list. For example, if the horizontal and vertical portions of the motion vectors in reference image list 0 are... and The motion vectors can be scaled to other reference images in reference lists 0 and 1. If the motion vectors are scaled to reference list 0... The reference image is Then the model can come from reference list 0. Referring to the block in the reference image, and with Δx and Δy set to... For example, if the horizontal and vertical portions of the motion vectors in reference image list 0 are... and The motion vectors are scaled to those in reference list 1. The reference image is The model can be derived from reference list 1. Referring to the block in the reference image, and with Δx and Δy set to... .

[0096] In one embodiment, if the current slice / image is an intra-frame or non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images. In one embodiment, the current block is located at (x, y), and the block size is... Two value sets and Defined as: , .

[0097] and All values ​​in the set are positive. Let The inherited model parameters can come from previously encoded / decoded slices / images at positions... The block.

[0098] In one sub-implementation, .For example, Inherited model parameters can come from... Figure 12 The block locations in the previously encoded / decoded slice / image are shown.

[0099] In another sub-implementation, .For example, = {1 / 2, 1, 3 / 2, 2, 5 / 2} and = {1,2, 3, 4, 5}.

[0100] In another embodiment, the current block is located at (x, y) and the block size is [missing value]. Inherited model parameters can come from previously encoded / decoded slices / images located in , The block of location.

[0101] In one sub-implementation, .For example, .

[0102] In another sub-implementation, .For example, and .

[0103] In one embodiment, from closer The location model is first added to the final merge candidate list. In another embodiment, models from closer... The location model is first added to the final merge candidate list.

[0104] In one embodiment, let and These are two fixed positive numbers. Inherited model parameters can be derived from previously encoded / decoded slices / images located at... The block of location.

[0105] In another embodiment, the current block is located at (x, y) and the block size is [value missing]. .make and Given two fixed positive numbers, the inherited model parameters can be derived from previously encoded / decoded slices / images at positions […]. The block.

[0106] In another embodiment, the current block is located at (x, y) and the block size is [value missing]. The inherited model parameters can come from blocks at some predefined locations in previously encoded / decoded slices / images. For example, these positions are within the corresponding region of the current coding block, i.e. and Inherited model parameters can come from position . For example, these positions are outside the corresponding region of the current encoded block, i.e. and Inherited model parameters can come from position . The block.

[0107] In one embodiment, the previously encoded / decoded image from which the inherited parameter model comes (i.e., the co-located image) is one of the images in the reference list.

[0108] The previously encoded or decoded images from which the inherited parameter model comes are hereafter referred to as collocated pictures.

[0109] In one embodiment, the co-location image is signaled in the image / slice header. The reference list and reference index are signaled in the image / slice header. For example, the co-location image is selected as L0[0]. As another example, the co-location image is selected as L1[0].

[0110] In one embodiment, the co-placed image is selected as the image with the smallest POC (Picture Order Count) difference from the current image in the reference list. For example, if the current image has a POC of 8, the images in reference list 0 have POCs of {7, 6, 5, 0}, and the images in reference list 1 have POCs of {7, 6, 5, 4}, then L0[0] (equivalent to L1[0]) is selected because it has the smallest POC difference. In another sub-embodiment, if two images have the smallest POC difference from the current image, the image with the smaller POC is selected. In another sub-embodiment, if two images have the smallest POC difference from the current image, the image with the larger POC is selected. In another sub-embodiment, if two images have the smallest POC difference from the current image, the image with the smaller QP (Quantization Parameter) difference from the current image is selected. Then L0[1] (POC = 4 and QP = 26) is selected. In another embodiment, if two images have the smallest POC difference from the current image, the image with the smaller QP is selected. In another embodiment, if two images have the smallest difference in POC with the current image, the image with the larger QP is selected.

[0111] In one embodiment, the co-located image is selected as the image with the smallest QP difference between the reference list and the current image. For example, if the QP of the current image is 28, the QP of the images in reference list 0 is {19, 26, 23}, and the QP of the images in reference list 1 is {23, 22, 21}, then L0[1] is selected. In another sub-embodiment, if multiple images in the reference list have the smallest QP difference with the current image, the image with the smaller QP is selected. In another sub-embodiment, if multiple images have the smallest QP difference with the current image, the image with the larger QP is selected. In another sub-embodiment, if multiple images have the smallest QP difference with the current image, the image with the smaller POC distance is selected. In another sub-embodiment, if multiple images have the smallest QP difference with the current image, the image with the smaller POC is selected. In another sub-embodiment, if multiple images have the smallest QP difference with the current image, the image with the larger POC is selected.

[0112] In one embodiment, the co-placed image is selected as the image with the smallest QP in the reference list. In another embodiment, the co-placed image is selected as the image with the largest QP in the reference list.

[0113] In one embodiment, the co-placed image is selected based on a combination of some or all of the methods described above.

[0114] In one embodiment, the previously encoded image (i.e., the co-located image) that inherits the parametric model is the most recently encoded I-image. The cross-component model information of the most recently encoded I-slice / image is stored in a long-term reference buffer.

[0115] In one embodiment, the positions of the co-located image and the inherited parameter model are determined by the motion vectors of neighboring blocks. For example, if the current block is located at (x, y) and the block size is... The inherited model parameters can come from blocks in the co-location image at positions (x', y'), (x', y' + h / 2), (x' + w / 2, y'), (x' + w / 2, y' + h / 2), (x' + w, y'), (x', y' + h), or (x' + w, y' + h), where x' = x + Δx and y' = y + Δy. Δx and Δy are set as the L0 horizontal and vertical motion vectors of the neighboring blocks, respectively, and the co-location image is an L0 reference image indicated by the L0 motion vectors of the neighboring blocks. In another embodiment, if the neighboring blocks are bidirectional predictions, Δx and Δy are set as the L1 horizontal and vertical motion vectors of the neighboring blocks, respectively, and the co-location image is an L1 reference image indicated by the L1 motion vectors of the neighboring blocks. In one embodiment, the neighboring block is the left block of the current block. In another embodiment, the neighboring block is the upper block of the current block.

[0116] In one embodiment, the inherited parameter model is derived from the position of a previously encoded slice / image, determined by the motion vectors of neighboring blocks. Let Δx and Δy be the horizontal and vertical displacements determined based on the selected neighboring block motion vectors, the current block's position being (x, y), and the block size being... The inherited model parameters can come from the block at position (x', y'), where x' = x + Δx and y' = y + Δy, or x' = x + w / 2 + Δx and y' = y + h / 2 + Δy.

[0117] In another embodiment, the inherited model parameters can also come from block positions in the pattern described in the preceding paragraphs. These positions are centered at (x', y'), where x' = x + Δx and y' = y + Δy, or x' = x + w / 2 + Δx and y' = y + h / 2 + Δy. That is, the predefined positions are represented as... The inherited model parameters can come from ,in and The horizontal and vertical displacements are determined based on the motion vectors of the selected neighboring blocks. For example, assuming the current block size is... These two sets of values and Defined as: , .

[0118] exist and All values ​​in the table are positive. Inherited model parameters can be taken from previously encoded slices / images located in... The location of the block. To give another example, let's say... and These are two fixed positive numbers. The inherited model parameters can be derived from previously encoded slices / images located at... The location of the block. To give another example, inherited model parameters can come from previously encoded slices / images relative to... Blocks at certain predefined locations. These locations can be... To give another example, these positions could be... .

[0119] In one embodiment, neighboring blocks can be located at predefined locations. For example, this location could be located at... Figure 7 The A0 position is shown. Predefined positions can also be located at... Figure 7 The positions A1, B0, B1, and B2 are shown. For example, if the block at a predefined position is not an intra-frame block, then neighboring blocks are not selected. For example, if the block is an intra-frame block, then the L0 motion vector is selected. If the L0 motion vector is unavailable, then the L1 motion vector is selected. Again, for example, if the block is an intra-frame block, then the L1 motion vector is selected; if the L1 motion vector is unavailable, then the L0 motion vector is selected.

[0120] In another embodiment, when selecting neighboring blocks, there may be a predefined list of locations. These locations are arranged according to the inspection order. For example, these locations could be... Figure 7 The spatial location is shown. The selected neighboring block can be the location of the first intra-block in the list. First, the L0 motion vector is selected. If the L0 motion vector is unavailable, the L1 motion vector is selected. For example, first, the L1 motion vector is selected. If the L1 motion vector is unavailable, the L0 motion vector is selected.

[0121] For example, the positions in the list are checked according to a predefined checking order. For each position, the L0 motion vector is checked first, then the L1 motion vector. Or, for example, the L1 motion vector is checked first, then the L0 motion vector. The selected motion vector is the first motion vector whose reference image is the co-location image. That is, the co-location image is determined before selecting neighboring motion vectors. The co-location image can be determined based on the method described above (i.e., the method described in the preceding paragraphs of this section).

[0122] In one embodiment, the horizontal and vertical displacements Δx and Δy are determined based on the selected motion vector of the neighboring block. In one sub-implementation, for example, if the reference image and the co-located image of the selected motion vector are the same image, then Δx is equal to the horizontal portion of the selected motion vector, and Δy is equal to the vertical portion of the selected motion vector. In another sub-implementation, for example, if the reference image and the co-located image of the selected motion vector are not the same image. The reference image can be an image in a reference list, where the co-located image is signaled in the image / slice header or determined based on the method described above. Let the POC distance between the current image and the reference image of the selected motion vector be tb, the POC distance between the current image and the co-located image be td, and the selected motion vector be (mv_x, mv_y). Δx = mv_x (td / tb) and Δy = mv_y (td / tb).

[0123] In one embodiment, the inherited model parameters are derived using the luminance and chrominance reconstruction samples of the co-located block. Let the current block position be (x, y), and the block size be... A co-located block is a block located at position (x', y') in a co-located image, and its size is [size missing]. When the inherited model comes from position (x', y'). For example, a co-location block can be a block located at position (x', y') in a co-location image, with a block size of... Where m and n are fixed positive values. For example, a co-placement block can be located at (x, y). As another example, if Δx and Δy are the L0 horizontal and vertical motion vectors of neighboring blocks, and the co-placement image is an L0 reference image indicated by the L0 motion vectors of neighboring blocks, then the co-placement block can be located within the co-placement image. (x', y') can be a block position in the pattern described in the preceding paragraph. For example, (x', y') could be... .

[0124] In one embodiment, the cross-component parameter model can be inherited from multiple previously encoded and decoded images. The number of co-located images (i.e., the previously encoded and decoded images from which the model parameters are inherited) can be signaled in PH, SH, PPS, or SPS, or can be determined by some predefined rules.

[0125] In one embodiment, the cross-component parameter model can be inherited from multiple previously encoded / decoded images. The cross-component parameter model can be inherited from an image set containing N previously encoded / decoded images. For example, the image set could be a reference list. The images from which the cross-component parameter model inherits can be determined using the methods described in the preceding paragraphs of this section. For example, these methods could be selecting images with a smaller QP, a larger QP, a smaller QP difference from the current image, a smaller POC difference from the current image, and / or a combination thereof.

[0126] In one embodiment, the cross-component parameter model can be inherited from multiple previously encoded / decoded images. The cross-component parameter model can be inherited from an image set containing N previously encoded / decoded images. An index can be signaled / parsed in the bitstream to indicate the selected image. The index ranges from 0 to N-1.

[0127] In one sub-implementation, the image with a smaller POC difference from the current image is associated with a smaller index. In another sub-implementation, the image with a smaller QP difference from the current image is associated with a smaller index. In yet another sub-implementation, the image with a smaller QP is associated with a smaller index. In yet another sub-implementation, the image with a larger QP is associated with a smaller index.

[0128] Inheriting the non-adjacent spatial proximity model In another embodiment, the inherited model parameters can come from non-adjacent spatial neighbor blocks. Models from blocks at predefined locations are added to the candidate list in a predefined order. For example, the predefined order is the same as the order of non-adjacent spatial neighbor candidates in the intra-frame merge mode. For example, the location and order can be as follows: Figure 10 As shown. The positions of the numbered squares are predefined. The numbers within each square represent a predefined order. The distance between each position is proportional to the width and height of the current encoding / decoding block.

[0129] In another embodiment, the maximum number of inherited models from non-adjacent spatial neighbors that can be added to the candidate list is less than the number of predefined locations. For example, if the predefined locations are as follows: Figure 13A -B is shown, which displays two patterns (1310 and 1320). The maximum number of inherited models from non-neighboring spatial neighbors that can be added to the candidate list is... Then only if the number of available models from the position of search pattern 1 is less than Only when this happens will search mode 2 be used (i.e., only add models from the mode 2 location).

[0130] Inheriting model parameters from history table In one embodiment, the inherited model parameters can come from a cross-component model history table. Cross-component models in the history table can be added to the candidate list in a predefined order. In one embodiment, the order of adding historical candidates can be from the beginning to the end of the table. In another embodiment, the order of adding historical candidates can be from a predefined position to the end of the table. In yet another embodiment, the order of adding historical candidates can be from the end to the beginning of the table. In yet another embodiment, the order of adding historical candidates can be staggered (e.g., the first added candidate comes from the beginning of the table, the second added candidate comes from the end of the table, and so on).

[0131] In one embodiment, a single cross-component model history table can be maintained to store previous cross-component models, and the cross-component model history table can be reset at the current image, current tile, current tile, every M CTU rows, or at the start of every N CTU, where N and M can be any values ​​greater than 0. In another embodiment, the cross-component model history table can be reset at the current image, current tile, current tile, current CTU row, or at the end of the current CTU.

[0132] Inherited from the fusion model Fusion mode refers to the mode in which two predictions are combined to generate a final prediction. In chroma intra-frame fusion mode, a chroma intra-frame prediction generated without using cross-component prediction (CCP) codecs (e.g., CCLM, MMLM, CCCM) is fused with another chroma intra-frame prediction generated using a cross-component prediction codec. For example, a non-CCLM-coded intra-frame prediction and a CCLM-coded intra-frame prediction are fused together to obtain the final intra-frame prediction.

[0133] In one embodiment, when inheriting cross-component model parameters from blocks / positions encoded by chroma intra-fraction mode, the model parameters for CCP-coded intra-prediction are inherited and further improved.

[0134] In one embodiment, in addition to inheriting and improving the CCP model parameters, the fusion weights and the encoding / decoding mode of non-CCP coded intra-frame prediction are also inherited. That is, the chroma intra-frame fusion mode is inherited.

[0135] Candidate list construction In one embodiment, the candidate list is constructed by adding candidates in a predefined order until a maximum number of candidates is reached. The added candidates may include, but are not limited to, all of the candidates described above. For example, the candidate list may include spatially nearby candidates, temporally nearby candidates, historical candidates, and non-neighborly nearby candidates. In another example, the candidate list may include the same candidates as in the previous example, but the candidates are added to the list in a different order.

[0136] In another embodiment, if all predefined neighboring and historical candidates are added but the maximum number of candidates is not reached, some default candidates are added to the candidate list until the maximum number of candidates is reached.

[0137] In one sub-implementation, the preset candidates include, but are not limited to, the candidates described below. Final scaling parameters. From collection And offset parameter Alternatively, it can be derived based on neighboring luminance and chrominance samples. For example, if the average values ​​of neighboring luminance and chrominance samples are lumaAvg and chromaAvg, respectively, then... Depend on Export.

[0138] In another embodiment, the preset candidate can be an earlier candidate refined with an incremental scaling parameter. For example, if the scaling parameter of the earlier candidate is... The preset candidate scaling parameter is ,in Can come from a set And the preset candidate offset parameter will be determined by... It is derived from the average values ​​of the luminance and chrominance samples of the current block.

[0139] Candidate index for signal transmission inheritance in the list An on / off flag can be signaled to indicate whether the current block inherits cross-component model parameters from neighboring blocks. This flag can be signaled by CU / CB, PU, ​​TU / TB, color component, or chroma color component. Advanced syntax can be signaled in SPS, PPS (Picture Parameter Set), PH (Picture Header), or SH (Slice Header) to indicate whether the proposed method is allowed for the current sequence, picture, or slice.

[0140] The maximum allowed candidate number is signaled to indicate the maximum size of the merge candidate list. This number can be signaled by CU / CB, PU, ​​TU / TB, color component, or chroma color component. Advanced syntax can be signaled in SPS, PPS, PH, or SH to indicate whether the proposed method is allowed for the current sequence, picture, or slice. The maximum allowed candidate number for the proposed method can be shared with the maximum allowed candidate number for intra-frame merge modes.

[0141] If the current block inherits cross-component model parameters from neighboring blocks, an inheritance candidate index is signaled. This index can be signaled (e.g., using truncated unary codes, Exp-Golomb codes, or fixed-length codes) and shared between the current Cb and Cr blocks. Another example is that the index can be signaled per color component. For instance, one inheritance candidate index can be signaled for the Cb component, and another for the Cr component. As yet another example, the inheritance index can be stored using chroma intra-prediction syntax (e.g., IntraPredModeC[xCb][yCb]).

[0142] If the current block inherits cross-component model parameters from neighboring blocks, the current chroma intra-prediction mode (e.g., IntraPredModeC[xCb][yCb] as defined in the VVC standard) is temporarily set to a cross-component mode (e.g., CCLM_LA) during the bitstream parsing phase. Subsequently, during the prediction or reconstruction phase, a candidate list is derived, and the inherited candidate model is determined by the inherited candidate index. After obtaining the inherited model, the codec information of the current block is updated according to the inherited candidate model. The codec information of the current block includes, but is not limited to, the prediction mode (e.g., CCLM_LA or MMLM_LA), the associated submode flags (e.g., CCCM mode flag), the prediction mode (e.g., GLM mode index), and the current model parameters. Then, the prediction for the current block is generated based on the updated codec information.

[0143] Any of the proposed methods described above can be implemented in the encoder and / or decoder. For example, any cross-component prediction model derivation method can be implemented in the intra-frame / intra-frame prediction module of the encoder and / or the intra-frame / intra-frame prediction module of the decoder. Alternatively, any proposed method can be implemented as a circuit coupled to the intra-frame / intra-frame prediction module of the encoder and / or the intra-frame / intra-frame prediction module of the decoder to provide the information required by the intra-frame / intra-frame prediction module. For example, an encoding / decoding system using a cross-component prediction model derivation method can be implemented in the reference encoder / decoder in Figures 1A-B. For example, any proposed cross-component prediction model derivation method can be implemented in the intra-frame / intra-frame encoding / decoding module in the decoder (e.g., Figure 1B Intra-frame prediction 150 / MC 152 in the encoder or intra-frame / intra-frame codec module in the encoder (e.g., Figure 1A The intra-frame prediction (110) and inter-frame prediction (112) are implemented in the model. Any proposed cross-component prediction model derivation method can also be implemented as a circuit, coupled to the intra-frame / intra-frame codec module of the decoder or encoder. However, the decoder or encoder can also use additional processing units to implement the required cross-component prediction processing. Although the intra-frame prediction unit (e.g., Figure 1AUnits 110 / 112 and Figure 1B Units 150 / 152 in the diagram are shown as independent processing units, but they may correspond to executable software or firmware code stored on media (such as hard disks or flash memory) for CPUs (central processing units) or programmable devices (such as DSPs (digital signal processors) or FPGAs (field programmable gate arrays)).

[0144] Figure 14 A flowchart illustrating an exemplary video codec system for deriving motion vectors from neighboring blocks to determine temporal candidates according to an embodiment of the present invention is provided. The steps shown in the flowchart can be implemented at the encoder end as program code executable on one or more processors (e.g., one or more central processing units (CPUs)). The steps shown in the flowchart can also be implemented based on hardware, such as one or more electronic devices or processors configured to perform the steps in the flowchart. According to the method, in step 1410, input data associated with a current block is received, the current block including a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data to be decoded at the decoder end associated with the current block, and wherein the current block is located in an intra-frame non-intra-frame slice or picture. In step 1420, a target motion vector is determined based on one or more neighboring blocks of the current block. In step 1430, a temporal candidate is determined based on the target motion vector, wherein a target cross-component model is generated for the second color block using one or more cross-component model parameters associated with the temporal candidate. In step 1440, a candidate list containing the temporal candidate is determined. In step 1450, the candidate list is used to encode or decode the second color patch, wherein when the time candidate is selected, cross-component prediction data is generated for the second color patch according to the target cross-component model.

[0145] Figure 15A flowchart illustrating an exemplary video codec system according to an embodiment of the present invention, which inherits a cross-component model from two or more previously encoded / decoded images, is provided. According to the method, in step 1510, input data associated with a current block is received, the current block comprising a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data to be decoded at the decoder end associated with the current block, and wherein the current block is located in an intra-frame non-intra-frame slice or image. In step 1520, a plurality of co-location images are determined for the current block, wherein the plurality of co-location images correspond to previously encoded / decoded images. In step 1530, temporal candidates are determined based on two or more co-location images, wherein a target cross-component model is derived by inheriting one or more cross-component model parameters from the two or more co-location images. In step 1540, a candidate list containing the temporal candidates is determined. In step 1550, the candidate list is used to encode or decode the second color block, wherein when the temporal candidate is selected, cross-component prediction data is generated for the second color block according to the target cross-component model.

[0146] The flowchart shown is intended to illustrate exemplary video encoding and decoding according to the present invention. Those skilled in the art can modify each step, rearrange the steps, split the steps, or combine the steps to practice the invention without departing from its spirit. Specific syntax and semantics are used in this disclosure to illustrate examples of implementing the invention. Those skilled in the art can practice the invention by using equivalent syntactic and semantic substitutions without departing from its spirit.

[0147] The foregoing description is intended to enable those skilled in the art to practice the invention in the context of specific applications and their requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the invention is not intended to be limited to the specific embodiments shown and described, but should be consistent with the broadest scope consistent with the principles and novel features disclosed herein. In the foregoing detailed description, various specific details have been set forth in order to provide a thorough understanding of the invention. However, those skilled in the art will understand that the invention can be practiced.

[0148] The embodiments of the present invention described above can be implemented in various hardware, software program code, or a combination of both. For example, one embodiment of the invention may be one or more circuits integrated into a video compression chip, or program code integrated into video compression software, to perform the processes described herein. Another embodiment of the invention may also be program code to be executed on a digital signal processor (DSP) to perform the processes described herein. The invention may also relate to multiple functions executed by a computer processor, digital signal processor, microprocessor, or field-programmable gate array (FPGA). These processors can be configured to execute machine-readable software program code or firmware program code to perform the specific methods embodied in the invention. The software program code or firmware program code can be developed in different programming languages ​​and different formats or styles. The software program code can also be compiled for different target platforms. However, different program code formats, styles, and languages, as well as other methods of configuring program code to perform the tasks according to the invention, do not depart from the spirit and scope of the invention.

[0149] This invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples should be considered illustrative in all respects, not restrictive. Therefore, the scope of the invention is indicated by the appended claims rather than the foregoing description. All modifications within the meaning and equivalence of the claims should be included within its scope.

Claims

1. A method for encoding and decoding color images, using an encoding and decoding tool that includes one or more cross-component model correlation modes, the method comprising: Receive input data associated with the current block, which includes a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data to be decoded at the decoder end associated with the current block, and wherein the current block is located in a non-intra-frame slice or image; The target motion vector is determined based on one or more neighboring blocks of the current block; A time candidate is determined based on the target motion vector, wherein a target cross-component model is generated for the second color block by using one or more cross-component model parameters associated with the time candidate; Determine a candidate list that includes candidates for that time period; as well as The candidate list is used to encode or decode the second color patch, wherein when the time candidate is selected, cross-component prediction data is generated for the second color patch according to the target cross-component model.

2. The method according to claim 1, characterized in that, The one or more neighboring blocks are located at one or more predefined locations.

3. The method according to claim 2, characterized in that, The selected target motion vector corresponds to either L0 or L1 motion vector.

4. The method according to claim 2, characterized in that, If the target neighboring block of one or more neighboring blocks is not an intra-frame block, then the target neighboring block of the one or more neighboring blocks will not be selected for the current block to determine the time candidate.

5. The method according to claim 1, characterized in that, One or more neighboring blocks of the current block correspond to a predefined list of locations, and the target neighboring block is selected from the predefined list of locations according to a predefined check order.

6. The method according to claim 5, characterized in that, The target neighboring block corresponds to the first position in the predefined location list, which is an intra-frame block.

7. The method according to claim 5, characterized in that, The target neighboring block corresponds to the first position in the predefined position list, and the first position has a reference image as a co-located image of the current block.

8. The method according to claim 1, characterized in that, The target cross-component model inherits from previously encoded and decoded images, which are selected from reference images in one or more reference lists.

9. The method according to claim 8, characterized in that, The target reference image selected from the reference images in one or more reference lists is signaled in the image or slice header.

10. The method according to claim 8, characterized in that, The target reference image is implicitly selected from multiple reference images in one or more reference lists according to a set of predefined rules.

11. The method according to claim 10, characterized in that, This set of predefined rules includes selecting the candidate reference image with the smallest picture order count (POC) difference or the smallest quantization parameter (QP) difference, or the candidate reference image with the smallest QP among the multiple reference images, as the target reference image.

12. The method according to claim 10, characterized in that, This set of predefined rules includes selecting a candidate reference image corresponding to the most recently encoded I-image as the target reference image.

13. An apparatus for video encoding and decoding, the apparatus comprising one or more electronic devices or processors configured to: Receive input data associated with the current block, including a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data to be decoded at the decoder end associated with the current block, and wherein the current block is located in a non-intra-frame slice or image; The target motion vector is determined based on one or more neighboring blocks of the current block; A time candidate is determined based on the target motion vector, wherein a target cross-component model is generated for the second color block by using one or more cross-component model parameters associated with the time candidate; Determine a candidate list that includes candidates for that time; and The candidate list is used to encode or decode the second color patch, wherein when the time candidate is selected, cross-component prediction data is generated for the second color patch according to the target cross-component model.

14. A method for encoding and decoding color images, using an encoding / decoding tool that includes one or more cross-component model correlation modes, the method comprising: Receive input data associated with the current block, which includes a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data to be decoded at the decoder end associated with the current block, and wherein the current block is located in a non-intra-frame slice or image; Multiple co-location images are determined for the current block, wherein the multiple co-location images correspond to previously encoded or decoded images; Time candidates are determined based on two or more co-located images, wherein a target cross-component model is derived by inheriting one or more cross-component model parameters from the two or more co-located images. Determine a candidate list that includes candidates for that time period; as well as The candidate list is used to encode or decode the second color patch, wherein when the time candidate is selected, cross-component prediction data is generated for the second color patch according to the target cross-component model.

15. The method according to claim 14, characterized in that, The total number of multiple co-located images is transmitted or parsed in the bitstream via signaling.

16. The method according to claim 15, characterized in that, The total number of multiple co-located images is transmitted or parsed in the bitstream's PH (image header), SH (slice header), PPS (image parameter set), or SPS (sequence parameter set).

17. The method according to claim 14, characterized in that, Multiple co-located images correspond to an image set consisting of N previously encoded images.

18. The method according to claim 17, characterized in that, Signaling or parsing a target index from a bitstream to indicate one of two or more co-located images selected from a set of images.

19. The method according to claim 18, characterized in that, If the difference in Picture Order Count (POC) between a member image in the image set and the current image is small, then the index value of the member image is small.

20. The method according to claim 18, characterized in that, If the difference in the image-level quantization parameter (QP) between a member image in the image set and the current image is small, then the index value of the member image is small.

21. The method according to claim 18, characterized in that, If the image-level QP of a member image in a set is small, then the index value of that member image is small.

22. The method according to claim 18, characterized in that, If the image-level QP of a member image in a set is large, then the index value of that member image is small.

23. The method according to claim 17, characterized in that, Two or more co-located images are selected from an image set according to a set of predefined rules.

24. The method according to claim 23, characterized in that, This set of predefined rules corresponds to selecting one or more candidate images from the image set, which have smaller QP, larger QP, smaller QP difference from the current image, smaller POC difference from the current image, and / or a combination thereof.

25. An apparatus for video encoding and decoding, the apparatus comprising one or more electronic devices or processors configured to: Receive input data associated with the current block, which includes a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data associated with the current block to be decoded at the decoder end, and wherein the current block is located in a non-intra-frame slice or picture; Multiple co-location images are identified for the current block, where the multiple co-location images correspond to images previously encoded or decoded; A temporal candidate is determined based on two or more co-located images, wherein the target cross-component model is derived by inheriting one or more cross-component model parameters from the two or more co-located images; Determine a candidate list that includes time-related candidates; as well as The candidate list is used to encode or decode the second color patch, wherein when a time candidate is selected, cross-component prediction data is generated for the second color patch according to the target cross-component model.