Video coding and decoding method and device for inheriting cross component model from scaled reference picture

By employing a cross-component model and reference image resampling technology in a multi-functional video coding system, the problem of low efficiency in color image coding is solved, achieving more efficient chroma component coding and better coding performance.

CN121773618APending Publication Date: 2026-03-31MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing multi-functional video coding technologies suffer from low coding efficiency and insufficient coding performance when processing color images, especially when using cross-component models, which cannot effectively utilize reference image resampling to improve the coding quality of chroma components.

Method used

A cross-component model is adopted, which derives the cross-component prediction model from the reference image by resampling, selects co-position images using predefined rules, and applies the cross-component prediction model to encode the chroma component, including using luminance sample gradients and convolutional filters for chroma prediction. Multiple models such as CCLM, CCCM, GLM and CCRM are combined to improve coding performance.

Benefits of technology

It improves the coding efficiency and quality of color images, reduces redundancy of chrominance components, and enhances the performance of the coding system, especially in terms of accuracy and efficiency in inter-frame prediction and intra-frame prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121773618A_ABST
    Figure CN121773618A_ABST
Patent Text Reader

Abstract

A method and apparatus for coding and decoding a color picture or movie using a coding and decoding tool including one or more modes associated with a cross-component model is disclosed. In the method, co-located pictures are selected from one or more reference picture lists according to one or more predefined rules. It is determined whether the selected co-located picture is a reference picture resampled picture. And deriving a current cross component prediction model from the co-located picture according to whether the co-located picture is a reference picture resampling picture or not. And coding and decoding the current second color block using a candidate list comprising the current cross-component prediction model, characterized in that when the current cross-component prediction model is selected for coding and decoding the current second color block, prediction data of the current second color block is generated by applying the current cross-component prediction model to the current first color block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a video encoding and decoding system that uses encoding and decoding tools that include one or more patterns related to Cross-Component Models (CCM). In particular, this application relates to encoding and decoding chroma components using a cross-component model derived from a Reference Picture Resampling (RPR) reference image. Background Technology

[0002] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the Joint Video Experts Team (JVET) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology—Encoding representation of immersive media—Part 3: Versatile Video Coding, published in February 2021. VVC was developed based on its predecessor, High Efficiency Video Coding (HEVC), by adding more coding and decoding tools to improve coding efficiency and handle various types of video sources, including 3D video signals.

[0003] Figure 1AAn exemplary adaptive intra / inter-frame movie coding system incorporating loop processing is illustrated. For intra-frame prediction 110, prediction data is derived from previously encoded movie data in the current frame. For inter-frame prediction 112, motion estimation (ME) is performed at the encoder, and motion compensation (MC) is performed based on the results of ME to provide prediction data derived from other frames and motion data. Switch 114 selects either intra-frame prediction 110 or inter-frame prediction 112 and provides the selected prediction data to adder 116 to form a prediction error, also known as a residual. The prediction error is then processed by transform (T) 118, followed by quantization (Q) 120. The residuals from the transform and quantization are then encoded by entropy encoder 122 to be included in the movie bitstream corresponding to the compressed movie data. The bitstream associated with the transform coefficients is then packaged with side information, such as motion and encoding / decoding modes associated with intra-frame and inter-frame prediction, and parameters associated with loop filters applied to the underlying image regions. Side information related to intra-frame prediction 110, inter-frame prediction 112, and in-loop filter (ILPF) 130, such as Figure 1A As shown, this is provided to the entropy encoder 122. When using inter-frame prediction mode, the encoder must also reconstruct the reference picture. Therefore, the residuals from transformation and quantization are processed by inverse quantization (IQ) 124 and inverse transformation (IT) 126 to recover the residuals. The residuals are then added back to the prediction data 136, and the video data is reconstructed at reconstruction (REC) 128. The reconstructed video data may be stored in the reference picture buffer 134 and used for prediction of other frames.

[0004] like Figure 1AAs shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subject to various degradations due to these processes. Therefore, before storing the reconstructed video data in the reference picture buffer 134, a loop filter 130 is typically applied to the reconstructed video data to improve video quality. For example, a deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF) may be used. It may be necessary to incorporate loop filter information into the bitstream so that the decoder can correctly recover the required information. Therefore, loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. Figure 1A In the process, the loop filter 130 is applied to the reconstructed video, and then the reconstructed samples are stored in the reference image buffer 134. Figure 1A The system described herein is intended to illustrate an exemplary architecture of a typical video encoder. It may correspond to a High Performance Video Encoder / Decoder (HEVC) system, VP8, VP9, ​​H.264, or VVC.

[0005] like Figure 1B As shown, the decoder can use the same or partially the same function blocks as the encoder, except for transform 118 and quantization 120, since the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses an entropy decoder 140 instead of an entropy encoder 122 to decode the movie bitstream into quantized transform coefficients and the required encoding / decoding information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). Intra-frame prediction 150 at the decoder end does not require mode search. Instead, the decoder only needs to generate intra-frame predictions based on the intra-frame prediction information received from the entropy decoder 140. Furthermore, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from the entropy decoder 140, without motion estimation.

[0006] To improve the encoding and decoding performance of systems using cross-component models, a method and apparatus for using cross-component models derived from reference picture resampling (RPR) reference pictures are disclosed. Summary of the Invention

[0007] This application discloses a method and apparatus for using encoding and decoding tools to include one or more pattern-encoded color images or videos related to a cross-component model. According to this method, input data related to the current block is received, including a first color block and a second color block. The input data includes pixel data to be encoded at the encoder end or data related to the current block to be decoded at the decoder end, and the current block is encoded and decoded in a non-intra-frame mode. A co-op image is selected from one or more lists of reference images according to one or more predefined rules. It is determined whether the selected co-op image is a Reference Picture Resampling (RPR) image. Based on whether the co-op image is a Reference Picture Resampling (RPR) image, the current cross-component prediction (CCP) model is derived from the co-op image. Encoding or decoding the current second color block using a candidate list including the current cross component prediction (CCP) model, characterized in that, when the current cross component prediction (CCP) model is selected for encoding or decoding the current second color block, prediction data for the current second color block is generated by applying the current cross component prediction (CCP) model to the current first color block.

[0008] In one embodiment, the one or more predefined rules and information include L0[0], L1[0], Picture Order Count (POC) distance, QP value, or a combination thereof. In one embodiment, whether the selected co-image is the Reference Picture Resampled (RPR) image is indicated by a flag emitted or parsed in the bitstream. In one embodiment, the selected co-image is the Reference Picture Resampled (RPR) image if the co-image and the current image containing the current block differ in one or more target parameters. In one embodiment, the one or more target parameters include the image width in the luminance sample, the image height in the luminance sample, the left offset of the scaling window, the right offset of the scaling window, the top offset of the scaling window, the bottom offset of the scaling window, the number of sub-images, or a combination thereof.

[0009] In one embodiment, the co-position image is selected from one or more unscaled images in one or more reference image lists. In another embodiment, a target image is selected as the co-position image from the unscaled images in the one or more reference image lists, characterized in that the target image is selected based on the image order count (POC) difference with the current image, the image order count (POC) value, the QP difference with the current image, the QP value, the reference list, or a combination thereof.

[0010] In one embodiment, if the selected co-image is a reference image resampled (RPR) image, the cross component model (CCM) information inherited from the co-image is disabled. In another embodiment, if the selected co-image is a reference image resampled (RPR) image, the CCM information from the co-image is retrieved from the zoom position based on the zoom ratio. In one embodiment, the zoom ratio is derived based on the zoom window of the current image and the co-image. In another embodiment, if the current position and the zoom ratio are represented as (x, y) and R, respectively, the zoom position is determined according to (x / R, y / R) after a rounding process. In one embodiment, the rounding process corresponds to rounding towards negative infinity, rounding towards positive infinity, rounding towards zero, or rounding to the nearest integer.

[0011] In one embodiment, the current cross-component prediction (CCP) model from the co-located image is determined based on the motion vectors of neighboring blocks, and the neighboring blocks are selected from a list of predefined locations. In another embodiment, the neighboring blocks are selected from the list of predefined locations according to a predefined checking order. In one embodiment, the predefined checking order corresponds to checking either L0 or L1 motion vectors first, and selecting the target motion vector associated with the unscaled reference image first.

[0012] In one embodiment, if the target reference image located by the motion vector is a reference image resampled (RPR) image, then the motion vector is considered to have no cross component model (CCM) information located by the motion vector. In another embodiment, if the target reference image located by the motion vector is the reference image resampled (RPR) image, then cross component model (CCM) information is retrieved from the target reference image at the scaled position according to the scaling ratio. Attached Figure Description

[0013] Figure 1A An exemplary adaptive inter-frame / intra-frame video coding system incorporating loop processing is demonstrated.

[0014] Figure 1B Showing the corresponding Figure 1A The decoder of the encoder.

[0015] Figure 2 It shows 16 gradient modes of GLM (Gradient Linear Model).

[0016] Figure 3It shows the 6-tap space terms corresponding to the 6 neighboring luminance samples (i.e., L0, L1, ..., L5) around the chrominance sample (i.e., C) to be predicted, for the Convolutional Cross-Component Model (CCCM) mode.

[0017] Figure 4 An exemplary system block diagram of the Cross-component residual model (CCRM) is shown.

[0018] Figure 5 Five neighboring blocks are shown for exporting VVC space merge candidates.

[0019] Figure 6 Exemplary patterns of candidate merge options for adjacent and non-adjacent spaces are shown.

[0020] Figure 7 An example of time-based candidate export is shown, characterized by exporting scaled motion vectors based on PictureOrder Count (POC) distance.

[0021] Figure 8 It shows the position of the time candidate when choosing between candidates C0 and C1.

[0022] Figure 9A (Mode 1) and Figure 9B (Mode 2) Two different non-adjacent spatial proximity candidate modes are shown based on predefined location and predefined order.

[0023] Figure 10 An example of CCM information dissemination is shown.

[0024] Figure 11 This demonstrates an example of mapping a position outside a corresponding CTU line to a position inside the corresponding CTU line.

[0025] Figure 12A and Figure 12B An n-tap pattern is demonstrated in a window region M x N surrounding / containing position (iL, jL) to derive sourceTermSet0(i, j), characterized by using only the center ( Figure 12A ) and using 5x5 cross ( Figure 12B ).

[0026] Figure 13 An example is shown of using Sobel filters to derive gradient information from the predicted and / or reconstructed samples of the source.

[0027] Figure 14A and Figure 14B This demonstrates an m-tap pattern within a window region M2 x N2 surrounding / containing position (iC, jC) to derive sourceTermSet1(i, j), characterized by using only the center ( Figure 14A ) and using 5x5 cross ( Figure 14B ).

[0028] Figure 15 An example of a neighboring spatial region is shown as a reference region for weight setting in a self-derived cross-component model.

[0029] Figure 16 A flowchart of an exemplary video encoding and decoding system is shown, which, according to one embodiment of this application, inherits a cross-component prediction model from a reference picture resampling (RPR) reference picture. Detailed Implementation

[0030] The components of this application, as generally described and depicted in the figures, can be arranged and designed in various different configurations. Therefore, the following more detailed description of embodiments of the systems and methods of this application, as illustrated, is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the invention. References to “one embodiment,” “an embodiment,” or similar language in this specification mean that a particular feature, structure, or characteristic that may be included in the description associated with said embodiment is included at least in one embodiment of this application. Therefore, the phrases “in one embodiment” or “in a kind of embodiment” appearing throughout this specification do not necessarily refer to the same embodiment.

[0031] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the invention can be practiced without using one or more specific details, or using other methods, components, etc. In other instances, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the invention. Embodiments of the invention can be best understood by referring to the drawings, characterized by the same parts being specified by the same numbers throughout. The following description is by way of example only and illustrates only certain selected embodiments of apparatus and methods consistent with the invention declared herein.

[0032] Cross-Component Linear Model (CCLM) Prediction To reduce redundancy in cross components, VVC uses a cross-component linear model (CCLM) prediction mode, characterized by predicting chrominance samples based on luminance samples reconstructed using the same coding unit (CU) of the linear model, as shown below: (1) Its features Represents the predicted chromaticity samples in CU. This represents the downsampled reconstructed brightness sample from the same CU.

[0033] CCLM parameters ( and The value is derived from at most four neighboring chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block size is W×H, then W' and H' are set to... W'= W, H'= H when LM_LA mode is applied; W' = W + H when LM_A mode is applied; H' = H + W when the LM_L mode is applied.

[0034] In this disclosure, the terms {LM_LA, LM_A, LM_L} and {CCLM_LT, CCLM_T, CCLM_L} are used interchangeably.

[0035] Multiple Model CCLM (MMLM) In JEM (J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Film Exploration Team (JVET), Jul. 2017), a multi-model CCLM (MMLM) mode was proposed to predict chromaticity samples of the entire CU from luminance samples using two models. In MMLM, neighboring luminance samples and neighboring chromaticity samples of the current block are classified into two groups, each group serving as a training set to derive a linear model (i.e., deriving specific α and β for a specific group). Furthermore, samples of the current luminance block are also classified according to the classification rules of neighboring luminance samples.

[0036] The threshold is calculated as the average of neighboring reconstructed brightness samples. Rec′ L Neighbor samples with [x,y]<= threshold are classified as group 1; while Rec′ LThe samples near the [x,y] threshold are classified as group 2.

[0037] (2) Convolutional Cross-Component Model (CCCM) In CCCM, a convolutional model is applied to improve chromaticity prediction performance. The convolutional model has a 7-tap filter, consisting of a 5-tap plus shapespace component, a nonlinear term, and a bias term.

[0038] The filter output is calculated as the convolution between the filter coefficients and the input values, and then clipped to the range of valid chromaticity samples.

[0039] The filter coefficients are calculated by minimizing the mean square error (MSE) between the predicted and reconstructed chromaticity samples in the reference region.

[0040] Gradient Linear Model (GLM) Compared to CCLM, GLM uses luminance sample gradients instead of downsampled luminance values ​​to derive a linear model. Specifically, when applying GLM, the input to the CCLM process is the downsampled luminance samples... The gradient of the brightness sample Replacement. Other parts of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged: .

[0041] Figure 2 The 16 gradient filters (210-240) used for gradient calculation are shown.

[0042] In-screen block copying Intra-block copy (IBC) is a tool used in HEVC extensions of screen content coding (SCC). It is well known for significantly improving the encoding and decoding efficiency of screen content materials. Since IBC mode is implemented as a block-level encoding / decoding mode, block matching (BM) is performed at the encoder end to find the optimal block vector (or motion vector) for each CU. Here, the block vector indicates the displacement from the current block to a reference block that has been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pel and 4-pel motion vector precision. IBC-encoded CUs are considered a third prediction mode in addition to intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.

[0043] Cross-component chromaticity model using non-down-sampled luminance samples (CCCM) The CCCM mode using a 3x2 filter with unsampled luminance samples is employed. It comprises a 6-tap spatial term, four nonlinear terms, and a bias term. The 6-tap spatial term corresponds to six neighboring luminance samples (i.e., L0, L1, ..., L5) surrounding the chrominance sample (i.e., C) to be predicted. The four nonlinear terms are derived from the luminance samples L0, L1, L2, and L3, as shown below. The positions of the unsampled luminance samples are displayed. Figure 3 .

[0044] Cross-Component Residual Model (CCRM) As described in JVET-AD0108 (Pekka Astola et al., “AHG12: Cross Component Residual Model (CCRM) for Inter-Frame Prediction,” Joint Video Experts Team (JVET) ITU-T SG 16 WP3 and ISO / IEC JTC 1 / SC 29, 30th Meeting, Antalya, TR, 21-28 April 2023, document: JVET-AD0108), it applies the Cross Component Residual Model (CCRM) to predict chroma samples from reconstructed luminance samples when blocks use inter-frame prediction or intra-frame block copying (IBC). Figure 4 The method at the decoder end is described. The cross-component filter is derived using the prediction signals for luminance and chrominance. The derived filter is applied to the reconstructed luminance signal to produce the final chrominance prediction. The filter coefficients are derived in step 420 for each chrominance component using the prediction signals (i.e., predY 410 and predCb 412 or predCr 414), and the filter is applied to the reconstructed luminance signal in step 430 as follows: Figure 4 As shown. The reconstructed luminance signal is formed by combining the luminance prediction (PredY) 410 and the residual luminance signal (resY) using adder 422. After applying the filter, step 430 generates the filtered predicted Cb 440 and the filtered predicted Cr 450. The reconstructed Cb signal is formed by combining the filtered predicted Cb (i.e., filtered-predicted Cb) 440 and the residual Cb signal (i.e., resCb) using adder 442. Similarly, the reconstructed Cr signal is formed by combining the filtered predicted Cr (i.e., filtered-predicted Cr) 450 and the residual Cr signal (i.e., resCr) using adder 452.

[0045] In-screen template matching Intra-template matching prediction (IntraTMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then issues a signal to use this mode and performs the same prediction operation at the decoder.

[0046] Extended Merge Prediction In VVC, the merge candidate list is constructed by sequentially including the following five types of candidates: 1. Spatial MVP (Motion Vector Predictor) from a spatially adjacent CU (2) Time MVP from co-occurring CU 1. Historical MVP from FIFO table 1. Paired average MVP (5) Zero MV Spatial Candidate Derivation The derivation of spatial merge candidates in VVC is the same as in HEVC, except that the positions of the first two merge candidates are swapped. A maximum of four merge candidates are selected for the current CU 510 (B 0, A 0, B1 and A1 are located in Figure 5 The location depicted in the text. The order of the derived text is B. 0, A 0, B 1, A1 and B2. Position B2 is considered only if positions B0, A0, B1, and A1 of one or more neighboring CUs are unavailable (e.g., belonging to another slice or tile) or if intra-frame coding is used. After a candidate for position A0 is added, the addition of remaining candidates is limited by a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency.

[0047] In addition to the aforementioned spatial candidates, non-proximity spatial merging candidates, as described in JVET-L0399 (Yu Han et al., “CE4.4.6: Improvements to Merge / Skip Modes,” ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 of the Joint Video Exploration Team (JVET), 12th Meeting: Macau, China, October 3-12, 2018, Document: JVET-L0399), are inserted after the TMVP (Temporal Motion Vector Predictor) in the regular merging candidate list. An example of a spatial merging candidate mode is shown in… Figure 6 The distance between non-nearest spatial candidates and the current decode block is based on the width and height of the current decode block. Line buffer constraints do not apply.

[0048] Temporal Candidates Derivation In this step, only one candidate is added to the list. Specifically, when deriving this time-merging candidate for the current CU 710, a scaled motion vector is derived from the corresponding CU 720 belonging to the same reference image, such as... Figure 7 As shown. The list of reference images and reference indices used to derive the isotopic CUs are explicitly indicated in the slice header. The scaled motion vectors of the temporal merge candidates are shown in Figure 730. Figure 7 As shown by the dashed line, the motion vector 740 of the co-located CU is obtained by scaling using the POC (Picture Order Count) distance, tb, and td. The characteristic is that tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the reference image and the co-located image. The reference image index of the time merging candidate is set to zero.

[0049] The position of the time candidate is selected between candidates C0 and C1, such as... Figure 8 As shown. If the CU at position C0 is unavailable, is intra-coded, or is outside the current CTU line, then position C1 is used. Otherwise, position C0 is used when merging candidates at derived time.

[0050] History-based Merge Candidate Derivation Following the spatial MVP and TMVP, historical MVP (HMVP) merge candidates are added to the merge list. In this approach, motion information from previously coded blocks is stored in a table and used as the MVP for the current CU. A table containing multiple HMVP candidates is maintained during encoding / decoding. The table is reset (cleared) when a new CTU row is encountered. Whenever a CU that is not coded between sub-blocks is encountered, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0051] Pair-wise Average Merge Candidates Derivation A pairwise average candidate is generated by averaging predefined candidate pairs in the existing merge candidate list, using the first two merge candidates. The first merge candidate is defined as p0Cand, and the second merge candidate can be defined as p1Cand. The average motion vector is calculated for each reference list based on the availability of motion vectors for p0Cand and p1Cand. If two motion vectors are available in a list, even if they point to different reference images, the two motion vectors are averaged, with the reference image set to the one for p0Cand; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, the list remains invalid. Furthermore, if the half-pixel interpolation filter indices of p0Cand and p1Cand are different, they are set to 0.

[0052] If the merge list is not full after adding the average number of merge candidates, insert the zero MVP until the maximum number of merge candidates is reached.

[0053] Reference image resampling (RPR) In HEVC, the spatial resolution of a picture cannot be changed unless a new sequence is started with a new SPS (Sequence Parameter Set) and an IRAP (Intra Random Access Point) picture is available; this is always intra-coded. VVC allows changing the picture resolution at a point in the sequence without encoding the IRAP picture. This feature is sometimes called Reference Picture Resampling (RPR) because when the reference picture has a different resolution than the current picture being decoded, the reference picture used for inter-frame prediction needs to be resampled. To avoid additional processing steps, the RPR process in VVC is designed to be embedded in the motion compensation process and performed at the block level. During the motion compensation phase, scaling is used along with motion information to locate reference samples in the reference picture used in the interpolation process.

[0054] In VVC, the scaling factor is limited to greater than or equal to 1 / 2 (i.e., 2x downsampling from the reference image to the current image) and less than or equal to 8 (i.e., 8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling factors between the reference and current images. These three sets of resampling filters are suitable for scaling ranges from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as in the case of motion-compensated interpolation filters. It is worth noting that the filter set for normal MC interpolation is used in the scaling range from 1 / 1.25 to 8. In fact, the normal MC interpolation process is a special case of the resampling process in the scaling range from 1 / 1.25 to 8. In addition to the traditional translation block motion, the affine mode has three sets of 6-tap interpolation filters for the luma component to cover different scaling factors in RPR. The horizontal and vertical scaling ratios are derived from the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference image and the current image.

[0055] Signalling of Resolution and Cropping Window In VVC, the maximum image resolution and the corresponding conformal cropping window are signaled in the Sequence Parameter Set (SPS), while the image resolution of each current image is signaled in the Picture Parameter Set (PPS). This signaling arrangement can be used to support Reference Picture Resampling (RPR). When the image resolution of the current image (from the PPS) is less than the maximum image resolution (from the SPS), a different conformal cropping offset can be used between the signal in the PPS and the signal in the SPS.

[0056] Miscellaneous Inter-Prediction Aspects To reduce memory bandwidth, 4x4 inter-frame coding units (CUs) are not allowed in VVC. For 4x8 / 8x4 CUs used in inter-frame coding, only unidirectional mode is permitted. When the motion information in merge mode is bidirectional, it is converted to unidirectional, retaining only the motion information from list 0.

[0057] Various schemes have been revealed to improve the encoding and decoding performance of cross-component prediction.

[0058] Cross-component information is used to improve the prediction accuracy of inter-frame blocks. To improve the prediction accuracy of the chroma components of inter-frame blocks, luminance information from the corresponding luma component and / or chroma information from the current chroma block and / or chroma information from previously encoded chroma components are used.

[0059] The first approach is to improve the prediction of Cb and / or Cr by applying a cross-component model to the information from Y (currently reconstructed or predicted) for code-decode units containing luminance (Y) and chrominance (Cb and / or Cr) components (under single-tree segmentation).

[0060] The second approach is to improve the prediction of Cr by applying a cross-component model to information from Cb (currently reconstructed or predicted) for codec units containing luminance (Y) and chrominance (Cb and / or Cr) components (under single-tree segmentation) or for codec units containing chrominance (Cb and / or Cr) components (under chrominance dual-tree segmentation).

[0061] The following are several embodiments related to the first scheme for using inherited cross-component modes for the current chroma block, including the following steps: step (1) establishing a candidate list (modelList) for the current block, characterized in that the candidate list includes cross-component models; step (2) selecting one or more model information sets in the list; and step (3) using the model information (similar to intra-frame chroma cross-component modes) to generate one or more prediction hypotheses for the current chroma component (Cb or Cr) by applying and / or modifying the selected model information to the reconstructed or predicted samples of the corresponding luminance component.

[0062] When the selected model information refers to the traditional cross-component linear model, the proposed method is called the inter-CCLM (inter-cross-component linear model) mode. When the selected model information refers to the convolutional cross-component model derived through a regression-based method (such as CCCM), the proposed method is called the inter-CCCM (inter-cross-component convolutional model) mode.

[0063] Furthermore, in some embodiments, a self-derived (or re-derived) cross-component pattern is proposed that can be added to the candidate list in step (1). In some embodiments, the self-derived cross-component pattern is not added to the list, and a choice is made to use the proposed inheritance pattern and / or the proposed self-derived pattern. In some embodiments, the choice to use the proposed inheritance pattern and / or the proposed self-derived pattern is determined according to an explicit rule, an implicit rule, or both. Further details are described in the section entitled “IV. Choice to use the proposed inheritance pattern and / or the proposed self-derived pattern”.

[0064] In one embodiment, the proposed embodiment can also be used in a second scheme by using the previously encoded chroma component (Cb) as the luminance component in the first scheme.

[0065] Model storage and inheritance In another embodiment, when the current inter-frame block uses model parameters from a self-derived cross-component mode, the model parameters used can be saved and / or referenced by subsequent codec blocks. For example, when the self-derived cross-component model is CCRM, all or any subset of the model parameters can be saved. In one embodiment, if the subsequent codec block is intra-frame, the saved model parameters are allowed to be used. If the subsequent codec block is inter-frame or any mode type (e.g., IBC), the saved model parameters are allowed to be used. In another embodiment, if the subsequent codec block and the current block have different mode types (e.g., one is an inter-frame block and the other is not), the saved model parameters are not allowed to be used.

[0066] In another embodiment, when the current inter-frame block uses an inherited cross-component mode, the model parameters used can be saved and / or referenced by subsequent codec blocks. For example, with inherited CCCM, all or any subset of model parameters can be saved. In one embodiment, the saved model parameters are allowed if the subsequent codec block is intra-frame. The saved model parameters are allowed if the subsequent codec block is inter-frame or any mode type (e.g., IBC). In another embodiment, the saved model parameters are not allowed if the subsequent codec block has a different mode type (e.g., not an inter-frame block).

[0067] In another embodiment, when the current inter-frame block uses any cross-component model (e.g., an inherited cross-component model, a self-derived cross-component model, a cross-component model used in chroma fusion, meaning that chroma prediction is based on adding one or more cross-component prediction assumptions to one or more existing non-cross-component prediction assumptions, or any combination thereof), the model parameters used can be saved and / or referenced by subsequent codec blocks. For example, with inherited CCCM, all or any subset of model parameters can be saved. In one embodiment, the saved model parameters are allowed if the subsequent codec block is intra-frame. The saved model parameters are allowed if the subsequent codec block is inter-frame or any mode type (e.g., IBC). In another embodiment, the saved model parameters are not allowed if the subsequent codec block has a different mode type (e.g., not an inter-frame block).

[0068] I. Establish a candidate list containing cross-component models In one embodiment, when building the merge class candidate model list (modelList), it includes a collection of one or more candidate model information items. For each candidate in the list, it refers to a candidate model information item. The definition of the model information can be found in the section titled "V.1. Inherited CCM Information".

[0069] Spatial model information from spatially neighboring blocks (corresponding to "spatial MVP from spatially neighboring CUs" between frames) Temporal model information from co-located blocks (corresponding to the "temporal MVP from co-located CUs" between frames) Historical model information from the FIFO table (corresponding to "historical MVP from the FIFO table" between frames) Pairwise average model information (corresponding to the "pairwise average MVP" between frames) Default model information (corresponding to "zero MV" between frames) In one sub-implementation of the candidate type list described above, where the candidate type is "spatial model information from spatially neighboring blocks," a valid spatially neighboring block can come from one of spatially adjacent and non-adjacent neighbors (or blocks from any subset of the current block's neighborhood search region) that satisfies a predefined condition. For example, for a non-adjacent neighbor, the predefined condition (e.g., a validity / availability check) refers to the non-adjacent neighbor being within the available region of the non-adjacent spatial candidate. For instance, the predefined condition is that the neighbor is encoded by or combined with a cross-component mode. A cross-component mode refers to modes such as CCLM, MMLM, CCCM, GLM, modes that inherit mode information from the merge class candidate list, MH CCLM, and / or any cross-component mode that belongs to a cross-component branch (containing many cross-component modes) and does not belong to a traditional intra-prediction mode. Combining with a cross-component mode refers to modes such as chroma fusion (or Angular / Planar modes named LM-assisted), inter-frame CCLM, inter-frame CCCM, and / or any traditional mode whose syntax does not belong to a cross-component branch but uses cross-component information to generate predictions. In another sub-implementation, when checking the validity of neighboring code-decode blocks, a second round of validity checks is further used when the aforementioned validity checks (e.g., the neighboring block is not in a cross-component mode or the neighboring block does not use / combines with a cross-component mode) are met. The motion vectors and / or block vectors of the neighboring blocks can be used to find the cross-component model. How to use motion vectors and / or block vectors to find model variations can be referred to the description of "Temporal Model Information from Co-location Blocks" in the candidate type list above. If a model is found, the second round of validity checks for the neighboring block is satisfied, and the found model can be inserted into the list; otherwise, the neighboring block is invalid for insertion. When scanning spatial neighboring blocks, if a candidate is valid, the candidate is added to the list.

[0070] In another sub-implementation where the candidate type is "temporal model information from a co-block," in a first case, the co-block comes from a block in a reference image or a predefined co-block image, as an inter-frame pattern, using the current block position and / or the current block motion; and / or in a second case, the co-block comes from a block in a reference image or a predefined co-block image, as an inter-frame pattern, using the current block position and / or the motion of neighboring blocks. In the first case, for example, when the current block is encoded by an inter-frame prediction pattern, the co-block is referenced by the motion information of the current block (containing motion vectors and a reference image indicated by a reference index). If the current block is a sub-block motion pattern (e.g., an affine pattern), each sub-block in the current block has its own co-temporal model information. Co-temporal information from all or any subset, and co-temporal model information referenced by different sub-block motions (each sub-block), is added to a list. Another example is when the reference image indicated by the reference index differs from a predefined co-op image. This co-op image can be used for inter-frame motion vector prediction or any co-op image specified in the standard to maintain motion or cross-component model information stored and available for the current block. Temporal information from the reference image is disabled. In another example, when the reference image indicated by the reference index differs from a predefined co-op image, this co-op image can be used for inter-frame motion vector prediction or any co-op image specified in the standard to maintain motion or cross-component model information stored and available for the current block. Motion vectors are scaled to reference the predefined co-op image, and the scaled motion vectors are used to find co-op blocks in the co-op image to obtain the cross-component model in the co-op block. The scaling process is shown in the "Inheriting Temporal Proximity Model Parameters" and "Temporal Candidate Derivation" sections. For the second case, some examples are described. For example, temporal model information can come from the motion information of a neighboring block referenced by the current block. Similar to the first case, the disabling method or scaling method can be used in the second case. If the proposed method is applied to IBC blocks or any pattern using block vectors (in the first case, the current block is an IBC; in the second case, neighboring blocks are IBCs), the block vector information is used as motion vectors. The block vector information is determined by signal and / or template matching within a predefined search range, such as intraTMP and / or any implicit or explicit predefined rules. Further details can be found in the section on "Inheriting Temporal Proximity Model Parameters".

[0071] In another sub-implementation where the candidate type is "history-based model information," a history-based table (FIFO table) is established to store model information from previous encoded blocks. This table can be reset at the beginning and / or end of a CTU, slice, image, tile, and / or sequence. One or more history-based candidates can be added to the candidate list in either head-to-tail or tail-to-head order.

[0072] In another sub-implementation where the candidate type is "pairwise averaged model information," the model information for this candidate is derived based on the model information of multiple previous candidates in the list. For example, it can average and / or modify the model parameters of multiple candidates as the model parameters to be applied. Another example is that it can combine multiple predictions as the final prediction, characterized in that each prediction is generated by applying a model from the candidate list.

[0073] In another sub-implementation, if the list is not full after inserting all predefined candidates, default model information is added. For example, the default model could be a CCLM model. The default alpha (or...) The value (a, or scaling parameter) is selected from {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, …}, while beta (or beta) is selected from {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, …}. (b, or offset parameter) is based on the selected default alpha, the average neighbor reconstructed luminance sample value, and the average neighbor reconstructed chromaticity (Cb / Cr) sample value.

[0074] In another sub-implementation, the candidate list for inter-frame chroma blocks is unified with the candidate list for intra-frame chroma blocks, and / or may further include inter-frame specific candidates (e.g., temporal model information referenced by the current motion) and / or may be any subset of the candidate list for intra-frame chroma blocks based on the candidate list for intra-frame chroma blocks.

[0075] In another embodiment, when constructing `modelList`, one or more self-derived cross-component candidates are included. These self-derived cross-component candidates are described in a section titled "Self-derived Cross-component Model". In another sub-implementation, self-derived cross-component candidates are added only if the list does not contain enough inherited candidates. For example, a self-derived candidate is added before or considered the default candidate. In another sub-implementation, self-derived cross-component candidates are added at any predefined position in `modelList`. For example, a position after spatially adjacent candidates. Another example is a position after spatially non-adjacent candidates. Yet another example is a position after all or a subset of any temporal-domain candidates.

[0076] After the list is constructed, in one embodiment, the list is reordered according to the method defined in the "Reorder Candidates in the List" section.

[0077] II. Enabled or disabled signals and one or more model information from the selection list (if enabled). In this section, the term "inter-frame CCLM" refers to "inter-frame CCLM or inter-frame CCCM".

[0078] When the suggested inter-frame CCLM (or inter-frame CCCM) is not applied, the prediction for the current block comes from the original inter-frame prediction.

[0079] In another embodiment, the choice to apply inter-frame CCLM or not depends on the signal.

[0080] In one sub-implementation, the signal refers to the encoded TU / TB / CU / CB level flags. These flags may or may not depend on the context to be encoded. For example, the TU / TB flag is emitted only when the TU / TB's luminance Cbf is non-zero and the inter-frame mode enable flag is true. Similarly, the CU / CB flag is emitted only when the CU / CB's luminance Cbf is non-zero and the inter-frame mode enable flag is true. The inter-frame mode enable flag means that when all proposed inter-frame CCLMs (or inter-frame CCCMs) support all inter-frame modes, the CU's predMode is MODE_INTER. When IBC-supported proposed inter-frame CCLMs (or inter-frame CCCMs) support IBCs, the IBC enable flag is checked first, and the inter-frame CCLM (or inter-frame CCCM) signal is encoded and decoded based on the CU's predMode being MODE_IBC. When only CIIP-supported proposed inter-frame CCLM (or inter-frame CCCM) is supported, the CIIP enable flag is checked first, and the inter-frame CCLM (or inter-frame CCCM) signal is encoded if the CIIP flag is true. When only merged proposed inter-frame CCLM (or inter-frame CCCM) is supported, the merge flag is checked first, and the inter-frame CCLM (or inter-frame CCCM) signal is encoded if the merge flag is true. When only AMVP-supported proposed inter-frame CCLM (or inter-frame CCCM) is supported, the merge flag is checked first, and the inter-frame CCLM (or inter-frame CCCM) signal is encoded if the merge flag is false. The proposed inter-frame CCLM (or inter-frame CCCM) can only support any predefined subset of merged modes, any predefined subset of inter-frame modes, or any predefined subset of non-intra-frame modes.

[0081] In another sub-implementation, when the signal indicates the application of inter-frame CCLM (or inter-frame CCCM), an additional signal is used to select one or more models from the total candidates. The candidate index is referred to as modelIdx in this disclosure. If the modelList containing the total candidates (e.g., candidates as described in the "Constructing a Candidate List Containing Cross-Component Models" section, CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T) or any subset of candidates is reordered according to the method described in the "Reordering Candidates in the List" section, the additional signal specifies the candidate index in the reordered list. For example, if an LM mode is selected, the LM prediction is generated from the selected LM. As another example, if multiple LM modes are selected, the LM prediction is generated by a mixture of prediction hypotheses from the multiple LM modes.

[0082] In another sub-implementation, no additional signal is required, and one or more models are selected based on implicit rules. For example, one or more selected models are implicitly determined, or one or more models for the current block are determined without emitting `modelIdx`. For instance, the first candidate in the list is used. If the list is reordered by template cost, the first candidate is the one with the lowest template cost.

[0083] In another embodiment, the original inter-frame prediction (generated by motion compensation) is used for luminance, and the prediction of the chrominance component is generated by CCLM and / or any other LM mode.

[0084] In one sub-implementation, the current CU is considered an inter-frame CU, an intra-frame CU, or a novel prediction mode (i.e., neither intra-frame nor inter-frame).

[0085] In another embodiment, one or more LM modes (i.e., cross-component modes) are used to generate one or more hypothesis predictions to assist Angular / Planar modes / inter-frame CCLM / inter-frame CCCM / MH CCLM, which are selected from a predefined merge candidate list (i.e., modelList). A modelIdx is signaled to select a candidate from the candidate list (modelList), and the selected candidate is used for the current block. modelList contains one or more candidates, each pointing to a model (or cross-component mode) information. If there is only one candidate in the list (i.e., the list size is 1), modelIdx is not signaled and / or can be inferred to be 0 or a default value. In one embodiment, modelIdx is implicitly determined, or one or more models for the current block are determined without signaling modelIdx. For example, the first candidate in the list is used. If the list is reordered by template cost, the first candidate is the one with the lowest template cost. Another example is that the candidate / model used is implicitly selected from the list using a predefined rule that depends on the block's codec information for the candidate to be used. This embodiment is referred to as "noteA".

[0086] In one embodiment, when modelList is built, one or more predefined candidates are added. The predefined candidates may include any subset / expansion of the following candidates and / or more candidates as described in the embodiment in “note A”.

[0087] CCLM family: CCLM_LT, CCLM_L, CCLM_T MMLM family: MMLM_LT, MMLM_L, MMLM_T CCCM family: CCCM_LT, CCCM_L, CCCM_T The method described above can also be applied to IBC blocks or blocks of any IBC submode (e.g., IBC merging, IBCAMVP, or any IBC mode under an IBC syntax). In this application, the term "inter-frame" can be replaced with IBC. That is, for chroma components, block vector prediction can be combined with or replaced by cross component prediction.

[0088] III. Use model information to generate one or more hypothetical predictions for the current chromaticity components. III.1. Concept In one embodiment, the prediction or reconstruction is based on a hypothetical prediction of the current chromaticity component, which is used to generate the model.

[0089] In one sub-example of prediction based on a linear model, the derived model parameters are applied to the prediction samples of the first component (Y) to obtain the prediction samples of the second or third component.

[0090] The predicted samples of the first component are downsampled by downsampling filters, which may be fixed to a predefined filter or selected from a number of candidate filters.

[0091] In a sub-implementation of reconstruction based on a linear model, the derived model parameters are applied to the reconstructed samples of the first component (Y) to obtain predicted samples of the second or third component.

[0092] The reconstructed samples of the first component are downsampled using downsampling filters, which may be fixed to a predefined filter or selected from a number of candidate filters.

[0093] The proposed methods for prediction or reconstruction based on convolutional models are similar to those based on linear models. The main difference is that the model coefficient patterns follow CCCM (instead of CCLM), and the brightness samples may or may not be downsampled first.

[0094] In another embodiment, cross-component predictions of multiple hypotheses (MHs) are blended, or multiple models are used to generate a single hypothesis prediction for the current block. A multi-hypothesis CCLM is proposed to blend predictions from multiple CCLM methods. The term "CCLM method" can refer to all cross-component modes. The CCLM methods to be blended can be from (but are not limited to) the CCLM methods mentioned above (e.g., CCLM, MMLM, CCCM, GLM, CCRM, …) and / or models defined in the embodiments described in “note A”. A weighting scheme is used for blending.

[0095] III.2. CCLM of Inter-Frame Blocks The term "CCLM of inter-block" can also be named "inter-CCLM", and "CCLM" can be extended to any LM mode (or any cross-component mode) or replaced by any LM mode (or any cross-component mode). When a convolutional cross-component model derived by a regression-based method is used, CCLM of inter-block can also be named inter-CCCM.

[0096] In one embodiment, for the chroma component, in addition to the original inter-frame prediction (generated by motion compensation, which may be a single prediction and / or a bidirectional prediction, multiple hypothetical predictions from multiple motion candidates, which may refer to one or more merged candidates, one or more AMVP candidates, any combination of the above, or may just be a single prediction), one or more hypothetical predictions (generated by CCLM and / or any other LM mode) are used to generate the current prediction.

[0097] In one sub-implementation, the current prediction is a weighted sum of inter-frame prediction and CCLM prediction.

[0098] In another embodiment, inter-frame prediction can be generated by any of the inter-frame modes described above. For example, the inter-frame mode can be a regular merging mode. Another example is that the inter-frame mode can be a CIIP mode. Yet another example is that the inter-frame mode can be GPM or any GPM variant (e.g., GPM intra-frame references a prediction unit using intra-frame prediction).

[0099] In another embodiment, inter-frame CCLM is supported only when the current block is compiled and decoded using one or more predefined inter-frame modes, or when one or more enable flags for enabling predefined inter-frame modes are indicated as enabled. The significance of supporting inter-frame CCLM is that prediction of the current block can be selected between applying inter-frame CCLM and not applying inter-frame CCLM.

[0100] When inter-frame CCLM is applied, the prediction for the current block is generated in the following way: In one sub-implementation: one or more prediction hypotheses (generated by CCLM and / or any other LM mode) are combined with the original inter-frame predictions. Hybrid chroma prediction from existing inter-frame modes and prediction from LM Mix: Predfinal = ( wInter * PredInter + wLM * PredLM + 2 )>>2 Weighting rules: wInter and wLM, for example: If the top and left sides are both intra-frame (or any cross-component mode), (wInter, wLM) = (1, 3) Otherwise, if the top and left sides are both within the frame, (wInter, wLM) = (2, 2). Otherwise, (wInter, wLM) = (3, 1) Another example is that the weights follow the CIIP weighting rules.

[0101] For example, predInter = inter-frame prediction after OBMC (if OBMC is used). Another example is predInter, where inter-frame prediction occurs before OBMC (OBMC can be applied after mixing). In another sub-implementation: the original inter-frame prediction is replaced with one or more prediction hypotheses (generated by CCLM and / or any other cross-component mode). Another example: if the CCLM model is used to generate chroma prediction samples, while luma prediction comes from inter-frame codec tools, a flag is used to indicate whether the CCLM model used for chroma prediction is inherited from the CCLM model used in a previous coding block or from a predefined CCLM model. If the CCLM model is inherited from the CCLM model used in a previous coding block, an index is used to indicate which model in the list was inherited or modified. Otherwise, a predefined CCLM model is used to implicitly derive the CCLM model for the current chroma prediction.

[0102] IV. Choose to use the proposed inheritance mode and / or the proposed self-export mode. In one embodiment, a flag can be issued to indicate / select whether to use the re-exported model. If the flag is 0, the cross-component model used to encode neighbor merge candidate blocks is inherited. If the flag is 1, the re-exported method is used.

[0103] In another embodiment, an implicit rule (without using additional flags) is used to determine whether to use the re-exported model.

[0104] In another embodiment, if no model can be inherited during the creation of modelList, or if spatially adjacent / non-adjacent candidate blocks, historical candidate blocks, temporal candidate blocks, or candidate blocks of all or any subset mentioned in this application (e.g., before the default candidate block) are unavailable, then a re-exported model is used.

[0105] In another embodiment, when using the proposed inheritance method, the candidate block with the lowest cost (e.g., the first candidate block in modelList) is implicitly selected to generate cross-component predictions. Another example is issuing an index to select one or more candidate blocks from modelList. More details can be found in Section 2.

[0106] V. Details of the cross component patterns (including Model information) in the candidate list V.1. Inheriting CCM Information In one embodiment, inherited Cross Component Model (CCM) information can be stored along with inherited model parameters. CCM information can be inherited along with the inherited model parameters. Prediction for the current block can be generated based on the inherited CCM information and the inherited model parameters. CCM information may include, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM, 2-parameter GLM, 3-parameter GLM (GLM model with a luminance term)), model indexes indicating which model shape to use in the convolutional model, classification thresholds for multiple models, information indicating the use of non-downsampled samples in the convolutional model, downsampling filter flags (whether downsampling is performed), downsampling filter indexes when using multiple downsampling filters, the number of neighboring lines used to derive the model, the template type used to derive the model, post-filter flags, and model parameters.

[0107] In one embodiment, hybrid cross-component models (CCCMs) consisting of various items (e.g., spatial, gradient, positional, nonlinear, and bias terms) can be inherited. In addition to storing model parameters, the prediction pattern can be stored in the cross-component model information to indicate that the inherited model is a hybrid CCCM model composed of various items. If there are multiple types of hybrid CCCM models, model indexes can also be stored in the cross-component model information to indicate which type of hybrid CCCM model is being inherited. For example, the Gradient and Location-based Cross-Component Model (GL-CCCM) proposed in JVET-AB0119 (Ramin G. Youvalari et al., “Non-EE2: Gradient and location-based convolutional cross-component model (GL-CCCM) for intra-frame prediction,” ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 28th meeting, Mainz, DE, October 20-28, 2022, document: JVET-AB0119) is a hybrid CCCM model. It consists of a spatial term for the center location, two gradient terms for the horizontal and vertical directions, two positional terms for the relative horizontal and vertical locations, a nonlinear term, and a bias term. The prediction pattern can be stored in the cross-component model information to indicate that the inherited model is a GL-CCCM model.

[0108] V.2. Inheritance Spatial Proximity Model Parameters In one embodiment, inherited model parameters can come from a directly neighboring block. Models from blocks at predefined locations are added to the candidate list in a predefined order. This predefined order can be any possible order of spatially neighboring blocks.

[0109] In one embodiment, the predefined positions and predefined order can be the same as the spatial candidates for the inter-frame merging mode.

[0110] In one embodiment, the predefined location can be Figure 5 The locations depicted (as described in the "Spatial Candidate Derivation" section). The predefined order can be B0, A0, B1, A1, and B2.

[0111] In one embodiment, assuming the current block's position, width, and height are (x, y), W, and H, respectively, the predefined position can include the position directly above the current block, such as (x + W>>1, y-1) or (x + (W+1)>>1, y-1), if W is greater than or equal to a threshold TH. The predefined position can also include the position to the left of the current block, such as (x-1, y+H>>1) or (x-1, y+(H+1)>>1), if H is greater than or equal to the threshold TH. TH can be 2, 4, 8, 16, 32, or 64. The predefined position includes the position directly above (W>>1) or ((W>>1)–1) if W is greater than or equal to TH, and the position directly to the left (H>>1) or ((H>>1)–1) if H is greater than or equal to TH.

[0112] In one embodiment, the number of models inherited from spatial neighboring blocks has a maximum value, and the maximum number that can be added to the candidate list is less than the number of predefined locations.

[0113] V.3. Inheritance Time Proximity Model Parameters In one embodiment, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded slices / images.

[0114] In one embodiment, the current block is located at (x, y), and the block size is... The inherited model parameters can come from blocks at some predefined locations in previously encoded slices / images.

[0115] In one sub-implementation, the predefined position can be the same as the predefined position of the time candidate in the inter-frame merge mode.

[0116] In one sub-implementation, the predefined location can be or Its characteristics Two value sets and Defined as: , and All values ​​in the range are positive.

[0117] For example, It can be .

[0118] For example, .For example, .

[0119] For example, .For example, and In one sub-implementation, a predefined location Located within the corresponding region of the current coded block, i.e. and Predefined positions can be .

[0120] In one sub-implementation, a predefined location Located outside the corresponding region of the current coding block, i.e. ,and Predefined positions can be .

[0121] In one embodiment, from closer The model at the location is first added to the final merge candidate list.

[0122] The previously encoded image, from which the inherited parametric model is obtained, is hereafter called a parsing image.

[0123] In one embodiment, the source of the inherited parameter model (i.e., the corresponding image) is one of the images in the reference list.

[0124] In one embodiment, the co-position image can be the same as the co-position image in the inter-frame merge mode.

[0125] In one embodiment, the corresponding image is marked in the image / slice header. A reference list and reference index are also marked in the image / slice header. For example, the corresponding image is selected as L0[0]. In another example, the corresponding image is selected as L1[0].

[0126] In one embodiment, the corresponding image is selected as the image in the reference list that has the smallest POC difference between the corresponding image and the current image.

[0127] In another sub-implementation, the corresponding image is selected as the image with a larger or smaller QP from the reference list.

[0128] In one embodiment, if an image in the reference list is rescaled (i.e., RprConstraintsActiveFlag of the co-op image is true), that image is not selected as the co-op image. A rescaled reference image means that the co-op image differs from the current image in one or more of the following seven parameters: 1) the image width in the luminance samples (pps_pic_width_in_luma_samples), 2) the image height in the luminance samples (pps_pic_height_in_luma_samples), 3) the left offset of the scaling window (pps_scaling_win_left_offset), 4) the right offset of the scaling window (pps_scaling_win_right_offset), 5) the top offset of the scaling window (pps_scaling_win_top_offset), 6) the bottom offset of the scaling window (pps_scaling_win_bottom_offset), and 7) the number of subpics - 1 (sps_num_subpics_minus1).

[0129] In one embodiment, the rules for selecting / not selecting co-op images can be combined. For example, co-op images can be selected from a reference list of images that have not been rescaled. The co-op image is selected as the one with the smallest difference in Proof of Concept (POC) between it and the current image.

[0130] In one embodiment, the inheritance of cross-component model information from the co-image is disabled when the co-image is rescaled.

[0131] In one embodiment, when the co-located image is rescaled, the source position of the inherited model can be scaled according to the scaling ratio. The scaling ratio is derived based on the scaling windows of the current image and the co-located image. Let the position be (x, y), the scaling position be (x', y'), and the scaling ratio be R. The scaling position can be (x / R, y / R) or (x / R, y / R) after rounding. The rounding method used can be, but is not limited to, the following: rounding to negative infinity, rounding to positive infinity, rounding to zero, or rounding to the nearest integer (e.g., rounding away from zero, rounding halfway up, rounding halfway down, ...).

[0132] In one embodiment, the position in the previously encoded slice / image, inherited from the source of the parametric model, is determined by the motion vectors of neighboring blocks. Let Δx and Δy be the horizontal and vertical displacements determined based on the selected neighboring block motion vectors, the current block position is (x, y), and the block size is... The inherited model parameters can come from a block at position (x', y'), characterized in that x' = x + Δx and y' = y + Δy, or x' = x + w / 2 + Δx and y' = y + h / 2 + Δy.

[0133] In one embodiment, when selecting neighboring blocks, there may be a predefined list of locations. The locations in the list are checked in a predefined order. For each location, the L0 motion vector is checked first, followed by the L1 motion vector. Another example is that the L1 motion vector is checked first, followed by the L0 motion vector. The selected motion vector is the first one whose reference image has not been rescaled.

[0134] V.4. Inheriting the proximity model of non-proximity spaces In one embodiment, the inherited model parameters can come from neighboring blocks in a non-neighboring space. Models from blocks at predefined locations are added to the candidate list in a predefined order.

[0135] In one sub-implementation, the predefined positions and predefined order are the same as those used for non-nearest spatial neighbor candidates in the inter-frame merging mode.

[0136] In one sub-implementation, the predefined positions and predefined orders are as follows: Figure 9A and Figure 9B As shown. The positions of the numbered squares are predefined. The number within each square represents a predefined order. Mode 1 ( Figure 9A The position in mode 2 () Figure 9B The positions in the list are added before the current block. The distance between each predefined position is proportional to the width and height of the current block.

[0137] In one embodiment, the maximum number of inheritance models from non-nearby spatial neighbors that can be added to the candidate list is less than the number of predefined locations.

[0138] V.5. Inheriting Model Parameters from History Tables In one embodiment, the inherited model parameters can come from a cross-component model history table. The history table stores CCM information for valid previously encoded blocks. A valid previously encoded block refers to any block containing valid CCM information. Cross-component models in the history table can be added to the candidate list in a predefined order. In one embodiment, the order in which historical candidates are added can be from the beginning to the end of the table. In another embodiment, the order in which historical candidates are added can be from the end to the beginning of the table.

[0139] In one embodiment, a cross-component model history table can be maintained to store previous cross-component models (i.e., CCM information), and the cross-component model history table can be reset at the beginning of the current image, current tile, current tile, each MCTU row, or each N CTU, where N and M can be any values ​​greater than 0. In another embodiment, the cross-component model history table can be reset at the end of the current image, current tile, current tile, current CTU row, or current CTU.

[0140] In another embodiment, multiple history tables are used to store different types of cross-component models. For example, the first history table stores a single model, and the second history table stores multiple models. Another example is that the first history table stores gradient models, and the second history table stores non-gradient models. Yet another example is that the first history table stores simple linear models (e.g., y = ax + b), and the second history table stores complex models (e.g., CCCM).

[0141] In one embodiment, when adding historical candidates to a candidate list from multiple historical tables, the addition order can be from the beginning to the end of a table, and then the next historical table can be added in the same or reverse order.

[0142] V.6. Vector propagation of CCM information In one embodiment, after encoding / decoding a block, the cross-component model (CCM) information of the current block is exported and stored in the current block. The stored CCM information can be referenced by subsequent encoding / decoding blocks. Subsequent encoding / decoding blocks can inherit the CCM information of the current block. The definition of the CCM information is in the "Inherit CCM Information" section. The stored CCM information can be, but is not limited to, the following types of candidates: spatial candidates (e.g., in the "Inherit Spatial Proximity Model Parameters" section), non-proximity candidates (e.g., in the "Inherit Non-Proximity Spatial Proximity Model" section), temporal candidates (e.g., in the "Inherit Temporal Proximity Model Parameters" section), and historical candidates (e.g., in the "Inherit Model Parameters from History Table" section).

[0143] In one embodiment, if the current block has not been cross-component prediction (CCP) coded and there are available motion vectors in the current block (e.g., the current luma block is coded across frames), the CCM information of the current block can be derived by copying the cross-component model (CCM) information of a reference block located in a reference image and positioned by the motion vectors of the current block. For example, as Figure 10 As shown, block B is not CCP-encoded and has available motion vectors. Reference block A is located by motion vectors. The CCM information of reference block A, using the cross-component model, is copied and stored in block B. In one embodiment, if a reference block located by motion vectors is also not CCP-encoded, but stores CCM information, the CCM information of the current block can be derived by copying the CCM information stored in the reference block. That is, even if the reference block is not CCP-encoded, as long as it has valid stored CCM information, the stored CCM information can be referenced by the current block. For example, as... Figure 10 As shown, current block C has available motion vectors, and its reference block B is not CCP-encoded but stores CCM information. The CCM information of block B is copied and stored in block C. Since the CCM information stored in block B is copied from block A, the CCM information stored in block C originally came from block A (i.e., the CCM information of block A was propagated to block C). Block C can retrieve the CCM information originally from block A simply by accessing block B. In one embodiment, if the reference block located by the motion vector is not CCP-encoded and does not store CCM information, then CCM information is not stored for the current block.

[0144] In one embodiment, when the current block is coded across frames using bidirectional prediction, in order to derive the CCM information of the current block, if only one reference block located by motion vectors has CCM information, then the CCM information of the reference block with CCM information is copied and stored in the current block. For example, as... Figure 10 As shown, assume block F is coded across frames using bidirectional prediction. The two reference blocks located by motion vectors are block G and block H. Block G stores CCM information, while block H does not. The CCM information of block G is copied and stored in block F.

[0145] In another embodiment, when the current block is coded across frames in a bidirectional prediction manner, and both reference blocks located by motion vectors store CCM information, the CCM information of the current block is derived by combining all or part of the CCM models of its reference blocks.

[0146] In another embodiment, when the current block is coded across frames using bidirectional prediction, and both reference blocks located by motion vectors store CCM information, a reference block is selected based on a set of predefined rules. The CCM information of the selected reference block is then copied and stored in the current block.

[0147] For a sub-implementation, a reference block encoded with CCP is selected.

[0148] For a sub-implementation, an intra-frame coded reference block is selected.

[0149] For one sub-implementation, a reference block that has been cross-frame encoded is selected.

[0150] For a sub-implementation, a reference block whose reference image (i.e., the image where the reference block is located) has a smaller distance from the current image's image order count (POC) is selected.

[0151] For a sub-implementation, a reference block with a smaller QP difference between its reference image and the current image is selected.

[0152] For one sub-implementation, a reference block with a smaller QP value in its reference image is selected. For another sub-implementation, a reference block with a larger QP value in its reference image is selected.

[0153] For one sub-implementation, a reference block indicated by the L0 motion vector is selected. For one sub-implementation, a reference block indicated by the L1 motion vector is selected. For a sub-implementation, the previously described rules can be used in combination, and it is not necessary to apply all of the previously described rules. For example, a CCP-encoded reference block is selected. If both blocks are CCP-encoded, the block with the smaller POC distance between its reference image and the current image is selected. If both blocks are CCP-encoded and have the same POC distance to the current image, the reference block with the smaller QP difference between its reference image and the current image is selected. If both blocks are CCP-encoded, have the same POC distance to the current image, and have the same QP difference, the reference block with the smaller QP value of its reference image is selected. Another example is to select the block with the smaller POC distance between its reference image and the current image. If the POC distances of the two blocks are the same, the reference block with the smaller QP difference between its reference image and the current image is selected. If the POC distances of the two blocks are the same and the QP differences are the same, the reference block with the smaller QP value of its reference image is selected.

[0154] In one embodiment, when a reference image located by motion vectors is rescaled (i.e., the reference image's RprConstraintsActiveFlag is true), it is considered that the motion vectors cannot locate CCM information. Therefore, CCM information is not retrieved and stored. Rescaling the reference image means that the reference image differs from the current image in one or more of the following seven parameters: 1) image width in the luminance samples (pps_pic_width_in_luma_samples), 2) image height in the luminance samples (pps_pic_height_in_luma_samples), 3) left offset of the scaling window (pps_scaling_win_left_offset), 4) right offset of the scaling window (pps_scaling_win_right_offset), 5) top offset of the scaling window (pps_scaling_win_top_offset), 6) bottom offset of the scaling window (pps_scaling_win_bottom_offset), and 7) number of subpics - 1 (sps_num_subpics_minus1).

[0155] In one embodiment, when a reference image located by motion vectors is rescaled, the position of the reference block can be scaled according to the scaling ratio. The scaling ratio is derived based on the scaling windows of the current image and the corresponding image. Let the position of the reference block be (x, y), the scaled position of the reference block be (x', y'), and the scaling ratio be R. The scaled position can be (x / R, y / R) or rounded to (x / R, y / R). The rounding method used can be, but is not limited to, the following: rounding toward negative infinity, rounding toward positive infinity, rounding toward zero, or rounding to the nearest integer (e.g., rounding away from zero, rounding to the nearest integer, rounding to the nearest integer, etc.).

[0156] In one embodiment, the position located by the motion vector of the co-location block 1130 must be in the co-location CTU (Coding Tree Unit) row 1120 in the reference image 1110. For example... Figure 11As shown, if the position 1140 located by the motion vector is above the co-located CTU row, the position is mapped to the corresponding position (marked as (Xm, Y1)) on the top line of the co-located CTU row. If the position 1142 located by the motion vector is below the current CTU row, the position is mapped to the corresponding position (marked as (Xm, Y2)) on the bottom line of the co-located CTU row. Then, the CCM information of the mapped position is copied and stored in the current block. Assume that the minimum and maximum vertical positions of the current CTU row are Y1 and Y2 respectively. Assume that the position located by the motion vector is (Xm, Ym). If Xm < Y1, the position is changed to (Xm, Y1). The CCM information of the position (Xm, Y1) is copied to and stored in the current block. If Ym > Y2, the position is changed to (Xm, Y2). The CCM information of the position (Xm, Y2) is copied to and stored in the current block.

[0157] V.7. Inheritance from the merge mode The merge mode refers to the mode of merging two predictions to generate a final prediction. In the chroma intra-frame merge mode, the chroma intra-frame prediction generated without using the cross-component prediction (CCP) codec tools (e.g., CCLM, MMLM, CCCM) is merged with another chroma intra-frame prediction generated using the cross-component prediction codec tools. For example, the intra-frame prediction encoded without CCLM and the intra-frame prediction encoded with CCLM are merged together to obtain the final intra-frame prediction.

[0158] In one embodiment, when inheriting the cross-component model parameters from a block / position encoded in the chroma intra-frame merge mode, the model parameters used to obtain the intra-frame prediction encoded with CCP are inherited and further refined.

[0159] In one embodiment, in addition to inheriting and refining the CCP model parameters, the merge weights and codec modes of the intra-frame prediction encoded without CCP are also inherited. That is, the chroma intra-frame merge mode is inherited.

[0160] VI. Constructing the candidate list VI.1. Reordering candidates in the list The candidates in the list can be reordered to reduce the syntax overhead when signaling the candidate indices for signal selection, or to bypass the syntax of signaling the candidate indices for signal selection by selecting one or more candidates using implicit rules.

[0161] In one embodiment, the reordering rules can depend on the codec information or model error of neighboring blocks. For example, if the neighboring block above or to the left is encoded with MMLM, the MMLM candidates in the list can be moved to the head of the current list.

[0162] In one embodiment, the reordering rule is based on applying a candidate model to the neighboring templates of the current block and then comparing the error with the reconstructed samples of the neighboring templates.

[0163] VII. Self-derived cross-component model In one embodiment, an example of a self-derived cross-component model is CCRM. When performing self-derived modeling, the model (filter shape / mode, parameter terms) is unified with the cross-component model in a regular intra-frame mode. For example, the CCRM model can be unified with any predefined existing intra-frame cross-component model (e.g., CCCM, GLM, MMLM using non-downsampled luma samples), and / or self-derived simply means that the input to the derived model parameters comes from the current chroma and iso-luma samples (e.g., motion-compensated results if the current block is interleaved).

[0164] In another embodiment, self-derived cross-component candidates refer to one or more models, which are used to generate cross-component predictions for the current block in the following manner. The cross-component predictions for the current block (used to generate target prediction samples) are formed by combining one or more proposed source terms and models (referring to the proposed weight settings). As shown in Equation (3), pred(i, j) is the target (predicted) sample in the current block, which can be obtained after our proposed mechanism, sourceTermSet0 includes one or more source terms from the luma component, sourceTermSet1 includes one or more source terms from the chroma component, and biasTermSet includes one or more bias terms.

[0165] Equation (3) is merely an example; our proposed mechanism can use any subset or extension of sourceTermSet0, sourceTermSet1, and biasTermSet. Each sample or sample of any subset in the current block obtains its target (predicted) sample according to Equation (3). The contents of sourceTermSet0 are described below in Section VII.1, “Contents of sourceTermSet0(i, j)”, the contents of sourceTermSet1 are described in Section VII.2, “Contents of sourceTermSet1(i, j)”, the contents of biasTermSet are described in Section VII.3, “Contents of biasTermSet”, and the predictor derivation using the proposed source terms and proposed weight settings is described in Section VII.4, “Predictor Derivation of Sample (i, j)”. Several examples of our proposed mechanism are shown in Section VII.4, “Predictor Derivation of Sample (i, j)”.

[0166] pred(i, j) = (sourceTermSet0(i, j) + sourceTermSet1(i, j) + … + biasTermSet) uses the proposed weight setting, characterized in that (i, j) is the sample position in the current block. (3) VII.1. Contents of sourceTermSet0(i, j) `sourceTermSet0(i, j)` includes one or more luminance source terms denoted as `sourceTerm00`, `sourceTerm01`, ..., and / or `sourceTerm0n-1`. The value of `n` represents the number of points in the source term set.

[0167] In one embodiment, the source term may be a linear term and / or a nonlinear term, a linear term only and / or a nonlinear term only.

[0168] In another embodiment, n is a predefined value, such as 1, 2, ... or any positive integer. For example, the predefined value is fixed in a standard.

[0169] In another embodiment, n is determined by the encoding / decoding information of the current block and / or the sample position (i, j). For example, when the current block is encoded by a specific encoding / decoding tool, n can be fixed to a predefined value of that specific encoding / decoding tool.

[0170] In another embodiment, the pattern of n points refers to a pattern defined as any subset of a window region M x N surrounding / containing the position (iL, jL), such as Figure 12A As shown. If the target sample is luminance, then (iL,jL) is (i, j). If the target sample is chrominance (e.g., Cb or Cr), then (iL, jL) is the luminance position from (i, j).

[0171] For example, using only the window center (iL, jL) as the source item, such as... Figure 12A As shown.

[0172] Another example is a 5x5 cross pattern that includes or does not exclude (iL, jL), such as Figure 12B As shown.

[0173] For source items in the source item set, the following examples are used to determine the generation of source content.

[0174] In one embodiment, the source content is based on predicted samples generated from the prediction pattern and / or reconstructed samples generated from the predicted samples and reconstruction residuals.

[0175] In another sub-implementation, the source content is a preprocessed source or any preprocessed source. For example, the source content is a predicted / reconstructed sample filtered by a predefined model or filter.

[0176] In another sub-implementation, the source content is gradient information from the predicted and / or reconstructed samples. If the target sample (i, j) belongs to chroma, and the gradient information of the corresponding brightness sample (as the center circle) is used... Figure 13 The gradient information of the source term for the target sample (i, j) is calculated using any of the Sobel filters (1310-1340) or any predefined filters. Each value around the central circle is multiplied by the corresponding predicted / reconstructed sample in the same brightness block, and then summed to form the gradient information of the source term for the target sample (i, j).

[0177] In another sub-implementation, since the target sample is a chroma sample (e.g., Cb or Cr), the predicted and / or reconstructed samples are located within the co-located (luminance) block from the current (chroma) block. The predicted and / or reconstructed samples are treated as initial samples and used as source content to generate the target sample.

[0178] In another embodiment, the source item may also include position information. For example, if the target sample refers to luminance, then the horizontal position (i) of (i, j) is used as the source item, and the vertical position (j) of (i, j) is used as the source item; otherwise, the horizontal position of the luminance block in the same position as sample (i, j) is used as the source item, and the vertical position of the luminance block in the same position as sample (i, j) is used as the source item.

[0179] In another embodiment, the source term may also include location information. For example, if the target sample refers to chromaticity, the horizontal position of the iso-brightness of sample (i, j) is used as the source term, and the vertical position of the iso-brightness of sample (i, j) is used as the source term.

[0180] VII.2. Contents of sourceTermSet1(i, j) `sourceTermSet1(i, j)` includes one or more chromaticity (Cb or Cr) source terms, denoted as `sourceTerm00`, `sourceTerm01`, ..., and / or `sourceTerm0m-1`. The value of `m` represents the number of points in the source terms set. In one embodiment, the source terms can be linear and / or nonlinear, linear only, or nonlinear only. In another embodiment, `m` is a predefined value, such as 1, 2, ..., or any positive integer. For example, the predefined value is fixed in a standard.

[0181] In another embodiment, m is determined based on the encoding / decoding information of the current block and / or the sample position (i, j). For example, when the current block is encoded by a specific encoding / decoding tool, m is fixed to a predefined value of that specific tool.

[0182] In another embodiment, the pattern of point m refers to a pattern defined as any subset of the M2 x N2 window region surrounding / containing the position (iC, jC), such as Figure 14A As shown. If the target sample is chromaticity (Cb or Cr), then (iC,jC) is (i, j). If the target sample is luminance, then (iC, jC) is the isotopic chromaticity position obtained from (i, j).

[0183] For example, only the center (iC, jC) of the window is used, such as Figure 14A As shown.

[0184] Another example of a pattern is a 5x5 cross: (containing or not excluding (iC, jC)), such as Figure 14B As shown.

[0185] For a source item in the source item set, the following examples are used to determine the generation of source content.

[0186] In one embodiment, the source content is based on a prediction sample generated from a prediction pattern and / or a reconstruction sample generated from the prediction sample, the prediction pattern, and the reconstruction residual.

[0187] In another sub-implementation, the source content is a filtered source or a source that has undergone any preprocessing. For example, the source content is a predicted / reconstructed sample filtered by a predefined model or filter.

[0188] In another sub-implementation, the source content is gradient information from the predicted and / or reconstructed samples. If the target sample (i, j) belongs to luminance, the gradient information of the isotopic chromaticity sample is calculated using any Sobel filter or any predefined filter.

[0189] In another sub-implementation, if the target sample is a chroma sample, the predicted and / or reconstructed samples are located within the current block. The predicted and / or reconstructed samples are treated as initial samples and used as source content for generating the target sample.

[0190] In another embodiment, the source item may also include location information. For example, if the target sample refers to chroma, then the horizontal position (i) of (i, j) is used for the source item, while the vertical position (j) of (i, j) is used for the source item.

[0191] VII.3. Content of biasTermSet The deviation term is a predefined value. In one embodiment, the deviation term is the midValue based on the bit depth specified in the standard. For example, the deviation term is set to... In another embodiment, the bias term for each sample is the same in the current block. That is, the bias term is independent of position (i, j).

[0192] VII.4. Predictor Derivation for Sample (i, j) VII.4.1. Suggested Weighting The proposed weighting method estimates the relationship between "predicted and / or reconstructed samples on the reference region of the current (chroma) block" and "predicted and / or reconstructed samples on the reference region of the corresponding luma block" using a predefined regression method (e.g., minimizing distortion), and generates weights (referring to model parameters) based on the regression method. The derived weights are then applied to the source terms to obtain the target (predicted) samples in the current block. In one embodiment, the predefined regression method may be the Linear Least Mean Squared Error (LMMSE) method of CCLM, or any unified method with the regression method used by CCLM. In another embodiment, the predefined regression method may be the LDL decomposition method of CCCM, or any unified method with the regression method used by CCCM. In yet another embodiment, the predefined regression method may be Gaussian elimination.

[0193] In one embodiment, the reference region of the current block is the spatial proximity region of the current block 1510, such as... Figure 15 As shown. The spatial proximity region of the current block includes the upper reference region 1520, the left reference region 1530, the upper left reference region 1540, and / or any subset thereof. The size of the upper reference region is A. w xA H The size of the reference region on the left is L. w xL H The size of the upper left reference area is AL. W xAL H Its characteristics A w = Current block width (W), k*W, W + current block height (H), any predefined value, or any adaptive value based on block position, block width, block height and / or block area.

[0194] A H or AL H = H, any predefined value (1, 2, 4, …), or any adaptive value based on block location, block width, block height and / or block area.

[0195] L W or ALW = W, any predefined value (1, 2, 4, …), or any adaptive value based on block location, block width, block height and / or block area.

[0196] L H = H, k*H, H + W, any predefined value, or any adaptive value based on block location, block width, block height, and / or block area.

[0197] The reference area for a corresponding luminance block is the spatially adjacent area of ​​that luminance block.

[0198] In another embodiment, the reference region of the current block is the vector co-location region of the current block, and the reference region of the corresponding luma block is the vector co-location region of the corresponding luma block. For a cross-coding unit containing luma and chroma blocks, the vector co-location region of the current block refers to the motion compensation result obtained using the motion information (motion vector and reference image) of the current block, while the vector co-location region of the corresponding luma block refers to the motion compensation result obtained using the motion information (motion vector and reference image) of the corresponding luma block. For IBC or intraTMP, the vector co-location region of the current block refers to the motion compensation result obtained using the motion information (e.g., block vector and current image) of the current block, while the vector co-location region of the corresponding luma block refers to the motion compensation result obtained using the motion information (e.g., block vector and current image) of the corresponding luma block.

[0199] In another embodiment, the two reference regions for the current block described above can be used together. For example, typically, when deriving model parameters, samples in the vector co-location region of the current block are used as input samples; however, for smaller blocks, samples in the spatially neighboring reference region are used as additional input samples when deriving model parameters.

[0200] In this application, the term “block” may refer to TU / TB, CU / CB, PU / PB, or CTU / CTB.

[0201] In this application, "LM" can be considered as one of the CCLM / MMLM modes or any other extension / variation of CCLM (e.g., the CCLM extension / variation proposed in this application). One variant is MMLM, which uses a threshold to determine different models for different samples in the current chroma component. Another variant, for Cb (or Cr), derives model parameters from multiple iso-luminance blocks. More possible variants are shown below. CCLM variants here mean that when the block indication reference uses one of the cross-component modes (e.g., CCLM_LT, MMLM_LT, CCLM_L, CCLM_T, MMLM_L, MMLM_T, and / or an intra-prediction mode that is not one of the traditional DC, planar, and angular modes), several optional modes can be selected. An example of using the Convolutional Cross-Component Mode (CCCM) as an optional mode is shown below. When this optional mode is applied to the current block, cross-component information of the model containing nonlinear terms is used to generate chroma predictions. Optional modes may follow the template selection of CCLM, so the CCCM family includes CCCM_LT, CCCM_L, and / or CCCM_T.

[0202] The method proposed in this application (for CCLM) can be used for any other cross component mode.

[0203] Any combination of methods proposed in this application may be applied.

[0204] Any of the cross-component prediction methods proposed above that use a cross-component prediction model derived from a reference image resampled (RPR) can be implemented in the encoder and / or decoder. For example, any proposed method can be implemented in the cross-component, intra-component, prediction, IBC, transform, quantization modules or combinations thereof at the encoder end, and / or in the cross-component, intra-component / prediction, IBC, transform, quantization modules or combinations thereof at the decoder end. Alternatively, any proposed method can be implemented as circuitry connected to the cross-component, intra-component, prediction, transform, quantization modules or combinations thereof of the encoder and / or decoder to provide the information required by the cross-component / intra-component / prediction / IBC / transform / quantization modules.

[0205] As described above, the cross-component prediction model derived from the reference image resampling (RPR) can be implemented at either the encoder or decoder end. For example, any proposed method can be implemented in the intra-frame / cross-decoding module of the decoder (e.g., Figure 1B Intra Pred. 150 / MC 152) or the intra-frame / cross-decoding module in the encoder (e.g., Figure 1AThe proposed method is implemented in Intra Pred. 110 / Inter Pred. 112. Any proposed method can also be implemented as a circuit connected to the intra-frame / cross-coding / decoding module of a decoder or encoder. However, the decoder or encoder may also use additional processing units to implement the proposed method. While the Intra Pred. / MC unit (e.g., Figure 1A Units 110 / 112 and Figure 1B Units 150 / 152 in the diagram are shown as separate processing units, which may correspond to executable software or firmware code stored on media such as hard disks or flash memory for CPUs (Central Processing Units) or programmable devices (e.g., DSPs (Digital Signal Processors) or FPGAs (Field Programmable Gate Arrays)).

[0206] Figure 16 A flowchart illustrating an exemplary video encoding / decoding system, according to one embodiment of this application, incorporates a cross-component prediction model derived from a reference image resampled (RPR) image. The steps shown in the flowchart can be implemented as program code executable on one or more processors (e.g., one or more CPUs) at the encoder or decoder end. The steps shown in the flowchart can also be implemented in hardware, such as one or more electronic devices or processors arranged to execute the steps in the flowchart. According to this method, in step 1610, input data associated with the current block is received, including a first color block and a second color block, characterized in that the input data includes pixel data to be encoded at the encoder end or data associated with the current block to be decoded at the decoder end, and characterized in that the current block is encoded / decoded in a non-intra-frame mode. In step 1620, a co-image is selected from one or more lists of reference images according to one or more predefined rules. In step 1630, it is determined whether the selected co-image is a reference image resampled (RPR) image. In step 1640, based on whether the co-image is an RPR image, a current cross-component prediction (CCP) model is derived from the co-image. In step 1650, the current second color block is encoded or decoded using a candidate list including the current CCP model. The feature is that when the current CCP model is selected to encode the current second color block, the current CCP model is applied to the current first color block to generate prediction data for the current second color block.

[0207] The flowchart shown is intended to illustrate an example of video encoding / decoding according to this application. Those skilled in the art can modify each step, rearrange the steps, split the steps, or combine the steps to practice this application without departing from its spirit. Specific syntax and semantics are used in this disclosure to illustrate examples of implementing embodiments of this application. Those skilled in the art can practice this application by substituting equivalent syntax and semantics without departing from its spirit.

[0208] The foregoing description is intended to enable those skilled in the art to practice this application according to specific applications and requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, this application is not intended to be limited to the specific embodiments shown and described, but rather to be given the broadest scope consistent with the principles and novel features disclosed herein. In the foregoing detailed description, various specific details have been depicted to provide a thorough understanding of this application. However, those skilled in the art will understand that this application can be practiced.

[0209] The embodiments of this application described above can be implemented in various hardware, software code, or a combination of both. For example, one embodiment of this application may be one or more circuits integrated into a video compression chip, or program code integrated into video compression software to perform the processes described herein. Another embodiment of this application may be program code to be executed on a digital signal processor (DSP) to perform the processes described herein. This application may also relate to multiple functions executed by a computer processor, digital signal processor, microprocessor, or field-programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the invention, defining the specific methods embodied in the invention by executing machine-readable software code or firmware code. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, styles, and languages ​​of the software code, as well as other means of configuring the code to perform tasks according to the invention, do not depart from the spirit and scope of the invention.

[0210] This application may be embodied in other specific forms without departing from its spirit or essential characteristics. The examples described should be considered illustrative in all respects, not restrictive. Therefore, the scope of this application should be indicated by the appended claims rather than the foregoing description. All variations within the meaning and equivalence of the claims should be included within its scope.

Claims

1. A method of coding a color picture using a coding tool comprising one or more modes related to cross-component modeling, the method comprising: receiving input data related to a current block, including a current first color block and a current second color block, wherein the input data comprises pixel data to be encoded at an encoder side or data related to the current block to be decoded at a decoder side, and wherein the current block is coded in a non-intra mode; selecting a collocated picture from one or more reference picture lists according to one or more predefined rules; determining whether the selected collocated picture is a reference picture resampled picture; deriving a current cross-component prediction model from the collocated picture depending on whether the collocated picture is the reference picture resampled picture; and coding or decoding the current second color block using a candidate list comprising the current cross-component prediction model, wherein when the current cross-component prediction model is selected for coding or decoding the current second color block, a prediction data of the current second color block is generated by applying the current cross-component prediction model to the current first color block.

2. The method of claim 1, wherein, the one or more predefined rules are related to information comprising L0[0], L1[0], picture order count distance, QP value, or a combination thereof.

3. The method of claim 1, wherein, whether the selected collocated picture is the reference picture resampled picture is indicated by a flag signaled or parsed in a bitstream.

4. The method of claim 1, wherein, the selected collocated picture is the reference picture resampled picture if the collocated picture and a current picture containing the current block differ in one or more target parameters.

5. The method of claim 4, wherein, the one or more target parameters comprise a picture width in luma samples, a picture height in luma samples, a scaling window left offset, a scaling window right offset, a scaling window top offset, a scaling window bottom offset, a number of sub-pictures, or a combination thereof.

6. The method of claim 1, wherein, the collocated picture is selected from one or more non-scaled pictures in the one or more reference picture lists.

7. The method of claim 1, wherein, a target picture is selected from the non-scaled pictures in the one or more reference picture lists as the collocated picture, and wherein the target picture is selected based on a picture order count difference from a current picture, a picture order count value, a QP difference from a current picture, a QP value, a reference list, or a combination thereof.

8. The method of claim 1, wherein, if the selected collocated picture is the reference picture resampled picture, cross-component model information inherited from the collocated picture is disabled.

9. The method of claim 1, wherein, if the selected collocated picture is the reference picture resampled picture, cross-component model information from the collocated picture is retrieved from a scaling location according to a scaling ratio.

10. The method of claim 9, wherein, the scaling ratio is derived based on a scaling window of a current picture and the collocated picture.

11. The method of claim 10, wherein, if a current location and the scaling ratio are represented as (x, y) and R, respectively, a scaling location is determined as (x / R, y / R) after a rounding process.

12. The method of claim 11, wherein, the rounding process corresponds to rounding to negative infinity, rounding to positive infinity, rounding to zero, or rounding to the nearest integer.

13. The method of claim 1, wherein, The current cross-component prediction model from the collocated picture is determined from motion vectors of neighboring blocks, and the neighboring blocks are selected from a list of predefined positions.

14. The method of claim 13, wherein, The neighboring blocks are selected from the list of predefined positions according to a predefined checking order.

15. The method of claim 14, wherein, The predefined checking order corresponds to checking first a L0 motion vector or a L1 motion vector, and selecting first a target motion vector related to an unscaled reference picture.

16. The method of claim 1, wherein, If a target reference picture located by a motion vector is the reference picture resampling picture, the motion vector is considered as not having cross-component model information located by the motion vector.

17. The method of claim 1, wherein, If a target reference picture located by a motion vector is the reference picture resampling picture, cross-component model information is retrieved from the target reference picture at a scaled position according to a scaling ratio.

18. An apparatus for encoding or decoding a color picture or video using an encoding tool including one or more cross-component model related modes, the apparatus comprising one or more electronic circuits or processors configured to: Receiving input data related to a current block, including a current first color block and a current second color block, characterized in that, the input data message containing pixel data to be encoded at an encoder side or data related to the current block to be decoded at a decoder side, and wherein the current block is encoded in a non-intra mode; selecting a collocated picture from one or more reference picture lists according to one or more predefined rules; determining whether the selected collocated picture is a reference picture resampling picture; deriving a current cross-component prediction model from the collocated picture according to whether the collocated picture is the reference picture resampling picture; and encoding or decoding the current second color block using a candidate list including the current cross-component prediction model, and wherein when the current cross-component prediction model is selected for encoding or decoding the current second color block, a prediction data of the current second color block is generated by applying the current cross-component prediction model to the current first color block.