Cross-component model propagation decision rules based on block vectors and motion vectors
By using a cross-component prediction method, the prediction of chroma blocks is generated by inheriting parameters from the reference block through a cross-component model. This solves the problem of low encoding and decoding efficiency of chroma and luminance components in the existing technology and achieves more efficient video coding results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video coding standards are inefficient in cross-component prediction and struggle to effectively utilize the correlation between chroma and luminance components for efficient encoding and decoding.
The cross-component prediction method is adopted. By generating a cross-component model that inherits model parameters from the reference block, it is applied to the reconstructed sample of the first color block to generate the prediction of the second color block. The cross-component model is then propagated using a selected further reference block. By combining multiple models such as CCLM, CCCM, GLM and CCRM, the prediction accuracy is improved.
It improves the efficiency and quality of video coding, enhances the accuracy of chroma and luminance component prediction, and reduces coding complexity and storage requirements.
Smart Images

Figure CN121753339A_ABST
Abstract
Description
[Technical Field] This disclosure generally relates to video encoding and decoding. In particular, this disclosure relates to methods for encoding and decoding pixel blocks by cross-component prediction, specifically by propagating a cross-component model. [Background Technology] Unless otherwise stated herein, the methods described in this section are not prior art to the following claims and are not admitted as prior art by virtue of their inclusion in this section.
[0003] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC is based on a hybrid block-based motion-compensated Discrete Cosine Transform (DCT-like) coding architecture. The basic unit of compression, called a coding unit (CU), is a 2Nx2N pixel block. Each CU can be recursively divided into four smaller CUs until a predefined minimum size is reached. Each CU contains one or more prediction units (PUs).
[0004] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from a reconstructed signal, which is derived from the image regions of the coded picture. The prediction residual signal is processed through block transform. The transform coefficients are quantized and entropy-coded in the bitstream along with other additional information. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal by performing an inverse transform on the dequantized transform coefficients. The reconstructed signal is further processed by loop filtering to remove coding artifacts. The decoded picture is stored in the frame buffer and used to predict future pictures in the input video signal.
[0005] In VVC, the encoded image is divided into non-overlapping block regions represented by relevant coding tree units (CTUs). Leaf nodes of the coding tree correspond to coding units (CUs). The encoded image can be represented by multiple slices, each containing an integer number of CTUs. The CTUs within a slice are processed in raster scan order. Bidirectional prediction (B) slices can be decoded using either intra-frame or inter-frame prediction, using at most two motion vectors and a reference index to predict the sample value for each block. Predictive (P) slices are decoded using either intra-frame or inter-frame prediction, using at most one motion vector and a reference index to predict the sample value for each block. Intra-frame (I) slices are decoded using only intra-frame prediction.
[0006] CTUs can be partitioned into one or more non-overlapping coding units (CUs) using quadtree (QT) and nested multi-type-tree (MTT) structures to accommodate various local motion and texture characteristics. CUs can be further partitioned into smaller CUs using one of five partitioning types: quadtree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical center-side ternary tree partitioning, and horizontal center-side ternary tree partitioning.
[0007] Each CU contains one or more prediction units (PUs). The prediction unit, along with the associated CU syntax, serves as the basic unit of signaling predictor information. The specified prediction process is used to predict the values of associated pixel samples within the PU. Each CU may contain one or more transform units (TUs) to represent prediction residual blocks. A transform unit (TU) consists of a transform block (TB) for one luma sample and two corresponding transform blocks for two chroma samples; each TB corresponds to a residual sample block from one color component. Integer transforms are applied to the transform blocks. The level values of the quantization coefficients, along with other additional information, are entropy-encoded in the bitstream. The terms coding tree block (CTB), coding block (CB), prediction block (PB), and transform block (TB) are defined as two-dimensional sample arrays specifying the monochromatic components associated with the CTU, CU, PU, and TU, respectively. Therefore, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. CUs, PUs, and TUs have similar relationships.
[0008] For each inter-frame predicted CU, motion parameters include motion vectors, reference picture indices, and reference picture list usage indices, as well as additional information for inter-frame predicted sample generation. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded motion vector increments, and no reference picture indices. A merging mode is specified where the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, and additional plans introduced in the VVC. The merging mode can be applied to any inter-frame predicted CU. An alternative to the merging mode is explicit signaling of motion parameters, where motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other necessary information are explicitly signaled in each CU. [Summary of the Invention] The following abstract is illustrative only and is not intended to be limiting in any way. That is, the following abstract aims to introduce the concepts, highlights, benefits, and advantages of the novel and non-obvious techniques described herein. Alternative embodiments will be further described in the detailed description. Therefore, the following summary is not intended to identify the essential features of the claimed subject matter, nor is it intended to determine the scope of the claimed subject matter.
[0010] Certain embodiments of this disclosure provide a method for cross-component prediction of pixel blocks. A current block has a first color block and a second color block. A video codec generates a reconstructed sample for the first color block. The video codec inherits a cross-component model from a reference block, wherein the cross-component model is propagated based on cross-component models of two or more further reference blocks for use by the reference block. The video codec applies the inherited cross-component model to the reconstructed sample of the first color block to generate a cross-component prediction for the second color block. The video codec uses the generated cross-component prediction to encode or decode the current block.
[0011] The two or more further reference blocks may be located by more than one motion vector or block vector of the reference block (e.g., when the reference block is a bidirectional inter-frame prediction block). In some embodiments, the reference block is not encoded by cross-component prediction.
[0012] In some embodiments, the video codec provides the cross-component model for the reference block to use by selecting a selected further reference block from two or more further reference blocks of the reference block, thereby propagating the cross-component model. The selected further reference block can be selected according to a set of predefined rules. The video codec can select the selected further reference block by identifying a block POC coded by cross-component prediction. The video codec can select the selected further reference block from two or more further reference blocks that uniquely possesses the cross-component model. The video codec can select the selected further reference block by identifying a block coded by the current picture reference (e.g., IBC mode), or by identifying a block coded by intra-frame prediction, or by identifying a block coded by inter-frame prediction.
[0013] In some embodiments, the selected further reference block is selected by identifying the block with the shortest spatial distance from the reference block among two or more further reference blocks. The spatial distance can be a vertical or horizontal distance, or a Euclidean distance, or a Manhattan distance, or a Minkowski distance. In some embodiments, the selected further reference block is selected by identifying the block with the shortest temporal distance from the reference block among two or more further reference blocks. The temporal distance of a block can be determined based on the point of view (POC) of the block's further reference images and the POC of the reference images.
[0014] The selected further reference block can be selected by identifying the block in each of the further reference blocks that has the closest quantization parameter QP to that reference block, or by identifying the block in each of the further reference blocks that has the largest QP, or by identifying the block in each of the further reference blocks that has the smallest QP, or by identifying the block in each of the further reference blocks that is indicated by the L0 (or L1) motion vector. [Attached Image Description] The accompanying drawings are included to provide a further understanding of the present disclosure and form part of it. The drawings illustrate embodiments of the present disclosure and, together with the detailed description, serve to explain the principles of the disclosure. It will be understood that the drawings are not necessarily drawn to scale, as some components may be shown out of proportion to their actual dimensions in order to clearly illustrate the concepts of the present disclosure.
[0016] Figure 1 A conceptual demonstration is presented of chromaticity and luminance samples used to derive parameters for linear models.
[0017] Figure 2 An example is shown where adjacent samples are classified into two groups.
[0018] Figure 3 The spatial components of a convolution filter are conceptually explained.
[0019] Figure 4 This demonstrates the gradient linear model (GLM) used by the gradient filter.
[0020] Figure 5 This demonstrates component sample reconstruction using a cross-component residual model (CCRM).
[0021] Figure 6 This indicates the position of the chromaticity sample relative to the luminance sample predicted by the CCRM filter.
[0022] Figure 7 It demonstrates the predefined search regions used for intra-frame template matching.
[0023] Figure 8 Displays the luminance blocks used to derive the direct block vectors of the corresponding chroma blocks.
[0024] Figure 9 This indicates that model parameters are inherited from the predefined locations of neighboring blocks in space.
[0025] Figure 10 This section provides a conceptual explanation of the parameters for the inheritance time proximity model.
[0026] Figure 11A -B indicates a current block and its non-adjacent spatial neighboring blocks from which model parameters can be inherited.
[0027] Figure 12 This conceptually demonstrates an example of CCM information propagation based on block vectors.
[0028] Figure 13 Explanation 1: The current block has two block directions used to identify two reference blocks, each with CCM information.
[0029] Figure 14 This section illustrates an example of CCM information propagation based on motion vectors.
[0030] Figure 15 Explanation 1: The current block has two motion directions used to identify two reference blocks, each with CCM information.
[0031] Figure 16 This illustrates an example of a video encoder that may implement cross-component prediction.
[0032] Figure 17 This describes a part of the video encoder used to implement the propagation of CCM information from multiple reference blocks.
[0033] Figure 18This conceptually illustrates the process of propagating CCM information from multiple reference blocks when encoding a block.
[0034] Figure 19 This illustrates an example of a video decoder that may implement cross-component prediction.
[0035] Figure 20 This describes a part of the video decoder used to implement the propagation of CCM information from multiple reference blocks.
[0036] Figure 21 This conceptually illustrates the process of propagating CCM information from multiple reference blocks when decoding a block.
[0037] Figure 22 This invention provides a conceptual description of electronic systems that implement certain embodiments of the present disclosure.
Detailed Implementation Methods
[0039] I. Cross-component prediction a. Cross Component Linear Model (CCLM) Cross-component linear model (CCLM) or linear model (LM) mode is a cross-component prediction mode in which the chromaticity component of a block is predicted from reconstructed luminance samples located at colocated positions using a linear model. The parameters of the linear model (e.g., scaling and offset) are derived based on reconstructed luminance and chromaticity samples from neighboring blocks. For example, in VVC, CCLM mode uses inter-channel dependencies to predict chromaticity samples from reconstructed luminance samples. This prediction system uses a linear model, expressed as: (1) In equation (1) This represents a predicted chromaticity sample in a CU (or a predicted chromaticity sample in the current CU). Represents the downsampled reconstructed luminance sample of the same CU (or the corresponding reconstructed luminance sample of the current CU).
[0040] The CCLM model parameters α (scaling parameter) and β (offset parameter) are derived based on up to four neighboring chromaticity samples and their corresponding downsampled luminance samples. In LM_A mode (also known as LM-T mode), only the upper or top neighboring template is used to calculate the linear model coefficients. In LM_L mode (also known as LM-L mode), only the left template is used to calculate the linear model coefficients. In LM-LA mode (also known as LM-LT mode), both the left and upper templates are used to calculate the linear model coefficients. In this disclosure, the terms {LM_LA, LM_A, LM_L} and {CCLM_LT, CCLM_T, CCLM_L} are used interchangeably. The terms CCLM_LT, LM_LA, and CCLM_LA are also used interchangeably.
[0041] Figure 1 This diagram conceptually illustrates the chromaticity and luma samples used to derive the parameters of a linear model. The figure shows a current block 100 with luma and chromaticity component samples in 4:2:0 format. Luma and chromaticity samples from neighboring blocks are reconstructed samples. These reconstructed samples are used to derive the cross-component linear model (parameters α and β). Because the current block is in 4:2:0 format, the luma samples are first downsampled before being used for linear model derivation. In this example, there are 16 pairs of reconstructed luma (downsampled) and chromaticity samples from neighboring blocks. These 16 pairs of luma and chromaticity values are used to derive the linear model parameters.
[0042] Assuming the current chroma block size is W×H, then W' and H' are set to... – W'= W, H'= H when the LM-LT mode is applied; – W' = W + H when the LM-T pattern is applied; – H' = H + W When the LM-L pattern is applied The aforementioned neighboring locations are labeled S[0,-1]...S[W'-1,-1], and the left neighboring locations are labeled S[-1,0]...S[-1,H'-1]. Then, four samples are selected as... – S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4] When the LM-LT mode is applied (both the upper and left neighboring samples are available); – S[W' / 8,-1], S[3 *W' / 8,-1], S[5*W' / 8,-1], S[7*W' / 8,-1] When the LM-T mode is applied (only the upper neighbor sample is available); – S[-1,H' / 8], S[-1,3*H' / 8], S[-1,5*H' / 8 ], S[-1, 7*H' / 8] when the LM-L mode is applied (only the left neighbor sample is available); Four neighboring brightness samples at the selected location are downsampled and compared four times to find the two larger values: And two smaller values: Their corresponding chromaticity sample values are labeled as Then X A X B Y A and Y B Exported as: (2) (3) The linear model parameters are obtained according to the following equation: (4) (5) In some embodiments, to obtain more samples for calculating the CCLM model parameters, the template above is expanded to include (W+H) samples for the LM-T mode, and the left template is expanded to include (H+W) samples for the LM-L mode. For the LM-LT mode, both the left and top templates are used to calculate the linear model coefficients.
[0043] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filters is determined by the sequence parameter set (SPS) level flag. These two downsampling filters correspond to "Type-0" and "Type-2" content, respectively.
[0044] (6) (7) b. Multi-model CCLM (MMLM) The Multi-Model CCLM mode (MMLM) uses two models to predict chromaticity samples from luminance samples and applies this to the entire CU. Similar to CCLM, three Multi-Model CCLM modes (MMLM_LA, MMLM_A, and MMLM_L) are used to indicate whether to use both the top and left neighbor samples together, use only the top neighbor samples, or use only the left neighbor samples to derive model parameters.
[0045] In MMLM, the neighboring luminance and chrominance samples of the current block are classified into two groups, each used as a training set to derive a linear model (i.e., derive specific α and β for a specific group). Furthermore, samples of the current luminance block are also classified according to the classification rules of the neighboring luminance samples.
[0046] Figure 2 This demonstrates an example of classifying neighboring samples into two groups. A threshold is calculated by averaging the brightness samples reconstructed from the nearest neighbors. Neighboring samples located in [x,y] are considered when Rec'... L When [x,y] <= this threshold, it is classified as group 1; while for neighboring samples located in [x,y], if Rec' L When [x,y] > the threshold, it is classified as group 2. Therefore, the multi-model CCLM prediction for chromaticity samples is: Pred c [x,y]=α1×Rec' L [x,y]+β1 when Rec' L [x,y]≤threshold Pred c [x,y]=α2×Rec' L [x,y]+β2 when Rec' L [x,y]>threshold c. Convolutional Cross-Component Model (CCCM) In some embodiments, a convolutional cross-component model (CCCM) is applied to improve cross-component prediction performance. In some embodiments, the convolutional model has a 7-tap filter with 5-tap spatial components of a sign shape, a nonlinear term, and a bias term. The 5-tap spatial components of the filter include a center (C) luminance sample (collocated with the chrominance sample to be predicted) and its top / north (N), bottom / south (S), left / west (W), and right / east (E) neighbors. Figure 3 The spatial components of a convolutional filter are conceptually explained. The nonlinear term (denoted as P) is represented as a square power of the center brightness sample C, scaled according to the range of sample values of the content: P = (C*C+midVal)>>bitDepth (8) Therefore, for 10-bit content, the nonlinear term P is calculated as follows: P = (C*C + 512) >> 10 The bias term (denoted as B) represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content). The filter output is calculated as the filter coefficients c. iConvolve the input values and crop them to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P +c6B (9) d. Gradient Linear Model (GLM) For the YUV 4:2:0 color format, the gradient linear model (GLM) method can be used to predict chromaticity samples from luminance sample gradients. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.
[0047] Compared to CCLM, the two-parameter GLM uses the gradient of luminance samples instead of downsampled luminance values to derive a linear model. Specifically, when applying the two-parameter GLM, the input to the CCLM process, i.e., the downsampled luminance sample L, is replaced by the luminance sample gradient G.
[0048] Other parts of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged. In three-parameter GLM, chromaticity samples can be predicted based on the luminance sample gradient and downsampled luminance values with different parameters: The model parameters of the three-parameter GLM are derived from the 6 row and column neighbor samples through LDL decomposition based on the MSE minimization method, just as used in CCCM.
[0049] For signal transmission, if the current CU is in CCLM mode, a flag will be sent to indicate whether GLM is enabled for both Cb and Cr components; if GLM is enabled, another flag will be sent to indicate which of the two GLM modes is selected, and a syntax element will be sent to select one of the four gradient filters for gradient calculation. Figure 4 The figure illustrates the gradient linear model (GLM) used by the gradient filter. Specifically, it shows four Sobel-based GLM gradient modes 401-404. The gray dot in the middle of each gradient mode represents the chromaticity position (C).
[0050] e. Cross-Component Residual Model (CCRM) Cross-component residual model (CCRM) can be used to predict chrominance samples from reconstructed luminance samples if the current block uses inter-frame prediction or intra-block copy (IBC, which refers to encoding and decoding pixel blocks by using block vectors to reference pixel positions within the same current image).
[0051] Figure 5 The component sample reconstruction using the Cross-Component Residual Model (CCRM) is shown. This figure illustrates the decoding method at the decoder end. The cross-component filter is derived using the predicted signals of luma and chroma. The derived filter is applied to the reconstructed luma signal to produce the final chroma prediction.
[0052] The derived filter may be an 8-tap filter consisting of 6 spatial brightness samples, a nonlinear term, and a bias term. Figure 6 This illustrates the position of the chromaticity samples predicted by the CCRM filter relative to the luminance samples. As shown in the figure, the spatial luminance samples (L0, L1, L2, L3, L4, L5) are obtained from the luminance grid by selecting the 6 luminance samples closest to the chromaticity position C, without downsampling. The chromaticity values predicted by the filter can be obtained as follows: predChromaVal = c0L0+c1L1+c2L2+c3L3+c4L4+c5L5 +c6nonlinear((L0+L3+1)>>1) + c7B, Where "nonlinear" is the nonlinear operator in CCCM, and B is the bias term.
[0053] II. Current Image Reference Intra-block copy (IBC) of the current image Intra-picture block copying (IBC) or current picture referencing (CPR) refers to encoding and decoding pixel blocks by using block vectors to reference pixel locations within the same current picture.
[0054] b. Intra-frame template matching Intra-Template Matching Prediction (IntraTMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, matching its L-shaped template with the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed at the decoder. The prediction signal is generated by matching the L-shaped, top-only, or left-only causal neighbor pixels of the current block with another block in the predefined search region.
[0055] Figure 7 This section describes the predefined search regions used for intra-frame template matching. As shown in the figure, for the current block 710 with the current template region 720, a search is performed within the current CTU 705 in several predefined search regions to find a matching reference template region 730 (thus finding the corresponding reference block 735). The predefined search regions include R1 (within the current CTU), R2 (upper left of the current CTU), R3 (above the current CTU), and R4 (to the left of the current CTU). Template matching uses the Sum of Absolute Differences (SAD) as the cost function.
[0056] c. Direct block vector of chroma blocks Within a dual-tree slice, direct block vectors can be used for chroma blocks. When dual-tree chroma is enabled, a flag is signaled to indicate whether the chroma blocks are encoded using IBC mode. Specifically, if a luma block at one of the five predefined locations is encoded using IBC or IntraTMP mode, its block vector (BV) is scaled and used as the block vector for the chroma block. Template matching is used to perform block vector scaling (from the luma domain to the chroma domain). Figure 8 The figure shows the luma blocks used to derive the direct block vectors of the corresponding chroma blocks. It also shows five predefined positions for determining whether to scale the BV of the luma block and use it as the BV of the chroma block.
[0057] III. Inheritance Cross-Component Model a. Inheriting neighbor model parameters When a cross-component predictive codec is applied to the current block to generate a predicted signal, cross-component model (CCM) information, including model parameters, can be inherited from neighboring blocks. In some embodiments, if the inherited neighboring blocks are encoded in CCLM mode, the final scaling parameters of the current block are inherited from the neighboring blocks. Once the final scaling parameters are determined, offset parameters (e.g., ...) are derived based on the inherited scaling parameters and the average values of the neighboring luma and chroma samples of the current block. (in CCLM).
[0058] In some embodiments, if the inherited neighboring block is encoded in CCLM mode, the offset parameters can be inherited or further refined by dB after inheriting the model parameters. For example, if the final offset parameters are inherited from selected neighboring blocks, and the inherited offset parameters are... Then the final offset parameter is ( + dB). For example, dB can be derived from signals or from neighboring reconstructed samples. For example, dB can be zero.
[0059] In some embodiments, if the inherited neighboring block is CCCM encoded, then the filter coefficients ( ) is inherited. Offset parameters (e.g., or In CCCM, the filter coefficients can be re-derived based on the inherited parameters and the average values of the luminance and chrominance samples at the corresponding positions of the current block. In some embodiments, if the inherited neighboring blocks are encoded using CCCM, the filter coefficients ( ) is inherited. Offset parameters (e.g., or (In CCCM) it is also inherited, and is not re-exported when encoding or decoding the current block.
[0060] In some embodiments, if the inherited candidate block applies a GLM gradient pattern to its brightness reconstruction samples, the current block may also inherit the GLM gradient pattern of the candidate block and apply it to the current brightness reconstruction samples. In some embodiments, if the inherited neighboring blocks are encoded using multiple cross-component models (e.g., MMLM, or CCCM with multiple models), the classification threshold is also inherited to classify the neighboring samples of the current block into multiple groups, and the inherited multiple cross-component model parameters are further assigned to each group.
[0061] b. Inherit CCM information In some embodiments, inherited cross-component model CCM information can be stored along with inherited model parameters. CCM information can be inherited along with the inherited model parameters. Predictions for the current block can be generated based on the inherited CCM information and / or the inherited model parameters.
[0062] CCM information may include, but is not limited to: prediction mode (e.g., CCLM, MMLM, CCCM, CCCM with multiple models, 2-parameter GLM, 3-parameter GLM (GLM with a luminance term), CCRM), information indicating whether a nonlinear term is used in the model, model index indicating which model shape is used in the convolutional model, classification threshold of the multiple models, information indicating whether non-downsampled samples are used in the convolutional model, downsampled filter flag, downsampled filter index when multiple downsampled filters are used, information indicating whether multiple downsampled filters are used, number of neighboring lines used to derive the model, template type used to derive the model (e.g., top left, top, left), multiple model flag, post-filter flag, or model parameters.
[0063] In some embodiments, various terms (e.g., spatial, gradient, positional, nonlinear, and bias terms) included in a hybrid CCCM model can be inherited. In addition to storing model parameters, a prediction pattern can be stored in the CCM information to indicate that the inherited model is a hybrid CCCM model including various terms. For example, gradient-based and position-based CCCM (GL-CCCM) is a hybrid CCCM model that includes a spatial term for the center position, two gradient terms for the horizontal and vertical directions, two positional terms X and Y relative to the horizontal and vertical positions, and a nonlinear and bias term. The prediction pattern can be stored in the CCM information to indicate that the inherited model is a GL-CCCM model. If there are multiple types of hybrid CCCM models, a model index can also be stored in the CCM information to indicate which type of hybrid CCCM model is being inherited.
[0064] c. Inheritance Spatial Proximity Model Parameters In some embodiments, inherited model parameters may come from a directly neighboring block. Models from blocks at predetermined locations are added to a CCP merge candidate list in a predetermined order. The CCP merge candidate list may include models from spatial, temporal, and non-adjacent neighbors, and / or various models from a history table and / or from a default model.
[0065] The video encoder may signal the index to select a CCP merge candidate from the list. The video decoder may generate a corresponding predictor for the current block based on the selected inherited candidate model parameters. In some embodiments, the video encoder may signal a flag to indicate whether a CCP merge mode is used (e.g., after cclm_mode_flag). In some embodiments, an on / off flag is signaled to indicate whether the current block inherits the cross-component model parameters of neighboring blocks (i.e., whether a CCP merge mode is used). This flag may be signaled by CU / CB, by PU, by TU / TB, by color component, or by chroma color component. In some embodiments, when the flag is true, an index is signaled to indicate which CCP merge candidate is selected. If the current block inherits the cross-component model parameters of neighboring blocks, the inherited candidate index is signaled. The index may be encoded and signaled (e.g., truncate unary code, Exp-Golomb code, or fixed length code) and shared between the current Cb and Cr blocks.
[0066] In some embodiments, the predefined positions and predefined order can be the same as the spatial candidate positions of the inter-frame merging mode. Figure 9 This demonstrates the inheritance of model parameters from predefined positions of neighboring blocks in space. The predefined order can be B. 0, A 0, B 1, A1 and B2 were added to the CCP merge candidate list because their corresponding model parameters were added.
[0067] In some embodiments, assuming the current block's position, width, and height are (x, y), W, and H, respectively, a predefined position may include the position directly above the current block, such as (x+W>>1, y-1) or (x+(W+1)>>1, y-1), if W is greater than or equal to a threshold TH. A predefined position may also include the position directly to the left of the current block, such as (x-1, y+H>>1) or (x-1, y+(H+1)>>1), if H is greater than or equal to a threshold TH. TH can be 2, 4, 8, 16, 32, or 64. Predefined positions include the position directly above (W>>1) or ((W>>1)-1) if W is greater than or equal to TH, and the position directly to the left (H>>1) or ((H>>1)-1) if H is greater than or equal to TH.
[0068] In some embodiments, there is a maximum number of models inherited from spatial neighbors that can be added to the CCP merging candidate list, and this maximum number is less than the number of predefined locations.
[0069] d. Inheritance time proximity model parameters In some embodiments, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters may come from blocks in previously encoded slices / images.
[0070] Figure 10 This conceptually illustrates the parameters of the inheritance temporal proximity model. As shown in the figure, the current position is (x, y), and the block size is... The inherited model parameters can come from blocks at positions (x',y'), (x',y'+h / 2), (x'+w / 2,y'), (x'+w / 2, y'+h / 2), (x'+w,y'), (x',y'+h), or (x'+w,y'+h) in previously encoded slices / images, where x'=x+Δx and y'=y+Δy. In some embodiments, if the prediction mode of the current block is intra-frame prediction, Δx and Δy are set to 0. If the prediction mode of the current block is inter-frame prediction, Δx and Δy are set to the horizontal and vertical motion vectors of the current block. In some embodiments, if the current block is inter-frame bidirectional prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 0. In some embodiments, if the current block is inter-frame bidirectional prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 1.
[0071] In some embodiments, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can be derived from blocks in previously encoded slices / images. In one embodiment, the current block is located at (x, y), and the block size is... Two value sets and Defined as: In some embodiments, and All values in the set are positive. Let... The inherited model parameters can come from positions in previously encoded slices / images. The block.
[0072] In some embodiments, the current block is located at (x, y), and the block size is [missing information]. The inherited model parameters can come from positions in previously encoded slices / images. The block.
[0073] In some embodiments, ,For example In some embodiments, ,For example and In some embodiments, distance Models located closer to the target location are first added to the final cross-component prediction (CCP) merging candidate list. In some embodiments, distance... Models located closer to the target location are first added to the final cross-component prediction (CCP) merging candidate list.
[0074] In some embodiments, it is set and For two fixed positive numbers, the inherited model parameters can be derived from positions in previously encoded slices / images. The block.
[0075] The current block is located at (x, y) and its size is [value missing]. .set up and These are two fixed positive numbers. The inherited model parameters can be derived from positions in previously encoded slices / images. The block.
[0076] In some embodiments, the current block is located at (x, y), and the block size is [missing information]. The inherited model parameters can come from certain predefined locations in previously encoded slices / images. The block. For example, the position is inside the region corresponding to the current coding block, i.e. and Inherited model parameters can come from location. The block. Another example is a position outside the region corresponding to the current coded block, i.e. and Inherited model parameters can come from location. The block.
[0077] In some embodiments, the inherited model parameters may come from blocks at certain predefined locations. The predefined locations and inclusion order may be the same as in inter-frame merging modes. The previously encoded picture from which the inherited parameter model originates is subsequently referred to as a collocated picture. In some embodiments, the previously encoded picture from which the inherited parameter model originates, i.e., the collocated picture, is one of the pictures in a reference list.
[0078] In some embodiments, the co-bit image can be the same as the co-bit image in the inter-frame merge mode. In some embodiments, the co-bit image is signaled in the image / slice header. Reference lists and reference indices are signaled in the image / slice header. For example, the co-bit image is selected as L0[0]. Another example is that the co-bit image is selected as L1[0]. In some embodiments, the co-bit image is selected as the image in the reference list whose image order count (POC) differs the least from the current image. (The image order count is a value that indicates the temporal order of an image in a series of images.) For example, if the current image's POC is 8, the images in reference list 0 have POCs of {7, 6, 5, 0}, and the images in reference list 1 have POCs of {7, 6, 5, 4}, then L0[0] (equivalent to L1[0]) is selected because its POC difference is the smallest. In some embodiments, if two images have the smallest POC difference with the current image, the co-located image with the smaller POC is selected. In some embodiments, if two images have the smallest POC difference with the current image, the co-located image with the larger POC is selected. In some embodiments, if two images have the smallest POC difference with the current image, the co-located image with the smaller QP value difference with the current image is selected. In some embodiments, if two images have the smallest POC difference with the current image, the co-located image with the smaller QP is selected. In some embodiments, if two images have the smallest POC difference with the current image, the co-located image with the larger QP is selected.
[0079] In some embodiments, the co-occurring image is the image with the smallest QP difference between the reference list and the current image. For example, if the QP of the current image is 28, the QP of the images in reference list 0 is {19, 26, 23}, and the QP of the images in reference list 1 is {23, 22, 21}, then L0[1] is selected. In some embodiments, if more than one image in the reference list has the smallest QP difference with the current image, the image with the smaller QP is selected. In some embodiments, if more than one image in the reference list has the smallest QP difference with the current image, the image with the larger QP is selected. In some embodiments, if more than one image has the smallest QP difference with the current image, the image with the smaller POC distance is selected. In some embodiments, the co-occurring image is selected as the image with the smallest QP in the reference list. In some embodiments, the co-occurring image is selected as the image with the largest QP in the reference list.
[0080] In some embodiments, the inherited parameter model is derived from a previously encoded image, i.e., a co-bit image, which is the most recently encoded I-image. The cross-component model information of the most recently encoded I-slice / image is stored in a long-term reference buffer.
[0081] In some embodiments, the predefined locations from which the co-location image and inherited parameter model originate are determined by the motion vectors of neighboring blocks. For example, if the current block position is (x, y) and the block size is... The inherited model parameters can come from blocks in the co-located image located at (x', y'), (x', y'+h / 2), (x'+w / 2, y'), (x'+w / 2, y'+h / 2), (x'+w, y'), (x', y'+h), or (x'+w, y'+h), where x' = x + Δx and y' = y + Δy. In some embodiments, Δx and Δy are set as the L0 horizontal and vertical motion vectors of the neighboring blocks, and the co-located image is an L0 reference image indicated by the L0 motion vectors of the neighboring blocks. In some embodiments, if the neighboring blocks are inter-frame bidirectional predictions, Δx and Δy are set as the L1 horizontal and vertical motion vectors of the neighboring blocks, and the co-located image is an L1 reference image indicated by the L1 motion vectors of the neighboring blocks. In some embodiments, the neighboring block is the block to the left of the current block. In some embodiments, the neighboring block is the block above the current block.
[0082] In some embodiments, the predefined position from which the inherited parametric model in the previously encoded slice / image originates is determined by the motion vectors of neighboring blocks. Let Δx and Δy be the horizontal and vertical displacements based on the motion vectors of selected neighboring blocks, the current block position being (x, y), and the block size being... The inherited model parameters can come from the block located at (x', y'), where x' = x + Δx and y' = y + Δy, or from the block located at (x' + w / 2 + Δx and y' + h / 2 + Δy).
[0083] In some embodiments, the inherited model parameters may also come from positions in the patterns described in the preceding paragraphs. These positions are centered at (x', y'), where x' = x + Δx and y' = y + Δy, or x' = x + w / 2 + Δx and y' = y + h / 2 + Δy. That is, the predefined positions are represented as... The inherited model parameters can come from ,in and It is based on the horizontal and vertical displacements of the motion vectors selected from neighboring blocks. For example, suppose the size of the current block is... Two value sets and Defined as: All in and The values in the table are all positive. Inherited model parameters can be derived from positions in the previous encoded slice / image. The block. For example, let... and These are two fixed positive numbers. The inherited model parameters can come from positions in the previously encoded slice / image. The block. Another example is that inherited model parameters can come from the previous encoded slice / image. Some predefined locations of blocks. These locations can be Another example is that these locations could be... .
[0084] In some embodiments, neighboring blocks can be located at a predefined location. For example, the location could be as follows: Figure 9 The predefined position is A0. Predefined positions can also be A1, B0, B1, or B2. If the block at the predefined position is not an inter-frame block, no neighboring blocks are selected.
[0085] In some embodiments, when selecting a neighboring block, there may be a predefined list of locations. These locations are placed according to the inspection order. For example, location B could be... 0, A 0, B 1, A1 and B2, such as Figure 9 As shown. The selected neighboring block can be the first inter-frame block in the list. Select the L0 motion vector. If the L0 motion vector is unavailable, select the L1 motion vector. Another example: select the L1 motion vector. If the L1 motion vector is unavailable, select the L0 motion vector.
[0086] In some embodiments, if a co-location image has already been determined (e.g., using the methods described in the preceding paragraphs of this section), positions in a predefined list of positions are checked in a predefined checking order. The selected motion vector is the first co-location image whose reference image belongs to it. For example, position B could be B. 0, A 0, B 1, A1 and B2, such as Figure 9 As shown. For each position, first check the L0 motion vector, then check the L1 motion vector. That is, the checking order is (B 0, L0), (B 0, L1), (A 0, L0), (A 0, L1), ..., (B) 2,L1). In some embodiments, the L1 motion vector is checked first, and then the L0 motion vector is checked.
[0087] In some embodiments, the inherited model parameters may also be derived from the positions in the above-described patterns. The positions are centered at (x', y'), where x' = x + Δx and y' = y + Δy. The horizontal and vertical displacements Δx and Δy are determined based on the selected motion vector from neighboring blocks. For example, if the reference image and co-location image of the selected motion vector are the same image, Δx equals the horizontal portion of the selected motion vector, and Δy equals the vertical portion. If the horizontal or vertical portion of the selected motion vector is a fraction, Δx equals the rounded value of the horizontal portion, and Δy equals the rounded value of the vertical portion. The rounding method used may include, but is not limited to, rounding towards negative infinity, rounding towards positive infinity, rounding towards zero, or rounding to the nearest integer (e.g., rounding away from zero, rounding to the nearest integer, rounding to the nearest whole number, etc.). Another example is if the reference image and co-location image of the selected motion vector are not the same image. The reference image can be one of the images in the reference list, while the co-location image is conveyed in the image / slice header. Let tb be the POC distance between the current image and the reference image of the selected motion vector, td be the POC distance between the current image and the co-location image, and (mv_x, mv_y) be the selected motion vector. Δx = mv_x * (td / tb) and Δy = mv_y * (td / tb). If mv_x * (td / tb) or mv_y * (td / tb) is a fraction, Δx is equal to the rounded value of mv_x * (td / tb) or the rounded value of the horizontal portion of the selected motion vector, and Δy is equal to the rounded value of mv_y * (td / tb) or the rounded value of the vertical portion of the selected motion vector. The rounding method used may include, but is not limited to, the following: rounding to negative infinity, rounding to positive infinity, rounding to zero, or rounding to the nearest integer (e.g., rounding away from zero, rounding to zero, rounding to six, etc.).
[0088] In some embodiments, the inherited model parameters are derived by reconstructing samples using the luminance and chrominance of collocated blocks. Let the current block position be (x, y) and the block size be... When the inherited model comes from position (x', y'), the co-location block is the block located at position (x', y') in the co-location image, and the block size is... Another example is a co-location block, which can be a block located at position (x', y') in a co-location image, with a block size of [size missing]. Where m and n are fixed positive values. For example, a co-located block can be located at (x, y). Another example is if Δx and Δy are the L0 horizontal and vertical motion vectors of neighboring blocks, and the co-located image is an L0 reference image indicated by the L0 motion vectors of neighboring blocks, then the co-located block can be located within the co-located image. (x', y') can be a position within the pattern described in the preceding paragraph. For example, (x', y') could be... .
[0089] In some embodiments, the cross-component parameter model can be inherited from more than one previously encoded image. The cross-component parameter model can be inherited from any image in a set of N previously encoded images. An index can be signaled / parsed in the bitstream to indicate the selected image. The index ranges from 0 to N-1. In some embodiments, images with smaller Picture Order Count (POC) values compared to the current images are associated with smaller indices. In another sub-implementation, images with smaller quantization parameters (QP) compared to the current images are associated with smaller indices. In some embodiments, images with smaller quantization parameters (QP) are associated with smaller indices. In some embodiments, images with larger quantization parameters (QP) are associated with smaller indices.
[0090] e. Model of inheriting non-adjacent spatial neighbor blocks In some embodiments, the inherited model parameters may come from non-adjacent spatial neighbor blocks (i.e., blocks not adjacent to the current block). Models from blocks at predetermined locations are added to the cross-component prediction (CCP) merging candidate list in a predetermined order. In some embodiments, the predetermined location and predetermined order are the same as the non-adjacent spatial neighbor candidates in the inter-frame merging mode.
[0091] Figure 11A-11B This diagram illustrates a current block 1100 and its non-adjacent spatial neighboring blocks from which model parameters can be inherited. The diagram also illustrates the predetermined positions and their predetermined order. Figure 11A The non-adjacent spatial locations and their predetermined order are shown according to the first mode (mode 1). Figure 11B This displays the non-adjacent spatial locations and their predetermined order according to the second mode (Mode 2). The positions of the numbered blocks are predetermined positions. The numbers within each block indicate the predetermined order. Positions in Mode 1 are added to the CCP merge candidate list before those in Mode 2. The distance between each predetermined position is proportional to the width and height of the current block.
[0092] In some embodiments, there is a maximum number of models inherited from non-adjacent spatial neighbors that can be added to the CCP merging candidate list, and this maximum number is less than the number of predetermined positions.
[0093] In some embodiments, the current block position is set to (x, y) and the block size is... Define two sets of values. and as follows: and All values in the set are positive. Let x' = x + Δx and y' = y + Δy. Inherited model parameters can come from positions determined by x' and y'. For example, inherited model parameters can come from positions at... The block. For example, suppose... and These are two fixed positive numbers. The inherited model parameters can come from positions at... The block. For example, inherited model parameters can come from a block relative to the previous encoded slice / image. Some predefined locations of blocks. These locations can be For example, these locations could be... For example, the position could be (x',y'), (x',y'+h / 2), (x'+w / 2,y'), (x'+w / 2,y'+h / 2), (x'+w,y'), (x',y'+h), or (x'+w,y'+h).
[0094] In some embodiments, if the prediction mode of the current block is based on a block vector referencing the current image (e.g., Intra-Block Copy (IBC) or IntraTMP), then Δx and Δy can be set according to the horizontal and vertical block vectors of the current block. For example, Δx and Δy can be equal to the horizontal and vertical block vectors of the current block. In some embodiments, Δx and Δy can be set according to the horizontal and vertical block vectors of neighboring blocks. For example, Δx and Δy can be equal to the horizontal and vertical block vectors of neighboring blocks.
[0095] f. Inheriting model parameters from the history table In some embodiments, the inherited model parameters may come from a cross-component model history table. The history table stores CCM information for valid previously encoded blocks. A valid previously encoded block refers to any block containing valid CCM information. Cross-component models in the history table can be added to the CCP merge candidate list in a predefined order. In some embodiments, the order in which historical candidates are added can be from the beginning to the end of the table. In some embodiments, the order in which historical candidates are added can be from the end to the beginning of the table.
[0096] In some embodiments, a cross-component model history table may be maintained to store previous cross-component models (i.e., CCM information), and the cross-component model history table may be reset at the beginning of the current image, current tile, current tile, every M CTU row, or every N CTU, where N and M can be any value greater than 0. In some embodiments, the cross-component model history table may be reset at the end of the current image, current tile, current tile, current CTU row, or current CTU.
[0097] In some embodiments, multiple history tables are used to store different types of cross-component models. For example, the first history table is used to store a single model, and the second history table is used to store multiple models. Another example is that the first history table is used to store gradient models, and the second history table is used to store non-gradient models. Yet another example is that the first history table is used to store simple linear models (e.g., y = ax + b), and the second history table is used to store complex models (e.g., CCCM).
[0098] In some embodiments, when adding historical candidates to the CCP merge candidate list from multiple historical tables, the addition order can be from the beginning to the end of a table, and then the next historical table can be added in the same or reverse order.
[0099] g. Inherit from the fusion pattern Fusion mode refers to the mode of fusing two predictions to generate a final prediction. In chroma intra-frame fusion mode, a chroma intra-frame prediction not generated using cross-component prediction (CCP) codecs (e.g., CCLM, MMLM, CCCM) is fused with another chroma intra-frame prediction generated using a cross-component prediction codec. For example, a non-CCLM-coded intra-frame prediction and a CCLM-coded intra-frame prediction are fused together to obtain the final intra-frame prediction.
[0100] In some embodiments, when inheriting cross-component model parameters from blocks / positions encoded by the chroma intra-fusion mode, the model parameters used to obtain CCP-encoded intra-prediction are inherited and further refined. In some embodiments, in addition to inheriting and refining the CCP model parameters, the fusion weights and the encoding / decoding mode for non-CCP-encoded intra-prediction are also inherited. That is, the chroma intra-fusion mode is inherited.
[0101] h. Construct a candidate list In some embodiments, the CCP merging candidate list is constructed by adding candidates in a predefined order until a maximum number of candidates is reached. The added candidates may include, but are not limited to, all of the aforementioned candidates. For example, the predefined order may be spatially adjacent candidates, temporal candidates, spatially non-adjacent candidates, historical candidates, and then default candidates.
[0102] In some embodiments, if all predefined neighboring and historical candidates have been added but the maximum number of candidates has not been reached, some default candidates are added to the CCP merge candidate list until the maximum number of candidates is reached. In some embodiments, the default candidates may be CCLM models. Scaling parameters It comes from the set {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, ..., N / 8, -N / 8}, where N is a positive integer.
[0103] The offset parameter β can be Alternatively, it can be derived based on neighboring luminance and chroma samples. For example, if the average values of neighboring luminance and chroma samples are lumaAvg and chromaAvg, respectively, then... .
[0104] In some embodiments, the inclusion order of the default candidates may depend on the scaling parameter. The absolute value and sign. For example, according to The default candidate is added to the CCP merge candidate list in the following order: 0, 1 / 8, –1 / 8, 2 / 8, -2 / 8, …, N / 8, -N / 8. In some embodiments, the default candidate may be an earlier candidate refined with an incremental scaling parameter. This earlier candidate is a CCLM model. If the scaling parameter of the earlier candidate is… The scaling parameter for the preset candidates is ( +Δ For example, Δ It can be 1 / 8, -1 / 8, 2 / 8, -2 / 8, ..., N / 8, -N / 8, where N is a positive integer. The offset parameter β can be based on ( +Δ The default candidate is derived from the average of the luminance and chrominance samples of the current block and its neighbors. In some embodiments, an earlier candidate is the first CCLM candidate added to the CCP merge candidate list. In some embodiments, the inclusion order of the default candidate may depend on the refined Δ The absolute value and sign. For example, for Δ The default candidates are added to the CCP merge candidate list in the following order: 0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, ..., N / 8, -N / 8.
[0105] i. Reorder the candidates in the list Candidates in the CCP merge candidate list can be reordered to reduce syntactic overhead when indexing the selected candidate. In some embodiments, the reordering rules may rely on the encoding / decoding information of adjacent blocks. For example, if the adjacent upper or left block is encoded with MMLM, MMLM candidates in the list can be moved to the head of the current list. Similarly, if the adjacent upper or left block is encoded with a single-model LM (i.e., CCLM) or CCCM, single-model LM or CCCM candidates in the list can be moved to the head of the current list. Likewise, if the adjacent upper or left block uses GLM, GLM-related candidates in the list can be moved to the head of the current list.
[0106] In some embodiments, the reordering rule is based on model error (i.e. template cost) by comparing the error of the adjacent templates applied to the current block with the error of the reconstructed samples of the adjacent templates (i.e., the difference with the reconstructed samples of the adjacent templates).
[0107] For example, suppose the size of the adjacent template above the current block is The size of the template adjacent to the left of the current block is Suppose there are K models (which can be CCLM or 2-parameter GLM) in the current candidate list, and and These are the final scaling and offset parameters inherited from candidate k. The model error of candidate k through the adjacent template above is: in, and These are the reconstructed samples of luminance (e.g., after downsampling or applying GLM mode) and chrominance at position (i, j) in the upper template, respectively. and Similarly, the model error of candidate k through the left-adjacent template is: in and These are the reconstructed samples of luminance (e.g., after downsampling or applying GLM mode) and chrominance at position (m, n) in the left template, respectively. and Then the model error for candidate k is: After calculating the model errors for all candidates, a list of model errors may be obtained. Then, the video codec can reorder the candidate indices in the inheritance candidate list (i.e., the CCP merge candidate list) by sorting the model error list by ascending powers. For example, a model error could also be the SATD between the predicted chroma sample on the template generated by applying the candidate model to the luminance samples on the adjacent template and the reconstructed chroma sample on the adjacent template.
[0108] In some embodiments, if candidate k is predicted using CCCM, then and It can be defined as: in and These are the final filtering coefficients after inheriting from candidate k. P and B are the nonlinear and bias terms, respectively. In some embodiments, if the aforementioned neighbor template is unavailable, then... Similarly, if the adjacent template on the left is unavailable, then If neither template is available, the candidate index reordering method using the incorrect model should not be applied.
[0109] j. Signaling succession candidate index In some embodiments, an on / off flag is signaled to indicate whether the current block inherits cross-component model parameters from neighboring blocks. This flag can be signaled by CU / CB, by PU, by TU / TB, by color component, or by chroma color component. A higher-order syntax can be signaled in SPS, PPS, PH, or SH to indicate whether the current sequence, image, or slice is allowed to inherit cross-component model parameters from neighboring blocks.
[0110] In some embodiments, the maximum allowed number of candidates is signaled to indicate the maximum size of the candidate list to be merged. This number can be signaled by CU / CB, by PU, by TU / TB, by color component, or by chroma color component. A higher-order syntax can be signaled in SPS, PPS, PH, or SH to indicate whether the maximum allowed number of candidates for the current sequence, picture, or slice is permitted. The maximum allowed number of candidates can be shared with the maximum allowed number of candidates in the inter-frame merging mode.
[0111] In some embodiments, if the current block inherits cross-component model parameters from neighboring blocks, an inheritance candidate index is signaled. This index can be encoded and signaled using (e.g., truncate unary code, Exp-Golomb code, or fixed length code) and shared between the current Cb and Cr blocks. For example, the index can be signaled per chroma component. Another example is that one inheritance index is signaled for the Cb component and another for the Cr component. Yet another example is that the video codec can use a chroma intra-prediction syntax (e.g., IntraPredModeC[xCb][yCb]) to store the inheritance index.
[0112] In some embodiments, if the current block inherits cross-component model parameters from neighboring blocks, the current chroma intra-prediction mode (e.g., IntraPredModeC[xCb][yCb] as defined in the VVC standard) is temporarily set to a cross-component mode (e.g., CCLM_LA) during the bitstream parsing phase. Subsequently, during the prediction or reconstruction phase, a candidate list is derived, and the inherited candidate model is determined by the inherited candidate index. Once the inherited model is obtained, the codec information of the current block is updated based on the inherited candidate model. The codec information of the current block includes, but is not limited to, the prediction mode (e.g., CCLM_LA or MMLM_LA), related sub-mode flags (e.g., CCCM mode flag), prediction mode (e.g., GLM mode index), and current model parameters. Then, a prediction for the current block is generated based on the updated codec information.
[0113] k. Vector propagation cross-component model In some embodiments, after an encoding / decoding block, cross-component model (CCM) information for the current block is derived and stored. The stored CCM information can be referenced by subsequent encoding / decoding blocks. Subsequent encoding / decoding blocks can inherit CCM information from the current block. The definition of CCM information is described in Section III.b, “Inheriting CCM Information.” Stored CCM information can be inherited from, but is not limited to, candidates of the following types: spatial candidates, non-adjacent candidates, temporal candidates, and historical candidates, as described in the previous section.
[0114] In some embodiments, if the current block is coded using Cross-Component Prediction (CCP), the cross-component model used by the current block can be stored and referenced by subsequent codec blocks. When a block is coded using CCP, it means that the block uses a cross-component model to generate the block's predictions. The block may use a cross-component model inherited from neighboring blocks, a cross-component model derived from neighboring luma and chroma prediction / reconstruction sample values (e.g., CCLM, MMLM, CCCM, CCRM), or, in chroma fusion, a cross-component model means that chroma prediction is based on adding one or more cross-component prediction assumptions to one or more existing non-cross-component prediction assumptions, or any combination thereof.
[0115] In some embodiments, if the current block is not CCP encoded and is located in a non-intra-frame slice / picture, the CCM information of the current block can be derived by copying the CCM information of the collocated block. Assuming the current block is located at (x, y) and its size is w*h, the collocated block can be a block in a collocated picture at (x, y), (x + w, y + h), (x + w - 1, y + h - 1), or (x + w / 2, y + h / 2).
[0116] In some embodiments, if the current block is not CCP encoded and a block vector is available in the current block (e.g., the current luma block is encoded in IBC or IntraTMP mode, and the co-bit luma block is encoded in IBC or IntraTMP mode), the CCM information of the current block can be derived by copying the CCM information of the reference block located by the block vector.
[0117] Figure 12 This diagram conceptually illustrates an example of CCM information propagation based on block vectors. Several blocks A through H are shown. Blocks A, E, and G are encoded using a cross-component model (e.g., CCLM, MMLM, GLM, CCCM, Chroma Fusion). Block B is not CCP-encoded, and block vectors are available at block B. A reference block A is located by a block vector. The CCM information of the cross-component model used by reference block A is copied and stored for use by block B. In some embodiments, if a reference block located by a block vector is also not CCP-encoded but stores CCM information, the CCM information of the current block can be derived by copying the CCM information corresponding to the reference block.
[0118] For example, such as Figure 12As shown, current block C has a block vector available, and its reference block B is not CCP-encoded, but reference block B stores CCM information. The CCM information of block B is copied and stored for use by block C. The CCM information stored in block B is copied from block A. Therefore, the CCM information of block A is propagated to block C. In some embodiments, if the reference block located by the block vector is not CCP-encoded and does not store CCM information, then CCM information is not stored for the current block.
[0119] In some embodiments, if multiple block vectors are available for the current block (e.g., block vectors can be bidirectional, a block can have multiple IntraTMP block vectors, or the current chroma block shares a position with multiple luma blocks and more than one luma block has a block vector), in order to derive the CCM information of the current block, if only one reference block located by a block vector has CCM information, then the CCM information of that reference block is copied and stored for use by the current block. For example, as... Figure 12 As shown, assume block F has two block vectors and two reference blocks G and H. Block G has CCM information, while block H does not. The CCM information of block G is copied and stored for use by block F.
[0120] In some embodiments, if the current block has multiple block vectors and more than one of the multiple reference blocks located by the block vectors has CCM information, the video codec determines / derives the CCM information to be stored for the current block based on the CCM information of the multiple reference blocks. Figure 13 This diagram illustrates a current block with two block vectors used to identify two reference blocks with CCM information. As shown, a current block (block X) has two reference blocks (block Y and block Z) within the same current image 1300. Reference block Y is identified by block vector BV0 and has CCM information 1315. Reference block Z is identified by block vector BV1 and has CCM information 1325. Blocks Y and Z may or may not be CCP encoded. The video codec determines the CCM information 1305 of the current block X based on CCM information 1315 and CCM information 1325.
[0121] In different embodiments, the video codec determines the current block CCM information 1305 in different ways. For example, in some embodiments, the current block CCM information 1305 is derived by combining all or part of the CCM models of its reference blocks (e.g., CCM models 1315 and 1325 of reference blocks 1310 and 1320).
[0122] In some embodiments, if the current block (e.g., Figure 13 The current block X) has multiple block vectors, and multiple reference blocks (e.g., located by the block vectors) are located by the block vectors. Figure 13If more than one of the reference blocks (Y and Z) has CCM information, one of the reference blocks is selected according to a set of predefined rules. The CCM information of the selected reference block is then copied and stored for use by the current block. In some embodiments, a reference block encoded by CCP is selected. In some embodiments, a reference block encoded intra-frame is selected. In some embodiments, a reference block encoded inter-frame or IBC is selected.
[0123] In some embodiments, a reference block with the smallest distance to the current block is selected. The CCM information of the selected reference block is copied and stored for use by the current block. These reference blocks are located at (x... r y r ) and (x c y c The distance between the reference block and the current block can be calculated using Euclidean distance. Calculate. (x) r y r ) and (x c y c The reference block and the current block can be located at their top-left, top-right, bottom-left, bottom-right, or center positions. The distance metric can also be Manhattan distance or Minkowski distance.
[0124] For some embodiments, the one with the minimum horizontal distance, |x r -x c The CCM information of the selected reference block is copied and stored for use by the current block. In some embodiments, the block with the smallest vertical distance, |y, is selected. r -y c The reference block is selected. The CCM information of the selected reference block is copied and stored for use by the current block.
[0125] In some embodiments, the previously described rules can be used in combination, and it is not necessary to apply all of the previously described rules. For example, selecting a CCP-encoded reference block. If there is more than one CCP-encoded reference block, the CCP-encoded reference block with the shortest distance to the current block is selected. If there is more than one CCP-encoded reference block with the smallest distance to the current block, the one with the smallest horizontal distance, |x r -x c |, the reference block. Another example is selecting a reference block encoded by the CCP. If there is more than one CCP-encoded reference block, the CCP-encoded reference block with the shortest distance to the current block is selected. If there is more than one CCP-encoded reference block with the smallest distance to the current block, the one with the smallest vertical distance, |y, is selected. r -y c The reference block is selected. The CCM information of the selected reference block is copied and stored for use by the current block.
[0126] In some embodiments, if the current block is not CCP-coded and motion vectors are available in the current block (e.g., the current luma block is inter-frame coded), the CCM information of the current block can be derived by copying the CCM information of a reference block in a reference image located by the motion vector of the current block.
[0127] Figure 14 An example of motion vector-based CCM information propagation is illustrated. The figure shows several blocks A through H. Blocks A, E, and G are encoded using a cross-component model (e.g., CCLM, MMLM, GLM, CCCM). As shown, block B is not encoded with CCP, and motion vectors are available at block B. A reference block A is located using motion vectors. The CCM information of the reference block A, using a cross-component model, is copied and stored for use by block B.
[0128] In some embodiments, if the reference block located by motion vectors is also not CCP-encoded, but the reference block stores CCM information, the CCM information of the current block can be derived by copying the CCM information corresponding to the reference block. Figure 14 In the example, block C has motion vectors available, and its reference block B is not CCP-encoded but stores CCM information. The CCM information of block B is copied and stored for use by block C. The CCM information stored in block B is copied from block A. Therefore, the CCM information of block A is propagated to block C. In some embodiments, if the reference block located by the motion vector is not CCP-encoded and does not store CCM information, then CCM information is not stored for the current block.
[0129] In some embodiments, if the current block is inter-frame coded using bidirectional prediction, in order to derive the CCM information of the current block, if only one of the reference blocks located by motion vectors has CCM information, then the CCM information of the reference block with CCM information is copied and stored for use by the current block. For example, as Figure 14 As shown, block F is inter-frame coded using bidirectional prediction. The two reference blocks located by the motion vectors (MV0 and MV1) of block F are block G and block H. Block G stores CCM information, while block H does not. The CCM information of block G is copied and stored for use by block F.
[0130] In some embodiments, if the current block is inter-frame coded using bidirectional prediction and both reference blocks located by motion vectors store CCM information, the video codec determines / derives the CCM information to be used for the current block based on the CCM information of the two or more reference blocks. Figure 15This diagram illustrates a current block with two motion vectors used to identify two reference blocks, each with its own CCM information. As shown, the current block (block P) in current frame 1502 is inter-coded using bidirectional prediction. Two reference blocks (block Q and block R) are located by motion vectors from different reference frames 1501 and 1503. Reference block Q is located by block vector MV0 and has CCM information 1515. Reference block R is located by block vector MV1 and has CCM information 1525. Reference blocks Q and R may or may not be coded using cross-component prediction. The video codec determines the CCM information 1505 of the current block P based on CCM information 1515 and CCM information 1525.
[0131] The video codecs in different embodiments determine the current block CCM information 1505 in different ways. For example, in some embodiments, the current block CCM information 1505 is derived by combining all or a subset of the CCM models 1515 and 1525 of reference blocks Q and R. In another example, in some embodiments, one of the reference blocks is selected according to a set of predefined rules. The CCM information of the selected reference block is then copied and stored for use in the current block.
[0132] In some embodiments, to select reference blocks from one or more reference blocks to obtain CCM information, reference blocks coded across component prediction are selected. In some embodiments, reference blocks coded intra-frame prediction are selected. In some embodiments, reference blocks coded inter-frame or IBC are selected.
[0133] In some embodiments, a reference block whose reference image (i.e., the image containing the reference block) has a small point of interest (POC) distance from the current image is selected. The CCM information of the selected reference block is then copied and stored for use by the current block. Figure 15 In the example, block P is inter-frame coded using bidirectional prediction. Two reference blocks Q and R are located by motion vectors and both store CCM information. The POC of reference image 1501 containing reference block Q is N1, the POC of the current image 1502 is N2, and the POC of reference image 1503 is N3. If |N1-N2| is less than |N3-N2|, block Q is selected, and the CCM information 1515 of block Q is copied and stored as the current block CCM information 1505 for block P.
[0134] In some embodiments, a reference block whose reference image has a small / minimum QP difference from the current image is selected. The CCM information of the selected reference block is then copied and stored for use by the current block. Figure 14In the example, block F in image 1401 is inter-coded using bidirectional prediction. The two reference blocks located by motion vectors are block G in image 1400 and block H in image 1402. Assume that both blocks G and H store CCM information. Assume the QP values for images 1400, 1401, and 1402 are 27, 32, and 33, respectively. Since |33-32| is less than |27-32|, block H is selected, and its CCM information is copied and stored for use by block F.
[0135] In some embodiments, a reference block with a smaller QP value is selected as its reference image (therefore in Figure 14 In the example, block G is selected because image 1400 has a smaller QP value of 27. In some embodiments, a reference block with a larger QP value is selected (therefore in...). Figure 14 In the example, since image 1402 has a large QP value of 33, block H will be selected.
[0136] In some embodiments, a reference block indicated by an L0 motion vector is selected. In some embodiments, a reference block indicated by an L1 motion vector is selected. In some embodiments, the previously described rules can be used in combination, and it is not necessary to apply all of the previously described rules. For example, a reference block that is cross-component predictive coding is selected. If both blocks are cross-component predictive coding, the block whose reference image has a smaller POC distance from the current image is selected. If both blocks are cross-component predictive coding and have the same POC distance from the current image, the reference block whose reference image has a smaller QP difference from the current image is selected. If both blocks are cross-component predictive coding, have the same POC distance from the current image, and have the same QP difference from the current image, the reference block whose reference image has a smaller QP value is selected. Another example is that a block whose reference image has a smaller POC distance from the current image is selected. If both blocks have the same POC distance from the current image and have the same QP difference, the reference block whose reference image has a smaller QP value is selected.
[0137] In some embodiments, if the current block is inter-coded or block vectors are available in the current block, the CCM information of the current block can be obtained by copying the CCM information of a reference block located by motion vectors or block vectors. For example, as... Figure 14 As shown, current block C has block vectors available, and its reference block B has motion vectors available. Block B's CCM information is copied from block A. Then, the CCM information of block B is copied to block C. Therefore, the CCM information of block A is propagated to current block C.
[0138] L. Inheriting multiple cross-component models In some embodiments, if the current candidate list size is N, the video codec may select k candidates from a total of N candidates (where k ≤ N). The k cross-component models can be combined into a final cross-component model by a weighted average of their corresponding model parameters. For example, if a cross-component model has M parameters, the j-th parameter of the final cross-component model is the weighted average of the j-th parameters of the k selected candidates, where j is 1…M. The final prediction is then generated by applying the final cross-component model to the corresponding brightness reconstruction samples. For example, in some embodiments, if two candidate models are… as well as The final cross-component model is... , in Weight values can be pre-set or implicitly derived from the costs of adjacent templates, and It is the x-th candidate model of the y-th generation.
[0139] In some embodiments, using the template cost defined in Section III.i (candidates in the reordering list), the template cost associated with two candidate models is expressed as follows: and ,in yes .
[0140] In some embodiments, the video codec can combine multiple cross-component models into a single final cross-component model. For example, the video codec can select one model from a first candidate and another from a second candidate to form a multi-model mode (e.g., MMLM or MM-CCCM). The selected candidates can be candidates encoded with CCLM / MMLM / GLM / CCCM. The multi-model classification threshold can be the average of the offset parameters of the two selected modes (e.g., the offset in CCLM / ...). or in CCCM or In some embodiments, the classification threshold is set as the average of the neighboring luminance and chrominance samples of the current block.
[0141] m. Store the time model in the index table. In some embodiments, CCM information for previously encoded slices / images is stored in a table, and an image-level index buffer is created to store the indexes of the table. The size of the index buffer is the same as that of the image. When CCM information is referenced at a position (x, y) within the co-bit image (as described in Section III.d, “Inheriting Temporal Proximity Model Parameters”), the index value is retrieved from the position (x, y) of the co-bit image's index buffer. The item indicated by the index value is obtained from the table as the CCM information to be referenced. If the index value indicates that no CCM information is available at position (x, y), the CCM information is not referenced.
[0142] In some embodiments, the table storing CCM information is an image-level table. CCM information for each image is stored in a separate table. In some embodiments, a single table is created to store CCM information for all images. In some embodiments, a table is created for each time ID. CCM information for layers with the same time ID is stored in the same table. In some embodiments, several tables are used to store the CCM information for a single image. An image can be divided into several regions, each corresponding to its own table.
[0143] In some embodiments, after storing the CCM information for position (x, y) of the current encoded / decoded image into the corresponding table, the table index value of the CCM information is stored at position (x, y) in the index buffer of the current encoded / decoded image. If no CCM information is available at position (x, y) (e.g., when the CU covering position (x, y) is not encoded using any CCLM, MMLM, CCCM, CCCM multi-model, chroma fusion, or other cross-component model), the value of position (x, y) in the index buffer is set to indicate that no CCM information is available.
[0144] In some embodiments, if no CCM information is available at the current encoding / decoding image position (x, y), the index value of the position (x, y) in the index buffer of the co-bit image can be stored at the current encoding / decoding image index buffer position (x, y).
[0145] In some embodiments, if no CCM information is available at position (x, y) of the current encoded / decoded image, and a block vector (Δx, Δy) is available at position (x, y) (e.g., the block at position (x, y) can be IBC encoded or IntraTMP encoded, or the co-bit luma block is encoded in IBC or IntraTMP mode), then the index value of position (x+Δx, y+Δy) in the index buffer can be stored at position (x, y) of the current encoded / decoded image's index buffer. In another embodiment, if no CCM information is available at position (x, y) of the current encoded / decoded image, and multiple block vectors are available, (Δxi, Δyi), 0
[0146] In some embodiments, if no CCM information is available at position (x, y) of the current encoded / decoded image, and the block at position (x, y) is inter-frame coded with a motion vector of (Δx, Δy), then the index value of position (x+Δx, y+Δy) in the index buffer of the reference image (also identified by the motion vector) can be stored at position (x, y) in the index buffer of the current encoded / decoded image. In some embodiments, if no CCM information is available at position (x, y) of the current encoded / decoded image, and the block at position (x, y) is inter-frame coded with a bidirectional motion vector of (Δxi, Δyi), i = 0 or 1, then the index value of position (x+Δxi, y+Δyi) in the index buffer of the reference image can be used to store the index value of position (x, y) in the index buffer of the current encoded / decoded image.
[0147] In some embodiments, if no CCM information is available at position (x, y) of the current encoded / decoded image, and the block at position (x, y) is inter-frame coded with a motion vector of (Δx, Δy), the index value of a position (x+Δx, y+Δy) in the index buffer of the reference image can be used to retrieve CCM information from the table corresponding to the reference image (also identified by the motion vector). The retrieved CCM information is then stored in the table corresponding to the current encoded / decoded image. The index value of the retrieved CCM information can be used to store the position (x, y) in the index buffer of the current encoded / decoded image.
[0148] In some embodiments, if no CCM information is available at position (x, y) of the current encoded / decoded image, and the block at position (x, y) is inter-frame coded and the motion vector is bidirectional (Δxi, Δyi), i = 0 or 1, the index value of a position (x + Δxi, y + Δyi) in the index buffer of the reference image (also identified by the motion vector) can be used to retrieve CCM information from the table corresponding to the reference image. The retrieved CCM information is then stored in the table corresponding to the current encoded / decoded image. The index value of the retrieved CCM information can be used to store the position (x, y) in the index buffer of the current encoded / decoded image. In some embodiments, a set of predefined rules can be used to determine how to select one of the motion vectors (i.e., retrieve the index value located by the motion vector (Δxi, Δyi) and store the corresponding cross-component model in the current table). These rules can be the same as those described in Section III.k, “Cross-component Models for Vector Propagation,” for selecting the reference block located by the motion vector.
[0149] In some embodiments, when a CCM entry is deleted from the table, all values in the index buffer indicating the use of the CCM entry to be deleted are reset to indicate that no CCM entry is available. Assuming the index value of the CCM entry to be deleted is N, all values in the index buffer greater than N will be decremented by 1.
[0150] In some embodiments, the table used to store CCM information has a maximum size limit. This maximum size limit can be indicated by sending high-level syntax in SPS, PPS, PH, or SH. If the table has reached its maximum size when attempting to store new CCM information, the new CCM information will not be stored. In some embodiments, if the table has reached its maximum size when attempting to store new CCM information, the oldest stored CCM information will be deleted to free up space in the table.
[0151] In some embodiments, the table may be reset at the start of encoding / decoding the IDR image. In some embodiments, the table may be reset after encoding / decoding the IDR image. In some embodiments, the table may be reset at the start of encoding / decoding the CRA image. In some embodiments, the table may be reset after encoding / decoding the CRA image. In some embodiments, the reset mechanism may be the same as that used in the parameter set or reference image.
[0152] In some embodiments, an index stored in the index buffer can only be referenced by units greater than or equal to the smallest decoding unit. For example, if the smallest decoding unit is 4x4, an index can be referenced by an 8x8 grid. That is, an 8x8 block has the same index value. To retrieve an index value at position (x, y), the position (x, y) can be rounded to a point on the grid (e.g., (x>>3)<<3, (y>>3)<<3) or to the nearest point on the grid.
[0153] In some embodiments, the CCM information to be stored in the table can be explicitly indicated in the SPS, PPS, PH, or SH of the bitstream. The corresponding position of the CCM information can also be indicated.
[0154] Any of the methods proposed above can be implemented in the encoder and / or decoder. For example, any of the proposed methods can be implemented in the inter-frame / intra-frame / prediction module of the encoder, and / or in the inter-frame / intra-frame / prediction module of the decoder. Alternatively, any of the proposed methods can be implemented as circuitry connected to the inter-frame / intra-frame / prediction module of the encoder and / or decoder to provide the information required by the inter-frame / intra-frame / prediction module.
[0155] IV. Example Video Encoder Figure 16A video encoder 1600 that may implement cross-component prediction is illustrated. As shown, the video encoder 1600 receives an input video signal from a video source 1605 and encodes the signal into a bitstream 1695. The video encoder 1600 has multiple components or modules for encoding the signal from the video source 1605, including at least some of the following components: self-transform module 1610, quantization module 1611, inverse quantization module 1614, inverse transform module 1615, intra-frame estimation module 1624, intra-frame prediction module 1625, motion compensation module 1630, motion estimation module 1635, loop filter 1645, reconstructed image buffer 1650, MV buffer 1665, MV prediction module 1675, and entropy encoder 1690. Motion compensation module 1630 and motion estimation module 1635 are part of inter-frame prediction module 1640. Intra-frame prediction module 1625 and intra-frame prediction estimation module 1624 are part of current-image prediction module 1620, which uses current-image reconstruction samples as prediction reference samples for the current block.
[0156] In some embodiments, modules 1610-1690 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 1610-1690 are modules of hardware circuitry implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 1610-1690 are depicted as separate modules, some modules may be combined into a single module.
[0157] Video source 1605 provides a raw video signal that represents the pixel data of each picture frame without compression. Subtractor 1608 calculates the difference between the raw picture pixel data from video source 1605 and the predicted pixel data from motion compensation module 1630 or intra-frame prediction module 1625 as a prediction residual 1609. Transform module 1610 transforms the difference (or residual pixel data or residual signal 1608) into transform coefficients 1616 (e.g., by performing a discrete cosine transform, or DCT). Quantization module 1611 quantizes the transform coefficients into quantized data (or quantization coefficients) 1612, which is encoded into a bitstream 1695 by entropy encoder 1690.
[0158] The inverse quantization module 1614 performs inverse quantization on the quantized data (or quantization coefficients) 1612 to obtain transform coefficients, and the inverse transform module 1615 performs inverse transform on the transform coefficients to generate a reconstruction residual 1619. The reconstruction residual 1619 is added to the predicted pixel data 1613 to generate reconstructed pixel data 1617. In some embodiments, the reconstructed pixel data 1617 is temporarily stored in an online buffer 1627 (or an intra-frame prediction buffer) for intra-frame prediction and spatial MV prediction. The reconstructed pixels are filtered by a loop filter 1645 and stored in a reconstructed image buffer 1650. In some embodiments, the reconstructed image buffer 1650 is external storage to the video encoder 1600. In some embodiments, the reconstructed image buffer 1650 is internal storage to the video encoder 1600.
[0159] Intra-frame estimation module 1624 performs intra-frame prediction based on reconstructed pixel data 1617 to generate intra-frame prediction data. The intra-frame prediction data is provided to entropy encoder 1690 to be encoded into bitstream 1695. The intra-frame prediction data is also used by intra-frame prediction module 1625 to generate predicted pixel data 1613.
[0160] The motion estimation module 1635 performs inter-frame prediction by generating MVs (Motion Values) to reference reference pixel data of previously decoded frames stored in the reconstructed image buffer 1650. These MVs are provided to the motion compensation module 1630 to generate predicted pixel data.
[0161] Instead of encoding the complete actual MV in the bitstream, the video encoder 1600 uses MV prediction to generate a predicted MV, and the difference between the MV used for motion compensation and the predicted MV is encoded as residual motion data and stored in the bitstream 1695.
[0162] The motion vector prediction module 1675 generates predicted motion vectors, i.e., motion-compensated motion vectors used for motion compensation, based on reference motion vectors previously generated for encoding previous image frames. The motion vector prediction module 1675 retrieves reference motion vectors from previous image frames in the motion vector buffer 1665. The video encoder 1600 stores the motion vectors generated for the current image frame in the motion vector buffer 1665 as reference motion vectors for generating the predicted motion vectors.
[0163] The motion vector prediction module 1675 uses reference motion vectors to create predicted motion vectors. The predicted motion vectors can be calculated through spatial motion vector prediction or temporal motion vector prediction. The difference (residual motion data) between the predicted motion vectors and the motion-compensated motion vectors (MC MVs) of the current frame is encoded into bitstream 1695 by the entropy encoder 1690.
[0164] The entropy encoder 1690 uses entropy encoding techniques such as Context-Adaptive Binary Arithmetic Encoding and Decoding (CABAC) or Huffman coding to encode various parameters and data into a bitstream 1695. The entropy encoder 1690 encodes various header elements, flags, quantization transform coefficients 1612, and residual motion data as syntax elements into the bitstream 1695. The bitstream 1695 is then stored in a storage device or transmitted to the decoder via a communication medium such as a network.
[0165] The loop filter 1645 filters or smooths the reconstructed pixel data 1617 to reduce encoding / decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering or smoothing operations performed by the loop filter 1645 include deblocking filtering (DBF), sample adaptive offset (SAO), and / or adaptive loop filtering (ALF). In some embodiments, luma-mapped chroma scaling (LMCS) is performed before the loop filter.
[0166] Figure 17 This describes a portion of the video encoder 1600 used to implement the propagation of CCM information from multiple reference blocks. Current image predictions (intra-frame prediction, IBC, and other predictions by referencing the current image) are used to generate a reconstruction 1715 for the luma component. A cross-component model 1710 can be applied to the reconstruction 1715 to generate a cross-component predictor 1725 for the chroma component. The cross-component predictor 1725 is then included in the predicted pixel data 1613. The cross-component model 1710 can also be stored in the CCM storage 1735 for use in subsequent coding blocks.
[0167] The cross-component model 1710 can be generated by the model builder 1705 based on reference samples and / or current samples (within and / or around the current block and / or reference block) retrieved from the reconstructed image buffer 1650 and / or the line buffer 1627. Section I above describes several types of cross-component models that can be used as the cross-component model 1710.
[0168] The cross-component model 1710 can also be provided by the CCM selection module 1730, which provides cross-component model (CCM) information or other cross-component prediction (CCP) information inherited from previously encoded blocks. In some embodiments, the CCM selection module 1730 can provide a list of CCP merging candidates, which may include models from spatial, temporal, and non-neighboring neighbors, and / or from a history table and / or from default candidates. The CCM selection module 1730 can provide the CCM information of the selected candidates to the current block. The propagated CCM information can be used as the cross-component model 1710 of the current block, or stored in the CCM storage 1735 for further propagation.
[0169] CCM propagation module 1740 can propagate CCM information of multiple reference blocks of a reference block to the reference block by updating the contents of CCM storage 1735. Reference blocks of a reference block can be identified by the motion vector and / or block vector of the reference block. When multiple further reference blocks located by multiple MVs or BVs all / have CCM information, CCM propagation module 1740 can combine the CCM information from the multiple further reference blocks. CCM propagation module 1740 can also select one further reference block from the multiple further reference blocks to obtain CCM information. In some embodiments, a set of predefined rules can be applied to select one further reference block from the multiple further reference blocks. Examples of these rules include: selecting the further reference block that is spatially or temporally closest to the reference block; selecting the further reference block that is most similar to the quantization parameters of the reference block / reference image; selecting the further reference block with the largest or smallest quantization parameter; selecting a further reference block that is internally encoded, or inter-frame encoded, or IBC encoded; selecting a further reference block located by L0 motion or L1 motion, etc.
[0170] CCM storage 1735 represents (or is constituted by) any form of storage for storing CCM information, including cross-component models generated by model builder 1705. CCM storage 1735 can be part of a block-level buffer, part of a picture-level buffer, or a CCM table (or index table) for different pictures, different time IDs, or different regions of different pictures. A CCM table can be associated with a video picture and store CCM information or other CCP information related to the associated picture. A CCM table can also be associated with a time ID and store CCM information or other CCP information for the picture with the associated time ID. A CCM table can be associated with a region in a video picture and store CCM information or other CCP information related to the associated region. The stored CCM information and / or CCP information can be used as merge pattern candidates to be inherited by subsequent blocks.
[0171] Each cross-component model (CCM) table has a corresponding index buffer used to map locations in the image to locations in the CCM table. To retrieve CCM information and / or cross-component prediction (CCP) information for a selected reference block (or a selected candidate block), the encoder identifies the CCM table and corresponding index buffer for the selected reference block (based on the image, time ID, or region in the image), and then uses the candidate's location in the image to find its index in the identified index buffer. This index is then used to access the selected CCM information and / or CCP information in the CCM table.
[0172] Figure 18A process 1800 is conceptually illustrated for propagating CCM information from multiple reference blocks when encoding a block. This process selects from multiple reference blocks containing CCM information during block encoding. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing encoder 1600 execute process 1800 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing encoder 1600 executes process 1800.
[0173] In step 1810, the encoder receives data to encode the pixels of a current block of a current image of the video. This current block has a first color block (e.g., a luma component block) and a second color block (e.g., a chroma component block). In step 1820, the encoder generates a reconstruction of the first color block.
[0174] In step 1830, the encoder inherits a cross-component model from a reference block; the cross-component model is propagated to the reference block based on cross-component information from two or more further reference blocks.
[0175] These two or more further reference blocks may be located using more than one motion vector or block vector of the reference block (e.g., when the reference block is a bidirectional inter-frame prediction block). In some embodiments, the encoder derives the cross-component model by combining all or part of the CCM model of the two or more further reference blocks.
[0176] In some embodiments, the encoder provides a cross-component model to the reference block by propagating the cross-component model by selecting one block from two or more further reference blocks. This further reference block (the selected further reference block) may be selected according to a set of predefined rules. The reference block may or may not be cross-component predictive coded. The encoder may select a further reference block as the selected further reference block by identifying a block that is cross-component predictively coded. The encoder may select the only further reference block among two or more further reference blocks that has a cross-component model. The encoder may select a further reference block by identifying a block that is cross-frame predictively coded (e.g., IBC mode), or by identifying a block that is cross-frame predictively coded, or by identifying a block that is cross-frame predictively coded.
[0177] In some embodiments, the selected further reference blocks are selected by identifying the block with the shortest spatial distance to the reference block among the further reference blocks. The spatial distance may be a vertical or horizontal distance, or a Euclidean distance, or a Manhattan distance, or a Minkowski distance. In some embodiments, the selected further reference blocks are selected by identifying the block with the shortest temporal distance to the reference block among one or more further reference blocks. The temporal distance of the block may be determined based on the POC containing each further reference image and the POC of the reference image.
[0178] Further reference blocks may be selected by identifying the block in each of the further reference blocks that has the QP closest to the quantization parameter of the reference block, or by identifying the block in each of the further reference blocks that has the largest QP, or by identifying the block in each of the further reference blocks that has the smallest QP, or by identifying the block in each of the further reference blocks that is indicated by the L0 (or L1) motion vector.
[0179] In step 1840, the encoder applies the inherited cross-component model to the reconstruction of the first color patch to generate a cross-component prediction for the second color patch. The encoder may use the generated cross-component prediction to encode the current patch in step 1850 (by generating a prediction residual), or in step 1860, store the determined cross-component model to encode and decode subsequent patches.
[0180] V. Example Video Decoder In some embodiments, the encoder may signal (or generate) one or more syntax elements in the bitstream, so that the decoder can parse one or more syntax elements from the bitstream.
[0181] Figure 19 An example of a video decoder 1900 that may implement cross-component prediction is illustrated. As shown, the video decoder 1900 is an image decoding or video decoding circuit that receives a bitstream 1995 and decodes the contents of the bitstream into pixel data of picture frames for display. The video decoder 1900 includes several components or modules for decoding the bitstream 1995, including some components selected from inverse quantization module 1911, inverse transform module 1910, intra-frame prediction module 1925, motion compensation module 1930, loop filter 1945, decoded picture buffer 1950, MV buffer 1965, MV prediction module 1975, and parser 1990. Motion compensation module 1930 is part of inter-frame prediction module 1940. Intra-frame prediction module 1925 is part of current picture prediction module 1920, which uses the reconstructed samples of the current picture as reference samples for predicting the current block.
[0182] In some embodiments, modules 1910-1990 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 1910-1990 are modules of hardware circuitry implemented by electronic devices of one or more ICs. Although modules 1910-1990 are depicted as independent modules, some modules may be combined into a single module.
[0183] The parser 1990 (or entropy decoder) receives the bitstream 1995 and performs initial parsing according to the syntax defined by the video-codec or image-codec standard. The parsed syntax elements include various header elements, flags, and quantization data (or quantization coefficients) 1912. The parser 1990 uses entropy-based encoding and decoding techniques such as Context-Adaptive Binary Arithmetic Codec (CABAC) or Huffman coding to parse the various syntax elements.
[0184] Inverse quantization module 1911 performs inverse quantization on quantized data (or quantization coefficients) 1912 to obtain transform coefficients, and inverse transform module 1910 performs inverse transform on transform coefficients 1916 to generate reconstructed residual signal 1919. Reconstructed residual signal 1919 is added to predicted pixel data 1913 from intra-frame prediction module 1925 or motion compensation module 1930 to generate decoded pixel data 1917. Decoded pixel data is filtered by loop filter 1945 and stored in decoded image buffer 1950. In some embodiments, decoded image buffer 1950 is external storage to video decoder 1900. In some embodiments, decoded image buffer 1950 is internal storage to video decoder 1900.
[0185] Intra-prediction module 1925 receives intra-prediction data from bitstream 1995 and generates predicted pixel data 1913 from decoded pixel data 1917 stored in decoded image buffer 1950 based on this data. In some embodiments, decoded pixel data 1917 is also stored in line buffer 1927 (or intra-prediction buffer) for intra-image prediction and spatial MV prediction.
[0186] In some embodiments, the contents of the decoded image buffer 1950 are used for display. The display device 1905 either retrieves the content directly from the decoded image buffer 1950 for display, or retrieves the contents of the decoded image buffer into a display buffer. In some embodiments, the display device receives pixel values from the decoded image buffer 1950 via pixel transfer.
[0187] The motion compensation module 1930 generates predicted pixel data 1913 from the decoded pixel data 1917 stored in the decoded image buffer 1950 based on the motion compensation MV (MC MV). These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 1995 to the predicted MV received from the MV prediction module 1975.
[0188] The MV prediction module 1975 generates a predicted MV based on a reference MV used for decoding a previous image frame, such as a motion-compensated MV used for motion compensation. The MV prediction module 1975 retrieves the reference MV of the previous image frame from the MV buffer 1965. The video decoder 1900 stores the motion-compensated MV used for decoding the current image frame in the MV buffer 1965 as a reference MV for generating the predicted MV.
[0189] Loop filter 1945 filters or smooths the decoded pixel data 1917 to reduce encoding / decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering or smoothing operations performed by loop filter 1945 include deblocking filter (DBF), sample adaptive offset (SAO), and / or adaptive loop filter (ALF). In some embodiments, luma-mapped chroma scaling (LMCS) is performed before the loop filter.
[0190] Figure 20 This describes a portion of the video decoder 1900 used to implement the propagation of CCM information from multiple reference blocks. Current image predictions (intra-frame prediction, IBC, and other predictions made with reference to the current image) are used to generate a reconstruction 2015 for the luma component. A cross-component model 2010 can be applied to the reconstruction 2015 to generate a cross-component predictor 2025 for the chroma component. The cross-component predictor 2025 is then included in the predicted pixel data 1913. The cross-component model 2010 can also be stored in the CCM storage 2035 for use in subsequent coded blocks.
[0191] The cross-component model 2010 can be generated by the model builder 2005 based on reference samples and / or current samples (within and / or around the current block and / or reference block) retrieved from the reconstructed image buffer 1950 and / or the line buffer 1927. Section I above describes several types of cross-component models that can be used as the cross-component model 2010.
[0192] The cross-component model 2010 can also be provided by the CCM selection module 2030, which provides cross-component model (CCM) information or other cross-component prediction (CCP) information inherited from previous coded blocks. In some embodiments, the CCM selection module 2030 can provide a list of CCP merging candidates, which may include models from spatial, temporal, and non-neighboring neighbors, and / or from a history table and / or from default candidates. The CCM selection module 2030 can provide the CCM information of the selected candidates to the current block. The propagated CCM information can be used as the cross-component model 2010 of the current block, or stored in the CCM storage 2035 for further propagation.
[0193] The CCM propagation module 2040 can propagate CCM information from multiple reference blocks of a reference block to another reference block by updating the contents of the CCM storage 2035. Reference blocks of a reference block can be identified by the motion vector and / or block vector of the reference block. When multiple further reference blocks located by multiple MVs or BVs all / both have CCM information, the CCM propagation module 2040 can combine the CCM information from the multiple further reference blocks. The CCM propagation module 2040 can also select one further reference block from the multiple further reference blocks to obtain CCM information. In some embodiments, a set of predefined rules can be applied to select one further reference block from the multiple further reference blocks. Examples of these rules include: selecting the further reference block that is spatially or temporally closest to the reference block; selecting the further reference block that is most similar to the quantization parameters of the reference block / reference image; selecting the further reference block with the largest or smallest quantization parameter; selecting the further reference block encoded by intra-frame prediction, inter-frame prediction, or IBC; selecting the further reference block located by L0 motion or L1 motion, etc.
[0194] CCM storage 2035 represents (or is constituted by) any form of storage for storing CCM information, including cross-component models generated by model builder 2005. CCM storage 2035 can be part of a block-level buffer, part of a picture-level buffer, or a CCM table (or index table) for different pictures, different time IDs, or different regions of different pictures. A CCM table can be associated with a video picture and store CCM information or other CCP information associated with that picture. A CCM table can also be associated with a time ID and store CCM information or other CCP information for pictures with the relevant time ID. A CCM table can be associated with a region in a video picture and store CCM information or other CCP information associated with that region. The stored CCM information and / or CCP information can be inherited by subsequent blocks as merge mode candidates.
[0195] Each CCM table has a corresponding index buffer used to map locations in the image to their corresponding locations in the CCM table. To retrieve CCM and / or CCP information for a selected reference block (or a selected candidate), the decoder identifies the CCM table and corresponding index buffer for the selected reference block (based on the image, time ID, or region within the image), and then uses the candidate's location in the image to look up its index in the identified index buffer. This index is then used to access the selected CCM and / or CCP information in the CCM table.
[0196] Figure 21 A process 2100 is conceptually illustrated for propagating CCM information from multiple reference blocks during the decoding of a block. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing decoder 1900 execute process 2100 by executing instructions stored in a readable computing medium. In some embodiments, an electronic device implementing decoder 1900 executes process 2100.
[0197] The decoder receives (in step 2110) data to decode the pixels of the current block of the current image as video. The current block has a first color block (e.g., a luma component block) and a second color block (e.g., a chroma component block). The decoder generates (in step 2120) a reconstruction of the first color block.
[0198] The decoder inherits (in step 2130) the cross-component model from the reference block; the cross-component model is based on the cross-component information of two or more further reference blocks propagated to the reference block.
[0199] Two or more further reference blocks can be located by more than one motion vector or block vector of the reference block (e.g., when the reference block is a bidirectional inter-frame prediction block). In some embodiments, the decoder derives the cross-component model by combining all or a subset of the CCM models of the two or more further reference blocks.
[0200] In some embodiments, the decoder provides a cross-component model to the reference block by propagating the cross-component model by selecting one block from two or more further reference blocks. A further reference block can be selected according to a set of predefined rules. The reference block may or may not be cross-component predictively encoded. The decoder can select a further reference block by identifying a block that is cross-component predictively encoded. The decoder can select a further reference block that is unique among two or more further reference blocks and has a cross-component model. The decoder can select a further reference block by identifying a block encoded by the current picture reference (e.g., IBC mode), or by identifying a block that is intra-predictively encoded, or by identifying a block that is inter-predictively encoded.
[0201] In some embodiments, a further reference block can be selected by identifying the block with the shortest spatial distance to the reference block among one or more further reference blocks. The spatial distance can be a vertical or horizontal distance, or a Euclidean distance, or a Manhattan distance, or a Minkowski distance. In some embodiments, a further reference block can be selected by identifying the block with the shortest temporal distance to the reference block among one or more further reference blocks. The temporal distance of the block can be determined based on the POC of the further reference picture containing the block and the POC of the reference picture.
[0202] A further reference block can be selected by identifying the block with the QP closest to the reference block's quantization parameter in one or more further reference blocks, or by identifying the block with the largest QP in one or more further reference blocks, or by identifying the block with the smallest QP in one or more further reference blocks, or by identifying the block indicated by the L0 (or L1) motion vector in one or more further reference blocks.
[0203] The decoder applies the inherited cross-component model (in step 2140) to the reconstruction of the first color patch to generate a cross-component prediction for the second color patch. The decoder can use the generated cross-component prediction (in step 2150) to reconstruct the current patch (e.g., in combination with the prediction residual) or store (in step 2160) the determined cross-component model to encode and decode subsequent patches. The decoder can then provide the reconstructed current patch for display as part of the reconstructed current image.
[0204] VI. Example Electronic System Many of the features and applications described above are implemented as software processes, which are specified as a set of instructions recorded on a computer-readable storage medium (also known as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, processor cores, or other processing units), they cause the processing unit to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard disks, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. Computer-readable media do not include carrier waves and electronic signals transmitted via wireless or wired connections.
[0205] In this specification, the term "software" means firmware residing in read-only memory or an application stored in magnetic memory that can be read into memory for processing by a processor. Furthermore, in some embodiments, multiple software inventions may be implemented as sub-parts of a larger program while remaining independent software inventions. In some embodiments, multiple software inventions may also be implemented as independent programs. Finally, any combination of independent programs that collectively implement the software inventions described herein is within the scope of this disclosure. In some embodiments, when a software program is installed on one or more electronic systems and runs, it defines one or more specific machine implementations that execute and perform the operations of the software program.
[0206] Figure 22 An electronic system 2200 implementing certain embodiments of this disclosure is conceptually illustrated. The electronic system 2200 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc.), a mobile phone, a PDA, or any other type of electronic device. Such an electronic system includes various types of readable computer media and interfaces for a variety of other types of readable computer media. The electronic system 2200 includes a bus 2205, a processing unit 2210, a graphics processing unit (GPU) 2215, system memory 2220, a network 2225, a read-only memory 2230, a permanent storage device 2235, an input device 2240, and an output device 2245.
[0207] Bus 2205 represents all system, peripheral, and chipset buses that communicate with numerous internal devices of electronic system 2200. For example, bus 2205 communicates with processing unit 2210 and GPU 2215, read-only memory 2230, system memory 2220, and permanent storage device 2235.
[0208] From these various memory units, processing unit 2210 retrieves instructions for execution and data for processing to perform the processes of this disclosure. The processing unit may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by GPU 2215. GPU 2215 may offload various computational or supplemental image processing provided by processing unit 2210.
[0209] Read-only memory (ROM) 2230 stores static data and instructions used by processing unit 2210 and other modules of the electronic system. On the other hand, permanent storage device 2235 is a read-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 2200 is turned off. Some embodiments of this disclosure use mass storage devices (such as magnetic disks or optical disks and their corresponding disk drives) as permanent storage device 2235.
[0210] Other embodiments use removable storage devices (such as floppy disks, flash memory devices, etc., and their corresponding disk drives) as permanent storage devices. Like permanent storage device 2235, system memory 2220 is a read-write memory device. However, unlike storage device 2235, system memory 2220 is a volatile read-write memory, such as random access memory. System memory 2220 stores some instructions and data used by the processor during runtime. In some embodiments, processes according to this disclosure are stored in system memory 2220, permanent storage device 2235, and / or read-only memory 2230. For example, various memory units include instructions for processing multimedia clips according to some embodiments. From these various memory units, processing unit 2210 retrieves instructions for execution and data for processing to perform the processes of some embodiments.
[0211] Bus 2205 is also connected to input and output devices 2240 and 2245. Input device 2240 allows the user to convey information and select commands to the electronic system. Input device 2240 includes an alphanumeric keypad and pointing device (also known as a "cursor control device"), a camera (e.g., a webcam), a microphone, or similar device for receiving voice commands, etc. Output device 2245 displays images generated by the electronic system or otherwise outputs data. Output device 2245 includes printers and display devices such as cathode ray tube (CRT) or liquid crystal display (LCD), and speakers or similar audio output devices. Some embodiments include devices such as touchscreens, which act as both input and output devices.
[0212] Finally, as Figure 22 As shown, bus 2205 also connects electronic system 2200 to network 2225 via a network adapter (not shown). In this way, the computer can be part of a computer network (e.g., a local area network (“LAN”), a wide area network (“WAN”), or an internal network, or a network of networks such as the Internet). Any or all components of electronic system 2200 can be used with this disclosure.
[0213] Some embodiments include electronic components such as microprocessors, storage, and memory that store computer program instructions in a machine-readable or computer-readable medium (also referred to as a computer-readable storage medium, machine-readable medium, or machine-readable storage medium). Examples of such computer-readable media include random access memory (RAM), read-only memory (ROM), read-only optical disc (CD-ROM), recordable optical disc (CD-R), rewritable optical disc (CD-RW), read-only digital multipurpose optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro-SD card, etc.), magnetic and / or solid-state drives, read-only and recordable Blu-ray® optical discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. A computer-readable medium may store a computer program that can be executed by at least one processing unit and includes a set of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as that generated by a compiler, and files containing higher-level code that are executed by a computer, electronic component, or microprocessor using an interpreter.
[0214] While the foregoing discussion primarily refers to microprocessors or multi-core processors that execute software, many of the features and applications described above are executed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, these integrated circuits execute instructions stored on the circuit itself. Furthermore, some embodiments execute software stored in programmable logic devices (PLDs), ROM, or RAM devices.
[0215] The terms “computer,” “server,” “processor,” and “memory” as used in this specification and any claim of this application refer to electronic or other technical devices. These terms do not include people or groups of people. For the purposes of this specification, the term “display” or “show” means “displayed on an electronic device.” The terms “computer-readable medium,” “machine-readable medium,” and “computer-readable medium” as used in this specification and any claim of this application are entirely limited to tangible, physical objects that store information in a form that can be read by a computer. These terms do not include any wireless signals, wired download signals, or any other transient signals.
[0216] While this disclosure has been described with reference to many specific details, those skilled in the art will recognize that this disclosure may be embodied in other specific forms without departing from its spirit. Furthermore, many figures (including...) Figure 18 and Figure 21This disclosure conceptually illustrates the process. Specific operations of these processes may not be performed in the exact order shown and described. Specific operations may not be performed consecutively in a series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process may be implemented using several sub-processes or as part of a larger macro process. Therefore, those skilled in the art will understand that this disclosure is not limited to the foregoing illustrative details, but should be defined by the appended claims.
[0217] Additional notes The topics described herein sometimes demonstrate different components contained within or connected to other components. It should be understood that the architectures depicted are merely examples, and many other architectures can actually achieve the same functionality. Conceptually, any arrangement of components to achieve the same function is actually “associated” in order to achieve the desired functionality. Therefore, any two components combined here to achieve a particular function can be considered “associated” with each other to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components so associating can also be considered “operably connected,” or “operably coupled,” to each other to achieve the desired functionality, and any two components that can be so associating can also be considered “operably coupled” to each other to achieve the desired functionality. Specific examples of operational coupling include, but are not limited to, physically matable and / or physically interactive components and / or wirelessly interactive and / or logically interactive components.
[0218] Furthermore, regarding any substantially plural and / or singular terms used herein, those skilled in the art can translate from plural to singular and / or from singular to plural depending on the context and / or application. For clarity, various singular / plural permutations may be explicitly defined herein.
[0219] Furthermore, those skilled in the art will understand that, in general, the terms used herein, particularly those in the appended claims, such as those in the body of the claim, are typically considered "open" terms. For example, the word "comprising" should be interpreted as "including but not limited to," the word "having" should be interpreted as "having at least," and the word "includes" should be interpreted as "including but not limited to," etc. Those skilled in the art will also further understand that if a particular number of claim statements are intentional, such intention will be explicitly stated in the claim, and without such a statement, no such intention exists. For example, as an aid to understanding, the appended claim may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim statements. However, the use of such phrases should not be interpreted as implying that any particular claim statement introduced by the indefinite article "a" or "an" contains only one implementation of such a statement, even if the same claim statement includes the introductory phrase "one or more" or "at least one" and indefinite articles such as "a" or "an," for example, "a" and / or "an" should be interpreted as "at least one" or "one or more"; the same applies to the use of definite articles used to introduce claim statements. Furthermore, even if a specific number of introduced claim statements are explicitly stated, those skilled in the art will recognize that such statements should be interpreted as including at least the number stated; for example, the statement "two statements" alone, without other modifiers, means at least two statements, or two or more statements. Moreover, in the use of conventions such as "at least one A, B, and C, etc.", such constructions are generally intended according to conventions understood by those skilled in the art; for example, "a system having at least A, B, and C" will include, but is not limited to, systems with only A, only B, only C, A and B together, A and C together, B and C together, and / or systems with A, B, and C together, etc. When using conventions such as "at least one A, B, or C," such a configuration is generally intended according to conventions understood by those skilled in the art. For example, "a system has at least A, B, or C" includes, but is not limited to, systems with only A, only B, only C, A and B together, A and C together, B and C together, and / or systems with A, B, and C together. Those skilled in the art will further understand that virtually any separate word and / or phrase presenting two or more alternative terms in a description, request, or drawing should be understood to account for the possibility of including one term, either term, or both terms. For example, "A or B" would be understood to include the possibility of including "A" or "B" or "A and B."
[0220] As can be seen from the foregoing, various embodiments of this disclosure have been described herein for illustrative purposes, and various modifications may be made without departing from the scope and spirit of this disclosure. Therefore, the various embodiments disclosed herein are not intended to be limiting, and the true scope and spirit are indicated by the following claims.
Claims
1. A video encoding / decoding method, comprising: Receive pixel data of a current block of a current image of one of the videos to be encoded or decoded, wherein the current block contains a first color block and a second color block; Generate a reconstructed block corresponding to the first color patch; Inherit a cross-component model from a reference block, wherein the cross-component model of the reference block is propagated from the cross-component models of two or more further reference blocks for use by the reference block; The inherited cross-component model is applied to the reconstructed block of the first color patch to generate a generated cross-component prediction for the second color patch; and Use the generated cross-component prediction to encode or decode the current block.
2. The video encoding / decoding method as described in claim 1, wherein the two or more further reference blocks are located via more than one motion vector or block vector of the reference block.
3. The video encoding / decoding method as described in claim 1, wherein the reference block is not encoded by cross-component prediction.
4. The video encoding / decoding method of claim 1, wherein propagating the cross-component model includes selecting a selected further reference block from two or more further reference blocks of the reference block to provide the cross-component model of the selected further reference block for use by the reference block.
5. The video encoding / decoding method of claim 4, wherein the selected further reference block is selected according to a set of predefined rules.
6. The video encoding / decoding method of claim 4, wherein selecting the selected further reference block includes identifying a block coded by cross-component prediction.
7. The video encoding / decoding method of claim 4, wherein the selected further reference block is the only further reference block among the two or more further reference blocks that has a cross-component model.
8. The video encoding / decoding method of claim 4, wherein selecting the selected further reference block includes identifying a block referenced and encoded by the current image.
9. The video encoding / decoding method of claim 4, wherein selecting the selected further reference block includes identifying a block coded by intra-frame prediction.
10. The video encoding / decoding method of claim 4, wherein selecting the selected further reference block includes identifying a block coded by inter-frame prediction.
11. The video encoding / decoding method of claim 4, wherein the selection of the further reference block is made by identifying the block with the shortest spatial distance to each of the further reference blocks.
12. The video encoding / decoding method as described in claim 11, wherein the spatial distance is a horizontal distance or a vertical distance.
13. The video encoding / decoding method of claim 4, wherein the selection of a further reference block is made by identifying the block with the shortest time distance between the reference block and each of the further reference blocks, wherein the time distance of one block is determined based on a picture sequence count (POC) of the further reference block and the picture sequence count (POC) of the reference block.
14. The video encoding / decoding method of claim 4, wherein the selection of the further reference block is made by identifying the block whose quantization parameter (QP) is closest to that of each of the further reference blocks.
15. The video encoding / decoding method of claim 4, wherein the selection of the further reference block is made by identifying the block with the largest quantization parameter (QP) among the further reference blocks.
16. The video encoding / decoding method of claim 4, wherein the selection of the further reference block is made by identifying the block with the smallest quantization parameter (QP) among the further reference blocks.
17. The video encoding / decoding method of claim 4, wherein the selection of the further reference block is made by identifying a block indicated by an L0 or L1 motion vector.
18. An electronic device comprising: A video codec circuit is configured to perform the following operations: Receive data of a current block of pixels to be encoded or decoded into one of the current images of a video, wherein the current block contains a first color block and a second color block; Generate a reconstructed block corresponding to the first color patch; Inherit a cross-component model from a reference block, wherein the cross-component model of the reference block is propagated from the cross-component models of two or more further reference blocks for use by the reference block; The inherited cross-component model is applied to the reconstructed block of the first color patch to generate a generated cross-component prediction for the second color patch; and Use the generated cross-component prediction to encode or decode the current block.
19. A video decoding method, comprising: Receive data of a current block of pixels to be decoded into a current image of a video, wherein the current block contains a first color block and a second color block; Generate a reconstructed block corresponding to the first color patch; Inherit a cross-component model from a reference block, wherein the cross-component model of the reference block is propagated from the cross-component models of two or more further reference blocks for use by the reference block; The inherited cross-component model is applied to the reconstructed block of the first color patch to generate a generated cross-component prediction for the second color patch; and Use the generated cross-component prediction to encode or decode the current block.