Cross-component prediction tool in video coding
By employing a cross-component prediction tool in video coding, and utilizing luminance-reconstructed pixels for chrominance prediction, the problem of low efficiency in chrominance and luminance prediction in existing technologies is solved, thereby improving coding efficiency and gain.
Patent Information
- Application Number
- CN202480032030.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-07
- Filing Date
- 2024-04-11
- Publication Date
- 2025-12-09
AI Technical Summary
Existing video coding technologies are inefficient in chroma and luminance prediction, especially in high-texture areas, making it difficult to effectively improve coding gain.
A cross-component prediction tool is employed, including chroma modeling based on inter-frame luminance, cross-component residual prediction, and convolutional cross-component intra-frame prediction. Chroma prediction is performed using luminance-reconstructed pixels, thereby improving coding efficiency through the improved chroma prediction tool.
It improves the compression efficiency of video encoding, especially in high-texture areas, and enhances encoding gain.
Smart Images

Figure CN121100524A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to Indian Provisional Patent Application Serial No. 202311027178 filed on April 12, 2023 and Indian Provisional Patent Application Serial No. 202311038996 filed on June 7, 2023, each of which is incorporated herein by reference in its entirety. Technical Field
[0002] This document generally relates to image and video coding. More specifically, embodiments of the invention relate to the application of cross-component prediction tools in video coding. Background Technology
[0003] In 2020, the MPEG expert group of the International Organization for Standardization (ISO) and the International Telecommunication Union (ITU) jointly released the first version of the Universal Video Coding Standard (VVC), also known as H.266 (Reference [1]). Recently, the MPEG expert group has been working on the development of the next generation of coding standards, which have improved coding performance compared with existing video coding technologies. As part of this investigation, new coding technologies have also been studied.
[0004] As the inventors understand herein, improved techniques for applying cross-component prediction tools in image and video coding are desired, and these techniques are described herein. As used herein, the term "cross-component prediction" means predicting luminance pixel values and / or chrominance pixel values using chrominance and luma chrominance respectively. That is, luminance prediction includes chrominance values, and chrominance prediction includes luminance values.
[0005] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be considered to be in any prior art based on this section. Attached Figure Description
[0006] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings: Figure 1 An example of pixel sample locations used in the cross-component linear model (CCLM) is depicted; Figure 2 An example of the spatial weights of the convolutional filters used in CCLM is depicted; Figure 3 An example of a reference region (including fill) used in CCLM is depicted; Figure 4 An example of a Sobel-based filter used in the gradient linear model (GLM) is depicted; Figure 5 An example of luminance-based chroma modeling for intra-frame coding units (CUs) according to existing techniques is depicted; Figure 6 An example of luminance-based chromaticity modeling for inter-frame CUs according to an embodiment of the present invention is described; Figure 7 An example process for data streaming for inter-frame LMC according to an embodiment of the present invention is described; Figure 8A An example process for data flow modeling across component residuals in an encoder, according to an embodiment of the present invention, is described; Figure 8B An example process for cross-component residual modeling in a decoder according to an embodiment of the present invention is described; and Figure 9 An example of a reference region for implicitly deriving cross-component residual modeling (CCRM) model parameters according to an embodiment of the present invention is depicted. Detailed Implementation
[0007] This document describes example embodiments involving the application of cross-component prediction tools in video coding. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of various embodiments of the invention. However, it will be apparent that various embodiments of the invention can be practiced without these specific details. In other instances, well-known structures and devices are not described in detail to avoid unnecessarily obscuring, obscuring, or confusing embodiments of the invention.
[0008] summary The example embodiments described herein relate to applying cross-component prediction tools for intra- or inter-frame prediction in image and video coding. The proposed methods include: 1) chroma modeling based on inter-frame luma (inter-frame luma and intra-frame luma-chroma of inter-coding units (CUs); 2) cross-component residual prediction; and 3) convolutional cross-component intra-frame prediction.
[0009] Cross-component prediction in video coding In enhanced compression model (ECM) software implementations, such as ECM 7 or later (Reference [2]), the cross-component linear model (CCLM) prediction mode, the convolutional cross-component intra-prediction model (CCCM) mode, and the gradient linear model (GLM) mode use modeling of luminance-reconstructed pixels to predict chrominance. A quick overview of these prediction modes follows.
[0010] CCLM: Prediction Mode of Cross-Component Linear Model In VVC (Reference [1]), the CCLM mode predicts chromaticity samples by using a linear model based on reconstructed luminance samples from the same coding unit (CU), as follows: (1) in, This represents the predicted chromaticity samples in the CU. This represents the downsampled reconstructed brightness samples from the same CU, and and Indicates model parameters.
[0011] It has three modes: Linear Model (LM) mode (using both top and left neighboring samples), LM-A mode (using only top neighboring samples), and LM-L mode (using only left neighboring samples).
[0012] and The parameters are estimated as follows. Figure 1 This illustrates an example of the positions of the left and top samples relative to the samples in the current block in CCLM mode. A group of four neighboring brightness samples at the selected location is downsampled and compared four times to find the two larger values: x 0 A and x 1 A and two smaller values: x 0 B and x 1 B Their corresponding chromaticity sample values (Rec C ) represents y 0 A , y 1 A , y 0 B and y 1 B .Then X a , X b , Y a and Y b It is deduced as: ; (2) ; , in, x >> n express x Binary right shift n bit or x Divide by 2 n .
[0013] Finally, the linear model parameters are obtained according to the following equation. and : (3) .
[0014] CCCM: Convolutional Cross-Component Intra-Frame Prediction Model CCCM is similar to CCLM, but uses a convolutional filter. The 7-tap filter consists of a 5-tap plus shape space component, a nonlinear term (P), and a bias term (B), as shown below: (4) Among them, c i These are the filter coefficients, where, i = 0, 1, ..., 6, B is the bias term, and P is the nonlinear term, calculated as follows: P = (C*C + midVal )>>bitDepth, Where bitDepth represents the bit depth of the luminance component, and midVal = 2 (bitDepth-1) .
[0015] The input of the spatial 5-tap component of the filter consists of the center (C) luminance sample that is in the same position as the chrominance sample to be predicted, and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighboring samples, such as... Figure 2 As shown. The bias term B represents the scalar offset between the input and output and is typically set to an intermediate chroma value (e.g., 512 for 10-bit content). Filter coefficients c i It is calculated by minimizing the mean square error (MSE) between the predicted chromaticity samples and the reconstructed chromaticity samples in the reference region (reference [2]). For example, as Figure 3 As shown, the prediction unit (PU) is surrounded by a reference region (305), which is defined by the area above and to the left of the PU. n (For example, n= 6) Consists of rows or columns of chroma samples. The reference area extends to the right of the PU boundary by one PU width and downwards by one PU height. The reference area is adjusted to include only available samples. The reference area, displayed in black, needs to be extended (by one row or column of pixels) to support... Figure 2 The cross-shaped spatial filter is shown as a "side sample point", and it is filled when the filter is applied to an unavailable region.
[0016] GLM: Gradient Linear Model GLM utilizes the gradient of luminance samples. To derive a linear model: (5) like Figure 4 As shown, four 3×2 gradient filters are enabled for GLM.
[0017] As described below, several novel cross-component prediction tools are proposed in exemplary embodiments of the present invention.
[0018] Chromaticity modeling based on inter-frame luma (inter-frame luma of inter-frame CU and intra-frame LM chroma) In traditional codecs, for inter-frame coding tools, luma and chroma share the same or scaled motion vectors (MVs) based on luma and chroma sampling formats (e.g., 4:2:0, 4:4:4, etc.). In the example embodiments of this disclosure, it is proposed to break these traditional rules. For example, luma uses MV for inter-frame prediction, while chroma uses luma-related cross-component prediction tools. The basic idea is that, for a specific content, for an inter-frame coding unit, the reconstructed luma (510, 520) can be used to better predict chroma (505) with these cross-component intra-frame chroma tools (e.g., CCLM, CCCM, GLM, etc.). Several cross-component prediction tools for intra-frame coding CUs are provided in ECM [2], among which, such as Figure 5 As shown, chromaticity (505) is predicted using luminance reconstruction samples (510) of intra-frame CUs, derived modeling parameters through techniques such as CCLM, CCCM and GLM.
[0019] In some contexts, inter-frame CUs use more chroma residual bits for encoding (e.g., high-texture regions), and cross-component prediction can further improve coding gain through improved chroma prediction. The example embodiments presented in this paper propose applying cross-component prediction to inter-frame CUs to improve compression efficiency.
[0020] like Figure 6As shown, the current luma block (605) is predicted using inter-frame prediction and reconstructed like any inter-frame CU, while the chroma block (610) is predicted using the current luma reconstructed samples (605), where appropriate modeling parameters are derived using techniques such as CCLM, CCCM, GLM, or any other similar modeling techniques. This coding tool will be referred to as inter-frame LMC (luminance-based chroma modeling of inter-frame CUs).
[0021] In another embodiment, instead of using neighboring luma and chroma reconstructed pixels, either the CCLM model or the CCCM model can be derived using luma and chroma prediction samples derived from the inter-frame mode of the current CU, and this model is applied to the luma reconstruction of the current CU to derive the inter-frame LMC predictor.
[0022] In another embodiment, instead of always using neighboring luma and chroma reconstructed pixels, either the CCLM model or the CCCM model can be derived using luma and chroma prediction samples derived from the current CU's inter-frame mode or neighboring luma and chroma reconstructed pixels, and then selected via signaling through a new syntax flag.
[0023] In another embodiment, based on the CU size, the CCLM or CCCM model is modified to use neighboring luma and chromaticity reconstructed pixels or inter-frame patterns of the current CU to derive luma and chromaticity prediction samples. Low-level decomposition (LLD) or Gaussian decomposition can be used to derive the model parameters.
[0024] In another embodiment, the CCLM or CCCM model is modified to use both neighboring luma and chromaticity reconstructed pixels and luma and chromaticity prediction samples derived from the inter-frame pattern of the current CU. The model parameters can be derived using low-level decomposition (LLD) or Gaussian decomposition.
[0025] In another embodiment, chromaticity prediction can be derived using a fusion of the CCLM and CCCM models. Fusion weights ( w ) can be fixed (for example, w = 0.5) or w This can be derived through a template matching process. For example, if the fusion is represented as: (6) Then used for derivation w The minimization model can be expressed as: , in, T Ref This represents the pixel used for chromaticity prediction with the reference template region, i.e., the (width) × 4 and 4 × (height) chromaticity reconstruction bands that are adjacent to the current CU; T Inter This represents the chroma prediction pixels derived using inter-frame prediction from the reference template region; and T LMC This represents the chromaticity prediction pixel derived using LMC prediction derivation from the reference template region.
[0026] In another embodiment, the encoder can select from a variety of chromaticity prediction modes (as previously described) and transmit the selected selection via a signal through appropriate syntax parameters. Examples of such syntax are depicted in Table 1.
[0027] Table 1. Example Syntax for Inter-Frame LMC
[0028]
[0029] An inter_lmc_flag value of 1 indicates that inter-frame LMC is used. An inter_lmc_flag value of 0 indicates that inter-frame LMC is not used.
[0030] inter_lmc_mode_idx specifies the index of the selected inter-frame LMC mode (e.g., choosing between LM_CHROMA_IDX and MMLM_CHROMA_IDX).
[0031] A value of 1 for cccm_flag indicates the use of CCCM. A value of 0 for cccm_flag indicates the use of CCLM.
[0032] Figure 7 An example data stream is depicted for deriving chroma predictions for inter-frame LMC modes. (e.g.) Figure 7 As shown, the decoder reconstructs pixels based on the inter-frame coding syntax parameters (705). If inter_lmc_flag == 1, the decoder uses additional syntax elements (such as inter_lmc_mode_idx and cccm_flag) to derive additional parameters for modeling the chroma reconstruction (710, 715). The derived model parameters are then applied to derive the reconstructed chroma samples (720).
[0033] In the experiments, as shown in Table 2, it was found that disabling sub-block partitioning for inter-frame LMC can improve coding efficiency. Therefore, in some embodiments, it is preferable not to use sub-block partitioning for inter-frame LMC. For example, in Class F, the compression gain of BD-Y is -0.18% without sub-block partitioning, while it is -0.01% with sub-block partitioning. In Table 2, Class D and Class F refer to the classes of video test sequences used in JVET.
[0034] Table 2. JVET SDR-CTC LDB results for inter-frame LMC with and without sub-blocks
[0035] Cross-component residual modeling (CCRM) and cross-component residual prediction (CCRP) This is an intra-frame prediction tool designed to use luminance residual modeling to achieve better chrominance prediction.
[0036] a. The reconstruction will have both low-frequency and high-frequency signals, among which traditional intra-frame prediction can help to better predict low-frequency signals.
[0037] b. The high-frequency component can be predicted using a cross-component residual model.
[0038] This method can be described as follows: Assuming PredL and ReconL are the predicted luminance data and reconstructed luminance data, respectively, and PredC is the chrominance prediction data corresponding to the intra-frame prediction mode for luminance, then: (7) in, alpha The values (modeling parameters) are derived in the decoder based on modeling using the neighboring residual properties of luminance and chrominance. Alternatively, the modeling parameters can be derived in the encoder and sent to the decoder via appropriate syntax parameters. In equation (7), using luminance residuals instead of luminance pixel values can improve prediction because the high-frequency components of the chrominance signal may be better predicted as modeling improves with only high-frequency residual signals.
[0039] In an embodiment, calculations can be performed based on modeling using neighboring luminance and chromaticity residual values. alpha Parameters. Let... Represents the chromaticity residual value and This represents the downsampled brightness residual value. Then, it can be derived using causal residual samples around the current block based on linear least squares error estimation. alpha ( ),as follows: (8) in, and These are the chromaticity residual samples and downsampled luminance residual samples surrounding the current block, and N This represents the total number of neighboring samples.
[0040] In another embodiment, alphaThe alpha values can be explicitly sent from the encoder to the decoder via signals. For example, a set of alpha values can be allowed, such as [0, ±1 / 8, ±¼, ±½, ±1]. The alpha values selected from that pre-specified table, along with a sign flag for each chroma component, can then be sent via signals using syntax parameters (e.g., alpha_idx).
[0041] Figure 8A An example of encoder-side data flow used to determine the CCRM / CCPM modeling parameters and how to derive the chroma residuals using cross-component residual modeling is depicted. The process uses raw (input) uncompressed chroma values (802), predicted chroma values (805), and predicted luminance residual values (830) as they are expected to be received in the decoder (e.g., after inverse quantization and inverse transform, not shown). Given the input chroma values (802) and predicted chroma values (805), subtractor 810 generates residual chroma values (808), which are combined with the predicted luminance residual values (830) to derive the CCRM / CCPM model parameters (e.g., alpha Further subtract (in 815) the output of "luminance residual modeling" (835) from the chrominance residual (808). alpha The luminance residual (837) and the final chrominance residual output (817) (together with the corresponding luminance residual, not shown) are passed to transform and quantize (820) and entropy encoding (840) to generate the output bitstream (827).
[0042] In an embodiment, in the unit "Model Parameter Derivation" (825), the modeling parameters of CCRM / CCRP can be derived using the neighboring residual properties of luminance and chrominance. These parameters can be regenerated in the decoder, or alternatively, they can be sent to the decoder via appropriate syntax parameters by considering the trade-off between bits and distortion.
[0043] Figure 8B The specification data stream for CCRM / CCRP-based decoding is shown. Given an encoded bitstream (827), after inverse quantization (IQ) and inverse transform (IT) (850), the luminance residual (854) and chrominance residual (852) are extracted. In cell 825B, CCRM modeling parameters (e.g., alpha The luminance modeling residuals (867) are either explicitly derived from the decoded information or extracted from the bitstream via appropriate syntax parameters. In unit 865, the luminance residuals (854) and modeling parameters are applied to derive the luminance modeling residuals (867) (e.g., alpha*(luminance residuals), and these luminance modeling residuals are added (870) to the chrominance values of the chrominance residuals (852) and the intra prediction data (862C) to obtain an improved chrominance prediction output value (877). Similarly, the luminance value of the intra prediction data (862L) is added (855) to the luminance residuals (854) to derive a luminance output value (857).
[0044] In an embodiment, the model parameters / coefficients (e.g., alpha ) can be derived by solving a system of linear equations corresponding to neighboring luminance and chrominance residuals using LDL decomposition or Gaussian elimination techniques. In another embodiment, in addition to the central sample, neighboring samples such as the left, right, top, and bottom samples can be used as a five-parameter model, with or without non-linear terms and bias terms.
[0045] Figure 9 An example of a neighboring residual sample region depicting M top rows and N left columns containing chrominance (e.g., M = 2, N = 2) is shown. The corresponding luminance reference residual region is twice the chrominance dimension of the location that is downsampled / subsampled and used for modeling. To improve the accuracy of the model parameters, continuity checks are performed along the top and left boundaries using the current luminance residual samples and neighboring luminance residual samples. For example, only when |(2 * R(X, -1) – R(X, -2) – R(X, 0))| < Th, are the neighboring residuals from R(X, -1) to R(X, -M) considered for modeling. Similarly, only when |(2 * R(-1, Y) – R(-2, Y) – R(0, Y))| < Th, are the residuals from R(-1, Y) to R(-N, Y) considered for modeling, where Th represents a threshold (e.g., Th = 10 << (Bitdepth – 8)) used to ensure the continuity of the luminance residuals at the CU boundary.
[0046] In another embodiment, the model parameters (e.g., alpha ) may be explicitly signaled for some CUs, but for other CUs, these parameters can be inferred from neighboring CUs with similar prediction modes and residual attributes to the current CU. This new mode can be represented as the CCRM_merge mode. For example, if the intra prediction mode of the current CU is a "non-LM" mode (i.e., chrominance prediction does not use luminance modeling modes such as CCLM, CCCM, or GLM), only neighboring CUs with non-LM intra prediction modes are considered. Among these neighboring CUs, as an example, the best neighboring CU can be selected by determining the minimum absolute difference between the average of the neighboring CU luminance residual samples and the average of the current luminance residual samples. If the neighboring CU uses CCRM, then the alphaModel parameters can be derived from their residuals. The pseudocode below provides an example of the proposed CCRM_merge pattern.
[0047]
[0048] In one embodiment, predefined model parameters can be used to reduce the signaling overhead of sending symbol flags and alpha_index values via signals. These predefined model parameters can also be sent all at once at the header level (e.g., in the image header or slice header).
[0049] In another embodiment, a CCRM flag can be sent via signaling for certain chroma intra-prediction modes. For example, the CCRM prediction mode (and therefore the CCRM flag) is only permitted for normal intra-prediction modes such as DC, Planar, and Angular, and not for cross-component prediction modes such as CCLM, CCCM, and GLM.
[0050] In another embodiment, model parameters can be implicitly derived using neighboring luminance and chromaticity prediction data.
[0051] CCCM Improvement The basic idea is to introduce more nonlinear elements (e.g., nonlinear kernels) as input to the 7-tap convolutional filter instead of the weighted sum of the original center luminance sample (C) and four neighboring luminance samples (N, S, W, E). The proposed nonlinear CCCM (NL-CCCM) considers not only the intensity of the reconstructed luminance sample but also the similarity between the neighboring samples and the center sample. The derivation process of the filter coefficients remains unchanged (reference [2]).
[0052] In one embodiment, a nonlinear kernel can be used instead of the weighted sum of four neighboring samples. The predicted chromaticity value can be derived as follows: (9) in, K ( ) is a nonlinear kernel function that takes into account the similarity of intensity and / or center brightness. For X In {N,S,E,W}, the kernel function K (X) can be defined as: 1) Distance: a.|XC|, (for example, K (X) = |XC|). in, th This represents the threshold.
[0053] For example, th= 0.5. maxVal is defined as 2. BitDepth -1, minVal is 0, and midVal = 2 (BitDepth-1) .
[0054] Note: x ? y : z Indicates if x If the value is TRUE or not equal to 0, the output value is y Otherwise, the output value is z .
[0055] 2) Square: .
[0056] In another embodiment, a higher-order polynomial (e.g., third-order) of the center brightness sample point can be used to replace the original nonlinear term. Then, the predicted chromaticity value can be defined as: (10) Among them, nonlinear terms .
[0057] In another embodiment, the nonlinear term can be retained. Simultaneously, a higher-order polynomial (e.g., 3rd order) of the center luminance sample is used to replace the sample farthest from the center luminance sample among the four neighboring samples. Then, the predicted chromaticity value can be defined as: (11) Among them, nonlinear terms ,and {N, S, W, E} excludes the nearest neighbor sample whose luminance value is furthest from the luminance value of the center (C) pixel. That is, if the luminance difference between all four neighbor samples is calculated as... , i = 1 to 4, Therefore, in equation (11), using Summation and Exclusion The largest neighboring pixel.
[0058] In another embodiment, in addition to replacing the original 7-tap filter, the modified NL-CCCM filter described above can be processed as an additional intra-prediction mode. Encoder side rate distortion (RD) decisions will be made based on RD costs. When CCCM is enabled, a CU-level flag is signaled indicating whether the additional NL-CCCM mode is used. If the NL-CCCM mode is used, the following filter index is signaled, indicating which nonlinear filter is applied to derive the final chroma prediction value. An example using the syntax nl_cccm_flag can be provided.
[0059] A value of 1 for nl_cccm_flag indicates that non-linear CCCM is used. A value of 0 for nl_cccm_flag indicates that non-linear CCCM is not used.
[0060] References Each of the references listed in this article is incorporated herein by reference in its entirety. The term JVET refers to the Joint Video Experts Group of ITU-TSG 16 WP 3 and ISO / IEC JTC 1 / SC 29.
[0061] [1] "Versatile Video Coding", Rec.ITU-T H.266, August 2020.
[0062] [2] JVET-AB2025, “Algorithm description of Enhanced Compression Model 7 (ECM 7)”, M. Coban et al., Mainz, Germany, October 2022.
[0063] Example computer system implementation Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuit systems and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may make, control, or execute instructions related to the application of cross-component prediction tools in image and video coding, such as those described herein. The computer and / or IC may calculate any of the various parameters or values related to the application of cross-component prediction tools in image and video coding as described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0064] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present invention. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement methods related to the application of cross-component prediction tools in image and video coding as described above by executing software instructions in a processor-accessible program memory. Embodiments of the present invention can also be provided in the form of a program product. The program product may comprise any non-transitory tangible medium carrying a set of computer-readable signals including instructions that, when executed by a data processor, cause the data processor to perform the methods of the present invention. The program product according to the present invention can take any of a variety of non-transitory and tangible forms. The program product may comprise, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0065] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, references to these components (including references to “devices”) should be interpreted as including any component that performs the function of the described component as an equivalent of that component (e.g., functionally equivalent), including components that are structurally different from those that perform the functions in the illustrated exemplary embodiments of the invention.
[0066] Various aspects of this disclosure can be understood from the following enumerated example embodiments (EEE): EEE1. A method for video decoding of inter-frame coded coding units, the method comprising: Receive coded units (CUs) encoded using inter-frame prediction mode. The first flag is used to determine whether chroma prediction is enabled via linear model luminance-correlated chroma inter-frame prediction; and If the first flag indicates that chromaticity prediction is enabled via linear model luminance correlation prediction, then: Read the first syntax parameter, which indicates a specific luminance-related chromaticity inter-frame prediction mode in two or more such modes; The second flag indicating whether to use the convolutional cross-component (CCCM) inter-frame prediction model; and The chromaticity pixels of the encoded unit are decoded based on the specific luminance-related chromaticity inter-frame prediction mode and the second flag.
[0067] EEE2. The method as described in EEE1, wherein the first syntax parameter indicates one of the cross-component linear model (CCLM) or gradient linear model (GLM) for chroma prediction of intra-frame coded CU.
[0068] EEE3. The method as described in EEE2, wherein the first syntax parameter indicates the use of the fusion of the CCLM model and the CCCM model.
[0069] EEE4. The method as described in EEE3, wherein the fusion can be represented as: , in, T Ref The reference template region represents the chromaticity reconstruction band that borders the encoded unit; T Inter This represents the chroma prediction samples derived using the inter-frame prediction of the reference template region. T LMC This represents the chromaticity prediction samples derived using the LMC prediction derivation of the reference template region, and w This represents the weight within the range [0, 1].
[0070] EEE5. A method for encoding a video bitstream using cross-component residual modeling (CCRM), the method comprising: Receive input video images, each frame including luminance pixels and chrominance pixels; For the region in the input video image to be encoded: Access the original chroma pixel values of the region; Access the predicted chroma pixel values and predicted residual luminance pixel values of the region; The chromaticity residual value is generated by subtracting the predicted chromaticity pixel value from the corresponding original chromaticity pixel value; Based on the predicted residual luminance pixel values and the chrominance residual values, the CCRM model parameters are derived. The CCRM model parameters are applied to the predicted residual brightness pixel values to generate adjusted predicted residual brightness pixel values. Subtracting the adjusted predicted residual luminance pixel value from the chromaticity residual value generates the adjusted chromaticity residual value; and The encoded bitstream is generated based at least on the adjusted chroma residual values.
[0071] EEE6. The method as described in EEE5, wherein the CCRM model parameters include an alpha value, wherein calculating the alpha value includes calculating the following formula: , in, Indicates the predicted chromaticity residual value and This represents the predicted brightness residual value sampled around the region, and N This represents the total number of neighboring samples.
[0072] EEE7. The method as described in EEE5 or EEE6, wherein the CCRM model parameters are sent to the decoder along with the encoded bitstream via a signal.
[0073] EEE8. A method for decoding a coded video bitstream using cross-component residual modeling (CCRM), the method comprising: Receive an encoded bitstream including encoded images, each encoded image comprising luminance pixels and chrominance pixels; For regions in encoded video images: Based on the encoded bitstream, generate the predicted residual luminance pixel value and the predicted residual chrominance pixel value of the region; Access CCRM model parameters based on the predicted residual luminance pixel value and the predicted residual chrominance pixel value; The CCRM model parameters are applied to the predicted residual brightness pixel values to generate adjusted predicted residual brightness pixel values. The adjusted predicted residual luminance pixel value is added to the predicted residual chrominance pixel value to generate an adjusted chrominance residual value; and The adjusted chroma residual value is added to the inter-frame predicted chroma value of the region to generate the output chroma pixel value of the region.
[0074] EEE9. The method as described in EEE8, wherein the CCRM model parameters include an alpha value, wherein calculating the alpha value includes calculating the following formula: , in, Indicates the predicted chromaticity residual value and This represents the predicted brightness residual value sampled around the region, and N This represents the total number of neighboring samples.
[0075] EEE10. The method as described in EEE8 or EEE9, wherein the CCRM model parameters are received together with the encoded bitstream.
[0076] EEE11. The method as described in any one of EEE8 to EEE10, wherein the CCRM model parameters further include a CCRM merging flag for indicating whether CCRM merging is enabled, wherein, If CCRM is enabled: When CCRM merging is enabled, CCRM model parameters are inferred from neighboring CUs; Otherwise, if CCRM merging is not enabled, the CCRM model parameters are extracted from the bitstream.
[0077] EEE12. The method of any one of EEE8 to EEE11, wherein the CCRM model parameters include an index to an array of possible alpha values and a sign indicating whether the alpha value selected via the index is positive or negative.
[0078] EEE13. A method for chromaticity prediction using a convolutional cross-component (CCCM) prediction model, said model comprising predicting chromaticity values using a nonlinear filter according to the following formula:
[0079] Among them, c i This represents the filter parameters, where i = 0 to 6. K ( . ) is a non-linear kernel function, and for the center brightness sample (C), N, S, W, and E represent its north, south, west, and east neighboring pixel samples, respectively. And B is between 0 and 2 bitDepth A fixed bias term between -1 and 1, where, for neighboring pixel X, K (X) includes the following distance functions:
[0080] in, th This represents the threshold, and maxVal = 2. bitDepth -1.
[0081] EEE14. The method of claim 13, wherein, K (X) includes one of the following:
[0082]
[0083] EEE15. A method for chromaticity prediction using a convolutional cross-component (CCCM) prediction model, said model comprising predicting chromaticity values using a nonlinear filter according to the following formula:
[0084] Among them, c i Let represent the filter parameters, where i = 0 to 6, and for the center brightness sample (C), N, S, W, and E represent its north, south, west, and east neighboring pixel samples, respectively.
[0085] And B is between 0 and 2 bitDepth Fixed bias terms between -1 and 1.
[0086] EEE16. A method for chromaticity prediction using a convolutional cross-component (CCCM) prediction model, said model comprising predicting chromaticity values using a nonlinear filter according to the following formula:
[0087] Among them, c i Let represent the filter parameters, where i = 0 to 6, and B is a variable between 0 and 2. bitDepth Fixed bias terms between -1 For the center brightness sample point (C) It is one of the neighboring samples in the east (E), north (N), west (W), and south (S) directions. Among them, the neighboring sample with the largest difference in brightness value from the central brightness sample (C) is excluded.
[0088] EEE17. A tangible computer-readable storage medium having stored thereon computer-executable instructions for executing, with one or more processors, any one of the methods described in EEE1 to EEE16.
[0089] EEE18. An apparatus comprising a processor and configured to perform any of the methods described in EEE1 through EEE16.
[0090] Equivalents, extensions, alternatives and miscellaneous Therefore, example embodiments involving the application of cross-component predictive coding tools in image and video coding are described. In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the invention and the applicant's inventive intent is the set of claims published in specific form according to this application, wherein such claim publication includes any subsequent corrections. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and drawings should be viewed in an illustrative rather than restrictive sense.
Claims
1. A method for video decoding of inter-frame coded coding units, the method comprising: Receive coded units (CUs) encoded using inter-frame prediction mode. The first flag is used to determine whether chroma prediction is enabled via linear model luminance-correlated chroma inter-frame prediction. as well as If the first flag indicates that chromaticity prediction is enabled via linear model luminance correlation prediction, then: Read the first syntax parameter, which indicates a specific luminance-related chrominance inter-frame prediction mode among two or more such modes; The second flag indicates whether the convolutional cross-component (CCCM) inter-frame prediction model is used; as well as The chromaticity pixels of the encoded unit are decoded based on the specific luminance-related chromaticity inter-frame prediction mode and the second flag.
2. The method as described in claim 1, wherein, The first syntax parameter indicates either the cross-component linear model (CCLM) or the gradient linear model (GLM) for chroma prediction of intra-frame coded CU.
3. The method as described in claim 2, wherein, The first syntax parameter indicates the use of the fusion of the CCLM model and the CCCM model.
4. The method of claim 3, wherein, The fusion can be represented as: , in, T Ref The reference template region represents the chromaticity reconstruction band that borders the encoded unit; T Inter This represents the chroma prediction samples derived using the inter-frame prediction of the reference template region. T LMC This represents the chromaticity prediction samples derived using the LMC prediction derivation of the reference template region, and w This represents the weight within the range [0, 1].
5. A method for encoding a video bitstream using cross-component residual modeling (CCRM), the method comprising: Receive input video images, each frame including luminance pixels and chrominance pixels; For the regions in the input video images to be encoded: Access the original chroma pixel values of the region; Access the predicted chroma pixel values and predicted residual luminance pixel values of the region; Chromaticity residual values are generated by subtracting the corresponding predicted chromaticity pixel values from the original chromaticity pixel values. Based on the predicted residual luminance pixel values and the chrominance residual values, the CCRM model parameters are derived. The CCRM model parameters are applied to the predicted residual brightness pixel values to generate adjusted predicted residual brightness pixel values. The adjusted predicted residual luminance pixel value is subtracted from the chromaticity residual value to generate the adjusted chromaticity residual value; as well as The encoded bitstream is generated based at least on the adjusted chroma residual values.
6. The method of claim 5, wherein, The CCRM model parameters include an alpha value, wherein calculating the alpha value includes calculating the following formula: , in, Indicates the predicted chromaticity residual value and This represents the predicted brightness residual value sampled around the region, and N This represents the total number of neighboring samples.
7. The method of claim 5 or 6, wherein, The CCRM model parameters, along with the encoded bitstream, are sent to the decoder via a signal.
8. A method for decoding a coded video bitstream using cross-component residual modeling (CCRM), the method comprising: Receive an encoded bitstream including encoded images, each encoded image comprising luminance pixels and chrominance pixels; For regions in encoded video images: Based on the encoded bitstream, generate the predicted residual luminance pixel value and the predicted residual chrominance pixel value of the region; Access CCRM model parameters based on the predicted residual luminance pixel value and the predicted residual chrominance pixel value; The CCRM model parameters are applied to the predicted residual brightness pixel values to generate adjusted predicted residual brightness pixel values. The adjusted predicted residual luminance pixel value is added to the predicted residual chrominance pixel value to generate the adjusted chrominance residual value. as well as The adjusted chroma residual value is added to the inter-frame predicted chroma value of the region to generate the output chroma pixel value of the region.
9. The method of claim 8, wherein, The CCRM model parameters include an alpha value, wherein calculating the alpha value includes calculating the following formula: , in, Indicates the predicted chromaticity residual value and This represents the predicted brightness residual value sampled around the region, and N This represents the total number of neighboring samples.
10. The method of claim 8 or 9, wherein, The CCRM model parameters are received together with the encoded bitstream.
11. The method according to any one of claims 8 to 10, wherein, The CCRM model parameters further include a CCRM merging flag to indicate whether CCRM merging is enabled, wherein... If CCRM is enabled: When CCRM merging is enabled, CCRM model parameters are inferred from neighboring CUs; Otherwise, if CCRM merging is not enabled, the CCRM model parameters are extracted from the bitstream.
12. The method according to any one of claims 8 to 11, wherein, The CCRM model parameters include an index to an array of possible alpha values and a sign indicating whether the alpha value selected via the index is positive or negative.
13. A method for chromaticity prediction using a convolutional cross-component (CCCM) prediction model, said model comprising predicting chromaticity values using a nonlinear filter according to the following formula: in, c i This represents the filter parameters, where i = 0 to 6. K ( . ) is a non-linear kernel function, and for the center brightness sample (C), N, S, W, and E represent its north, south, west, and east neighboring pixel samples, respectively. And B is between 0 and A fixed bias term between them, where, for neighboring pixel X, K (X) includes the following distance functions: in, th Represents the threshold, and .
14. The method of claim 13, wherein, K (X) includes one of the following: 。 15. A method for chromaticity prediction using a convolutional cross-component (CCCM) prediction model, said model comprising predicting chromaticity values using a nonlinear filter according to the following formula: , in, c i Let represent the filter parameters, where i = 0 to 6, and for the center brightness sample (C), N, S, W, and E represent its north, south, west, and east neighboring pixel samples, respectively. And B is between 0 and Fixed bias terms between them.
16. A method for chromaticity prediction using a convolutional cross-component (CCCM) prediction model, said model comprising predicting chromaticity values using a nonlinear filter according to the following formula: in, c i Let represent the filter parameters, where i = 0 to 6, and B is a variable between 0 and 6. The fixed bias term between them For the center brightness sample point (C) It is one of its neighboring samples to the east (E), north (N), west (W), and south (S), among which the neighboring sample with the largest difference in brightness value from the central brightness sample (C) is excluded, and 。 17. A tangible computer-readable storage medium having stored thereon computer-executable instructions for executing, with one or more processors, any one of the methods described according to claims 1 to 16.
18. An apparatus comprising a processor and configured to perform any one of the methods described in claims 1 to 16.