Adaptive cross-component prediction for inter-coded blocks
By using cross-component prediction mode to reconstruct the chrominance component from the luminance component, the problem of low inter-frame coding efficiency in existing technologies is solved, and more efficient video coding is achieved.
Patent Information
- Application Number
- CN202480031640.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-22
- Filing Date
- 2024-05-07
- Publication Date
- 2025-12-19
AI Technical Summary
Existing video coding technologies struggle to effectively utilize the correlation between luminance and chrominance components when employing inter-frame prediction, resulting in low coding efficiency.
A cross-component prediction mode is adopted, which encodes the chrominance component by reconstructing the luminance component. The syntax elements in the cross-component prediction mode set are used for encoding and decoding to achieve joint prediction of luminance and chrominance components.
It improves the compression efficiency of video coding and enhances the performance of inter-frame coding, especially in utilizing the correlation between luminance and chrominance components.
Smart Images

Figure CN121176007A_ABST
Abstract
Description
[0001] This application claims priority to European Patent Application No. 23315190.1, filed on 11 May 2023, and European Patent Application No. 23305991.4, filed on 22 June 2023, which are incorporated herein by reference in their entirety. Technical Field
[0002] This embodiment generally relates to video compression. Specifically, it relates to methods and apparatus for encoding or decoding images or videos. More specifically, this embodiment relates to improved inter-frame coding. Background Technology
[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in video content. Intra-frame or inter-frame prediction is usually used to fully leverage intra-frame or inter-frame image correlations, followed by transform, quantization, and entropy coding of the differences between the original and predicted blocks (often referred to as prediction error or prediction residual). In inter-frame prediction, motion vectors for motion compensation are typically predicted using a motion vector predictor. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to entropy coding, quantization, transform, and prediction. Summary of the Invention
[0004] According to one aspect, a method for video coding is provided. The method includes: reconstructing the luma component of a block of video, the video block being encoded in an inter-frame coding mode; determining a cross-component prediction mode for the video block; encoding one or more syntax elements for the video block, the one or more syntax elements being used to derive a cross-component prediction mode from a set of cross-component prediction modes at the decoder side; and encoding the chroma component of the video block based on the reconstructed luma component and the determined cross-component prediction mode.
[0005] According to another aspect, an apparatus for video encoding is provided. The apparatus includes one or more processors operable to: reconstruct the luma component of a block of video, the video block being encoded in an inter-frame coding mode; determine a cross-component prediction mode for the video block; encode one or more syntax elements for the video block, the one or more syntax elements being used to derive a cross-component prediction mode from a set of cross-component prediction modes at the decoder side; and encode the chroma component of the video block based on the reconstructed luma component and the determined cross-component prediction mode.
[0006] According to another aspect, a method for video decoding is provided. The method includes: decoding one or more syntax elements for a block of video encoded in an inter-frame coding mode; using the one or more syntax elements to determine a cross-component prediction mode for the video block from a set of cross-component prediction modes; reconstructing a luma component for the video block; and using the cross-component prediction mode to reconstruct one or more chroma components for the video block from the reconstructed luma component.
[0007] According to another aspect, an apparatus for video decoding is provided. The apparatus includes one or more processors operable to: decode one or more syntax elements for a block of video encoded in an inter-frame coding mode; use the one or more syntax elements to determine a cross-component prediction mode for the video block from a set of cross-component prediction modes; reconstruct a luma component for the video block; and use the cross-component prediction mode to reconstruct one or more chroma components for the video block from the reconstructed luma component.
[0008] This document describes further embodiments that can be used alone or in combination.
[0009] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform a method for video encoding / decoding according to any embodiment described herein. One or more embodiments also provide a non-transitory computer-readable medium and / or a computer-readable storage medium storing instructions thereon for performing video encoding / decoding according to the method described herein.
[0010] One or more embodiments also provide a computer-readable storage medium storing a bit stream generated according to the method described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving a bit stream generated according to the described method. Attached Figure Description
[0011] Figure 1 The figure illustrates a system block diagram that can implement aspects of this embodiment.
[0012] Figure 2 A block diagram illustrating an embodiment of a video encoder that can implement aspects of this embodiment is shown.
[0013] Figure 3 A block diagram illustrating an embodiment of a video decoder that can implement aspects of this embodiment is shown.
[0014] Figure 4 The illustration shows an example of sample locations used to derive the α and β parameters of a cross-component linear model.
[0015] Figure 5 The illustration shows the effect of the slope adjustment parameter "u", with the left side showing the model created using regular CCLM and the right side showing the model updated using slope adjustment.
[0016] Figure 6 The illustration shows an example of the spatial portion of a convolution filter in a convolution cross-component model.
[0017] Figure 7 The illustration shows a reference region and an example of its filling used to derive the filter coefficients of the cross-component model.
[0018] Figure 8 The illustration shows examples of four Sobel-based gradient modes for use in linear gradient models.
[0019] Figure 9 The illustration shows an example of the spatial portion of a gradient linear CCCM convolution filter.
[0020] Figure 10 An example of the spatial portion of the non-subsampled CCCM luminance term used in a convolutional filter is illustrated.
[0021] Figure 11 The illustration shows an example of a downsampling filter applied to a luminance sample.
[0022] Figure 12 An example of the inter-frame CU prediction and reconstruction process using cross-component prediction is illustrated.
[0023] Figure 13 The illustration shows examples of luminance samples L0, ..., L5 associated with chrominance sample C.
[0024] Figure 14 An example flowchart of a method for video block encoding according to an embodiment is illustrated.
[0025] Figure 15 An example flowchart of a method for video block decoding according to an embodiment is illustrated.
[0026] Figure 16 An example flowchart of a method for determining the inter-frame coding mode and cross-component prediction mode for video block coding, according to a first variant, is illustrated.
[0027] Figure 17 An example flowchart of a method for video block encoding according to a first variant is illustrated.
[0028] Figure 18 An example flowchart illustrating a method for decoding syntax associated with video blocks, based on the first variant, is shown.
[0029] Figure 19 An example flowchart of a method for video block reconstruction based on a first variant is illustrated.
[0030] Figure 20 An example flowchart illustrating a method for decoding syntax associated with video blocks according to a second variant is shown.
[0031] Figure 21 An example flowchart of a method for video block reconstruction according to the second variant is illustrated.
[0032] Figure 22 An example flowchart of a method for video block reconstruction based on a third variant is illustrated.
[0033] Figure 23 An example flowchart of a method for video block reconstruction according to the fourth variant is illustrated.
[0034] Figure 24 The figure illustrates a system block diagram that can implement aspects of this embodiment according to another embodiment.
[0035] Figure 25 This illustrates two remote devices communicating over a communication network, as exemplified by this principle.
[0036] Figure 26 The syntax of a signal exemplified according to this principle is shown. Detailed Implementation
[0037] This application describes various aspects, including tools, features, embodiments, models, solutions, etc. Many of these aspects are described in detail and, at least for the sake of showing individual characteristics, are generally described in a manner that may sound restrictive. However, this is for the sake of clarity of description and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Moreover, these aspects can also be combined and interchanged with aspects described in earlier applications.
[0038] The aspects described and considered in this application can be implemented in many different forms. The following... Figure 1 , 2 Sections 3 and 4 provide some embodiments, but other embodiments are taken into consideration, and Figure 1 , 2 The discussion in section 3 does not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least another aspect generally relates to the transmission of a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media storing instructions for encoding or decoding video data according to any described method, and / or computer-readable storage media storing bitstreams generated according to any described method.
[0039] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0040] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of steps and / or actions can be modified or combined. Furthermore, terms such as "first," "second," etc., can be used in various embodiments to modify elements, components, steps, operations, etc., for example, "first decoding" and "second decoding." The use of these terms does not imply the order of the modified operations unless specifically required. Therefore, in this example, the first decoding does not necessarily have to be performed before the second decoding and can occur before, during, or within overlapping time periods of the second decoding.
[0041] This aspect is not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations (whether pre-existing or developed in the future) and any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically impractical, the aspects described in this application may be used alone or in combination.
[0042] Figure 1 The diagram illustrates a block diagram of an example system that can implement various aspects and embodiments. System 100 can be embodied as a device including various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100 can be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.
[0043] System 100 includes at least one processor 110 configured to execute instructions loaded thereon, for example, to implement the various aspects described in this application. Processor 110 may include embedded memory, input / output interfaces, and a variety of other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. Storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices, as non-limiting examples.
[0044] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 130 may be implemented as a separate element of system 100, or may be incorporated into processor 110 in a combination of hardware and software, as is known to those skilled in the art.
[0045] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded into memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more items of various kinds during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or a portion thereof, bitstreams, matrices, variables, and intermediate or final results from processing of equations, formulas, operations, and operational logic.
[0046] In some embodiments, the memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external dynamic volatile memory (such as RAM) is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222 and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Variety Video Coding, a standard developed by JVET).
[0047] Inputs to elements of system 100 can be provided via a variety of input devices indicated in box 105. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Other examples (not listed) Figure 1 (As shown in the image) includes composite video.
[0048] In various embodiments, the input device of block 105 has respective associated input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal); (ii) down-converting the selected signal; (iii) band-limiting the signal again to a narrower band to select, for example, a signal band, which may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements to perform these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs these multiple functions, such as down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In many embodiments, the RF section includes an antenna.
[0049] Additionally, USB and / or HDMI terminals may include their respective interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, if necessary, in a separate input processing IC or in processor 110. Similarly, aspects of USB or HDMI interface processing may be implemented in a separate interface IC or in processor 110. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in combination with memory and storage elements to process the data streams as needed for presentation on the output device.
[0050] Various elements of system 100 can be provided within an integrated housing. Within the integrated housing, various elements can be interconnected and transfer data between each other using suitable connection means 115, such as internal buses known in the art, including I2C buses, wiring, and printed circuit boards.
[0051] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.
[0052] In various embodiments, data is streamed to system 100 using a Wi-Fi network (such as IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers)). In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks (including the Internet) to allow streaming applications and other over-the-top communication. Other embodiments use a set-top box to provide streaming data to system 100, delivering data via an HDMI connection to input box 105. Still other embodiments use an RF connection to input box 105 to provide streaming data to system 100. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0053] System 100 can provide output signals to a variety of output devices, including a display 165, a speaker 175, and other peripheral devices 185. The display 165 in various embodiments includes one or more of the following: for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be used in a television, tablet computer, laptop computer, mobile phone, or other device. The display 165 can also be integrated with other components (e.g., in a smartphone) or separate (e.g., an external monitor for a laptop computer). Other peripheral devices 185 in various example embodiments include one or more of the following: a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functionality based on the output of system 100. For example, a disk player performs the function of playing the output of system 100.
[0054] In various embodiments, control signals communicate between system 100 and display 165, speaker 175, or other peripheral devices 185 using signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention). Output devices can be communicatively coupled to system 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 can be integrated with other components of system 100 into a single unit of an electronic device (e.g., a television). In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0055] The display 165 and speaker 175 may also be separate from one or more other components, for example, if the RF section of input 105 is part of a separate set-top box. In various embodiments where the display 165 and speaker 175 are external components, the output signal may be provided, for example, via a dedicated output connection including an HDMI port, a USB port, or a COMP output.
[0056] The embodiments may be executed by computer software implemented by processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. Memory 120 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based storage devices, fixed memory, and removable memory, as non-limiting examples. Processor 110 may be of any type suitable for the technical environment and may include one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures, as non-limiting examples.
[0057] Figure 2 An example of a block-based hybrid video encoder 200 is illustrated. Variations of this encoder 200 are taken into account, but for clarity, the encoder 200 is described below without describing all anticipated variations.
[0058] In some embodiments, Figure 2 The illustration also shows an encoder that improves upon the HEVC or VVC standard (Multi-Functional Video Coding, standard ITU-T H.266, ISO / IEC 23090-3, 2020) or an encoder that uses a similar technology to HEVC or VVC (such as the encoder ECM developed by JVET (Joint Video Exploration Team)).
[0059] Before encoding, the video sequence may undergo pre-coding (201), such as applying color transforms to the input color images (e.g., transforming from RGB 4:4:4 to YCbCr 4:2:0), performing remapping of the input image components to obtain a more compression-resistant signal distribution (e.g., using histogram equalization of the color components), or resizing the images (e.g., downsampling). Metadata may be associated with pre-processing and appended to the bitstream.
[0060] In encoder 200, the image is encoded by the encoding elements described below. The image to be encoded is divided (202) and processed, for example, into CUs (coding units) or blocks. In this disclosure, different expressions may be used to refer to such units or blocks resulting from image division. Such terms may be encoding unit or CU, encoding block or CB, luminance CB or block. CTU (coding tree unit) may refer to a set of blocks or a set of units. In some embodiments, a CTU itself may be considered as a block or unit.
[0061] Each unit is encoded, for example, using either intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) whether to use intra-frame or inter-frame mode to encode the unit, and indicates the intra-frame / inter-frame decision, for example, via a prediction mode flag. The encoder can also mix (263) intra-frame prediction results and inter-frame prediction results, or mix results from different intra-frame / inter-frame prediction methods. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.
[0062] The motion refinement module (272) refines the motion field of a block using an already available reference image, without referencing the original block. The motion field for a region can be viewed as the set of motion vectors for all pixels within that region. If the motion vectors are based on sub-blocks, the motion field can also be represented as the set of motion vectors for all sub-blocks within that region (all pixels within a sub-block have the same motion vector, and the motion vectors may differ from sub-blocks). If a single motion vector is used for a region, the motion field for that region can also be represented by that single motion vector (the same motion vector for all pixels in the region).
[0063] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is directly encoded without applying either the transform or quantization process.
[0064] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inversely transformed (250) to decode the prediction residuals. The image blocks are reconstructed by combining (255) the decoded prediction residuals and the predicted blocks. A loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).
[0065] Figure 3 A block diagram of a video decoder 300 is shown. In decoder 300, the bitstream is decoded by the decoding elements described below. Video decoder 300 typically performs operations similar to... Figure 2 The encoding process is described as the inverse of the decoding process. Encoder 200 typically also performs video decoding as part of the video data encoding.
[0066] Specifically, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoded information. Image segmentation information indicates how the image is segmented. The decoder can therefore segment (335) the image based on the decoded image segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. By combining (355) the decoded prediction residuals and the predicted blocks, image blocks are reconstructed.
[0067] The predicted blocks can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). The decoder can mix (373) the intra-frame prediction results and the inter-frame prediction results, or mix the results from multiple intra-frame / inter-frame prediction methods. Before motion compensation, the motion field can be refined (372) using an already available reference image. A loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).
[0068] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., from YCbCr4:2:0 to RGB 4:4:4) or inverse remapping (the reverse of the remapping process performed in pre-encoding processing (201), or resizing the reconstructed image (e.g., upsampling). Post-decoding processing can utilize metadata derived in pre-encoding processing and signaled in the bitstream.
[0069] Some of the embodiments described herein relate to inter-frame prediction, and more specifically to adaptive cross-component prediction for inter-frame coded blocks.
[0070] Any of the embodiments described herein can be implemented, for example, in the inter-frame prediction module of a video encoder or video decoder. For instance, the embodiments described herein can be implemented in the motion compensation module 270 of the video encoder 200 or the motion compensation module 375 of the video decoder 300.
[0071] Cross-component linear model (CCLM) for intra-frame prediction in ECM 7 ( M. Coban, F. Le Leannec, RL. Liao, K.Naser, J.Ström, L.Zhang, "Algorithm description of Enhanced Compression Model 7 (ECM 7) (Algorithm Description of Enhanced Compression Model 7 (ECM7)), document JVET-AB2025, 28th meeting, passed. By conference call, October 2022 This is implemented in CCLM. In CCLM, a linear model is used to predict chromaticity samples based on reconstructed luminance samples from the same CU, as follows: pred_C (i,j)=α·rec_L'(i,j)+β (Equation 1) Where pred_C (i,j) represents the predicted chromaticity sample in the CU, and rec_L (i,j) represents the reconstructed luminance sample after downsampling in the same CU.
[0072] The CCLM parameters (α and β) are derived from the adjacent chromaticity samples (top row and left column) and their corresponding downsampled luminance samples (LM mode).
[0073] In one variant, a maximum of four adjacent chroma samples are used. In another variant, adjacent chroma samples are selected only from the top (LM-A) or only from the left (LM-L), and the selected mode is signaled to the decoder.
[0074] Selected adjacent brightness samples at the selected location are downsampled and compared to find the two smaller values: x 0 A and x 1 A and two larger values: x 0 B and x 1 B Their corresponding chromaticity sample values are represented as y 0 A , y 1 A , y 0 B and y 1 B .Then X a , X b , Ya and Y b The derivation is as follows: X a =( x 0 A + x 1 A +1)>>1 X b =(x 0 B + x 1 B +1)>>1 Y a =( y 0 A + y 1 A +1)>>1 Y b =( y 0 B + y 1 B +1)>>1 Finally, the linear model parameters and The following equation is used to obtain: ;
[0075] Figure 4 (400) shows the left and top samples Rec used for the luminance component. ’ L and the left and top samples Rec used for the chromaticity components c (Samples on the left and top) Figure 4 (Illustrated with circles above) and an example of the location of a sample in the current 2N×2N block involving CCLM mode.
[0076] In another variant, three multi-model LM (MMLM) modes were added. In each MMLM mode, reconstructed neighboring samples are classified into two classes using a threshold that is the average of the brightness of the reconstructed neighboring samples. The linear model for each class is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model.
[0077] In another variant, slope adjustment is applied to both cross-component linear model (CCLM) and multi-model LM predictions. This adjustment tilts the linear function that maps luminance values to chrominance values relative to a center point determined by the average luminance values of the reference samples, such as... Figure 5 As shown in (500), the left side shows the model created using a regular CCLM with parameters a and b, and the right side shows the model updated using the slope adjustment parameter u. t It is the average brightness value of the reference sample.
[0078] In ECM7, a Convolutional Cross-Component Model (CCCM) for intra-frame prediction is also implemented. The CCCM predicts chroma samples from reconstructed luma samples in a manner similar to CCLM. When chroma subsampling is used, the reconstructed luma samples are downsampled to match a lower-resolution chroma grid.
[0079] Additionally, there is an option to use a single-model or multi-model variant of CCCM. A multi-model CCCM mode can be selected for blocks with at least 128 available reference samples. The multi-model variant uses two models: one applied to sample values above the average brightness reference value, and the other for the remaining samples (following the spirit of the CCLM design).
[0080] CCCM uses a convolutional filter consisting of 7 parameters weighted by 7 input samples, where the sample is {s}. i} i=0、..6 Five coefficients are applied to the luminance pixel values corresponding to the plus sign shape, one coefficient is applied to the square term (P), and the last coefficient is applied to the bias term (B). The input to the spatial 5-tap component of the filter consists of the center (C) luminance sample (co-located with the chrominance sample to be predicted) and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as shown below. Figure 6 As shown in (620).
[0081] Term P is represented as the square of the center brightness sample C, scaled to the range of sample values of the content: P = (C) C + midVal )>>bitDepth Where midVal is the rounding term. That is, for 10-digit content, it is calculated as follows: P = (C) C + 512)>>10 The bias term B represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).
[0082] The output of the filter is calculated as the filter coefficients c. i Convolve the input values and crop them to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B (Equation 2) Filter coefficients c i (710) is calculated by minimizing the MSE between the predicted chromaticity sample and the reconstructed chromaticity sample in the reference region. Figure 7 (710) illustrates an example of a reference region consisting of 6 rows / columns of chroma samples above and to the left of the block / PU. The reference region extends one PU width to the right and one PU height downward from the PU boundary. The region is adjusted to include only available samples. The extended portion of the reference region (shown in blue) requires “side samples” to support the plus shapespace filter and fills in unavailable areas.
[0083] MSE minimization is performed by computing the autocorrelation matrix for the luminance input and the cross-correlation vector between the luminance input and the chrominance output.
[0084] make ,in
[0085] In order to derive the coefficient {c i The autocorrelation matrix is inverted. For example, the autocorrelation matrix is inverted by LDL. T The filter coefficients are decomposed and finally calculated using back-substitution. This process roughly follows the calculation of ALF filter coefficients in ECM, however, LDL is chosen. T Decomposition (also known as alternative Cholesky decomposition) avoids the use of square root operations. In another variant, Gaussian elimination is used to invert the matrix, as in the VVC standard. In ECM, computation uses 64-bit integer arithmetic.
[0086] ECM7 also provides a gradient linear model. For the YUV 4:2:0 color format, the gradient linear model (GLM) method can be used to predict chromaticity samples from the gradient of luminance samples. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.
[0087] Compared to CCLM, the two-parameter GLM uses luminance sample gradients instead of downsampled luminance values to derive a linear model. Specifically, when applying the two-parameter GLM, the input to the CCLM process (i.e., downsampled luminance samples) is... The ) was replaced with the brightness sample gradient Other parts of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged.
[0088]
[0089] In a three-parameter GLM, chromaticity samples can be predicted using different parameters based on the luminance sample gradient and downsampled luminance values. The model parameters of a three-parameter GLM are derived from adjacent 6 rows and 6 columns of samples using an MSE minimization method based on LDL decomposition, as used in CCCM.
[0090]
[0091] For signal notification, when CCLM mode is enabled for the current CU, a flag is signaled to indicate whether GLM is enabled for the Cb and Cr components; if GLM is enabled, another flag is signaled to indicate which of the two GLM modes to select, and a syntax element is further signaled to select one of the four gradient filters for gradient calculation.
[0092] Four gradient filters were enabled for GLM, such as Figure 8 As shown.
[0093] ECM7 also provides a gradient- and location-based convolutional cross-component model (GL-CCCM). This is a variant of CCCM where convolution is applied to gradient and location information, rather than the four spatial neighbor samples in the CCCM filter (as in VVC). This variant is shown in... Figure 9 (930) Above. The GL-CCCM filter used for prediction is: predChromaVal = c0C + c1Gy + c2Gx + c3Y + c4X + c5P + c6B (Equation 3) Where Gy and Gx are the vertical and horizontal gradients of brightness, respectively, and are calculated as (see...). Figure 9 (930) Gy = (2N + NW + NE) - (2S + SW + SE) Gx = (2W + NW + SW) - (2E + NE + SE) In addition, parameters Y and X are the vertical and horizontal coordinates of the center brightness sample position.
[0094] The remaining parameters are the same as those in the CCCM tool. The reference area used for parameter calculation is the same as that in the CCCM method.
[0095] In another variant, the reconstructed luminance samples are not downsampled to match a lower-resolution chroma grid, and the co-located reconstructed six luminance samples are used directly, such as... Figure 10As shown. In this variant, four terms constructed from L0, L1, L2, and L4 are also used, plus a bias B (VVC). The number of parameters is then 10.
[0096] In another variant, multiple (e.g., four) downsampling filters can be used to derive the input samples to be weighted by coefficients. The downsampling filter model is signaled for each CU, and the prediction derivation for the chroma samples is as follows: Model 1: predChroma = c0 H(C) + c1 G1(C) + c2 G2(C) + c3 G3(C) + c4 G4(C) + c5 P + c6 B Model 2: predChroma = c0 H(C) + c1 H(W) + c2 H(E) + c3 G1(C) + c4 G1(W) + c5 G1(E) + c6 B Model 3: predChroma = c0 H(C) + c1 H(N) + c2 H(S) + c3 G2(C) + c4 G2(N) + c5 G2(S) + c6 B Model 4: predChroma = c0 H(C) + c1 H(NE) + c2 H(SW) + c3 G4(C) + c4 G4(NE) + c5 G4(SW) + c6 B Where H(·), G1(·), G2(·), G3(·), and G4(·) are downsampling filters applied to the luminance samples, such as... Figure 11 As shown in (1150).
[0097] exist K.Zhang, L.Zhang, Z.Deng, Chia-Ming Tsai, Hsin-Yi Tseng, Cheng-Yen Chuang, Chih-Wei Hsu, Ching-Yeh Chen, Tzu-Der Chuang, Olena Chubach, Yi-Wen Chen, Yu-Wen Huang, Shaw-Min Lei, "EE2-1.6: Non-local cross-component prediction and "Cross-component merge mode (EE2-1.6: Non-local cross-component prediction and cross-component merging mode)" document JVET-AD0188, 30th Meeting, via Teleconference, April 21-28, 2023 In this context, nonlocal cross-component prediction and cross-component merging modes are provided.
[0098] In ECM, for intra-frame blocks, CU-level selection of the luma-to-chroma cross-component prediction mode and several syntax elements (LMCMode flag, cclmFlag, mmlmFlag, cccmFlag, cccmNoSubFlag, glCccmFlag, glmFlag) are enabled. Therefore, the JVET-AD0188 proposal allows the derivation of cross-component prediction (CCP) for a given block from blocks already encoded in the same image.
[0099] Construct a CCP merging candidate list. This list may include spatially adjacent candidates, spatially non-adjacent candidates, or history-based candidates. After including these candidates, a default model is further included to fill any remaining blank spaces in the merge list. To remove redundant CCP models from the list, pruning is applied. After constructing the list, the CCP models are reordered according to their SAD costs, which are obtained using the neighboring templates of the current block. More details are described below: Spatial adjacent and non-adjacent candidates: The positions and inclusion order of spatial adjacent and non-adjacent candidates are the same as those defined for regular inter-frame merge prediction candidates in ECM.
[0100] History-based candidates: A history-based table is maintained to include recently used CCP models, and this table is reset at the start of each CTU row. If the current list is not full after including spatially adjacent and non-adjacent candidates, CCP models from the history-based table are added to the list.
[0101] Default Candidate: CCLM candidates with the default scaling parameters are considered only when the list is not full after including spatially adjacent, spatially non-adjacent, or history-based candidates. If there are no candidates using the single-model CCLM mode in the current list, the default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8}. Otherwise, the default scaling parameters are {0, the scaling parameter of the first CCLM candidate + {1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8}}. The offset parameters are derived based on the default scaling parameters, the average reconstructed brightness sample values from adjacent candidates, and the average reconstructed Cb / Cr sample values from adjacent candidates.
[0102] A flag is signaled to indicate whether CCP merging mode is applied. If CCP merging mode is applied, an index is signaled to indicate which candidate model is used for the current block. Additionally, CCP merging mode is not allowed for the current chroma code block when the current CU is encoded using intra-frame sub-partitioning (ISP) with a single tree, or when the current chroma code block size is less than or equal to 16.
[0103] exist P. Astola, J. Lainema (Nokia), "AHG12: Cross-component residual model (CCRM) for inter prediction (AHG12: Cross-Component Residual Model for Inter-Frame Prediction (CCRM)), document JVET-AD0108, April 2023 The paper presents a cross-component residual model (CCRM) for inter-frame prediction. An 8-tap filter consisting of 6 spatial luma samples, a nonlinear term, and a bias term is used for chroma prediction of inter-frame coded blocks using only the luma channel. Spatial luma samples (L0, ..., L5) are obtained from the luma grid, where the 6 luma samples closest to the chroma position C are selected without downsampling, as shown below. Figure 13 As shown. Figure 13 The display shows the luminance samples L0, ... L5 associated with the chrominance sample C in the half-pixel luminance grid.
[0104] The predicted chromaticity values are as follows: predChromaVal = c0L0 + c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6nonlinear((L0+L3+1)>>1) + c7B, where nonlinear is the nonlinear operator of CCCM and B is the bias.
[0105] The filter coefficients are derived using the division-free Gaussian elimination method of ECM, and the necessary offsets are applied to the samples before the filter derivation.
[0106] When a block has fewer than 64 chroma samples, intra-frame reference samples are used as additional input samples in filter derivation. The design utilizes up to 6 rows and 6 columns of intra-frame reference samples from the CCCM.
[0107] A block with 256 or more chromaticity samples is divided into sub-blocks with a maximum of 256 chromaticity samples. Sub-blocks containing zero luminance residuals are skipped.
[0108] The CCRM tool is currently used to replace temporal chroma prediction (i.e., chroma prediction with motion compensation) commonly used in chroma component inter-frame prediction. One limitation of the CCRM method in JVET-AD0108 is that a single CCP method is used to predict chroma from luma in inter-frame blocks, while several possible CCP methods are allowed for intra-frame blocks.
[0109] Therefore, the compression performance of inter-frame blocks may be limited due to the lack of flexibility in CCPs within inter-frame blocks.
[0110] One objective of the embodiments provided herein is to introduce greater flexibility in the selection of CCP methods for inter-frame blocks in order to improve the compression efficiency of existing video codecs.
[0111] In some embodiments, the CCP of inter-frame blocks is improved by introducing the possibility of using one of a set of multiple CCP methods (hereinafter also referred to as CCP mode) for inter-frame blocks. In this document, the term "block" may also be referred to as a prediction unit (PU) or coding unit (CU), as commonly used in video coding standards or implementations.
[0112] Figure 14 An example of a method (1400) for video block encoding according to an embodiment is illustrated. In this embodiment, the video block is encoded using an inter-frame mode or inter-frame prediction. Therefore, it is referred to as an inter-frame block below. At 1401, the luminance portion of the considered inter-frame block is predicted and reconstructed according to an inter-frame prediction mode determined for the inter-frame block. For example, a reference block is determined in a reference picture by using motion compensation of the inter-frame block using motion data determined for the inter-frame block. On the encoder side, at 1402, a CCP mode is selected for the inter-frame block from a variety of CCP modes. Then, the selected CCP method for the inter-frame block is signaled (1403) in the bitstream encoding the video. This signaling can be explicit (i.e., in the form of one or more syntax elements for identifying the CCP mode) or implicit (i.e., in the form of syntax elements on the decoder side for deriving the CCP mode of the current inter-frame block). In some embodiments, signaling is only provided for the use of the CCP mode, which is determined in the same manner on both the encoder and decoder sides, for example, using a template for the inter-frame block (a reconstructed sample around the inter-frame block) or a reference block for the inter-frame block. At 1404, the chroma portion of the inter-frame block, along with the inter-frame block data (such as inter-frame prediction modes, luma data, etc.), is encoded based on the selected CCP mode.
[0113] Figure 15An example of a method (1500) for video block decoding according to an embodiment is illustrated. On the decoder side, the syntax elements associated with the video block are parsed (1501), i.e., decoded from the bitstream. In this embodiment, the video block is encoded using an inter-frame mode or inter-frame prediction. Therefore, it is referred to as an inter-frame block below. The parsed syntax elements for the inter-frame block are used to identify the CCP mode for that inter-frame block. In some embodiments, only the use of the CCP mode is signaled, and the selected CCP mode is determined in the same way on both the encoder and decoder sides, for example, using a template or reference block of the inter-frame block. At 1502, the CCP mode for a given inter-frame block is derived from the decoded syntax elements. The luma portion of the considered inter-frame block is reconstructed (1503), and the chroma portion (Cb and Cr blocks) of the inter-frame block is predicted and reconstructed from the reconstructed luma blocks according to the CCP mode for the current inter-frame block (1504).
[0114] Several variations of the embodiments provided above may include the following aspects, which may be used alone or in combination.
[0115] When inter-frame blocks are encoded in AMVP mode, the CCP mode and the selected CCP mode are explicitly signaled along with the motion data. For example, AMVP mode is an inter-frame predictive coding mode known in the VVC standard. More generally, AMVP mode is an inter-frame predictive coding mode in which the motion vector difference (mvd) relative to the motion predictor is encoded. The motion predictor can be signaled in the bitstream, for example, by an index indicating a motion predictor candidate in a motion candidate list.
[0116] In another variant, when inter-frame blocks are encoded in merge mode, the CCP mode of the inter-frame blocks can be derived on the decoder side in the same manner as the motion data derivation process used for inter-frame blocks. For example, merge mode is an inter-frame predictive coding mode known in the VVC standard. More generally, merge mode is an inter-frame predictive coding mode in which the motion vector difference (mvd) relative to the motion predictor is inferred to be zero. The motion predictor can be signaled in the bitstream, for example, by an index indicating a motion predictor candidate in the merge candidate list.
[0117] For example, the merge candidate list designed in existing video codecs is used for block-to-block motion information derivation and CCP mode information derivation. Thus, the merge_idx syntax element of an inter-frame block is used for both motion information derivation and CCP mode derivation. When an inter-frame block inherits the motion of a merge candidate, it also inherits the CCP method of that merge candidate.
[0118] Alternatively, a separate merge index dedicated to CCP mode derivation can be signaled at the PU or CU level.
[0119] In another variant, CCP pattern derivation can also be applied to temporal merging derivation. When inheriting motion vectors from TMVP (Temporal Motion Vector Predictor) or SbTMVP (Sub-Block Temporal Motion Vector Predictor), the associated CCP pattern can also be inherited from the temporal candidate. For example, TMVP and SbTMVP motion vectors can be obtained as in the VVC standard or ECM implementation.
[0120] In another variant, CCP mode derivation can be applied only when inter-frame blocks are in skip mode. In this case, inter-frame blocks can inherit the CCP mode from a merge candidate selected from the merge candidate list, similar to motion information.
[0121] In another variant, CCP mode derivation can be applied only when the merge candidate used to derive motion data corresponds to a spatially adjacent block of the current inter-frame block. In this variant, the CCP method for a given inter-frame block can be derived without using non-adjacent merge candidates.
[0122] In another variant, the CCP mode selected for the inter-frame block can be the mode that minimizes the distortion between the filtered predicted luma block and the predicted chroma block. In other words, the CCP mode does not need to be signaled because the same process is performed on both the encoder and decoder sides to select the CCP mode. This distortion is determined between the chroma portion of the temporally predicted inter-frame block and the chroma portion of the inter-frame block predicted from the temporally predicted luma portion of the current inter-frame block using candidate CCP modes. The CCP mode that causes the least distortion among all candidate CCP modes is selected on both the encoder and decoder sides.
[0123] In another variant, in addition to the CCP prediction mode itself, CCP mode inheritance can further include the inheritance of CCP filter parameters.
[0124] In a further variant, these CCP parameters may be inherited only if the current inter-frame block is in skip mode. In another further variant, these CCP parameters may be inherited only if derivation is performed using spatial adjacency merge candidates.
[0125] In another variant, the set and number of available CCP modes can depend on the inter-frame mode; for example, they can depend on whether the inter-frame mode is AMVP, merge, TMVP, or SbTMVP.
[0126] In another variant, CCP modes can be reordered using a template from the current inter-frame block. For each (or a subset thereof) of the CCP mode candidates, the CCP mode is used to predict the reconstructed sample of the template, and the cost of that CCP mode on this template is determined. The CCP mode cost is used to reorder the CCP mode candidates.
[0127] In some embodiments, the inter-frame block is signaled at the block level to the cross-component prediction mode from a plurality of possible cross-component modes. These plurality of cross-component modes may include all or part of the following modes further described above: CCLM, MMLM, CCCM, GLM, GLCCCM.
[0128] Regarding modifications to the coding unit syntax of existing video codec implementations (e.g., the ECM7 under investigation), this signal notification can take the form of Table 1 below.
[0129] As shown in Table 1, in the case of inter-frame CUs, some syntax elements are added to the existing bitstream syntax to identify the cross-component prediction mode for a given inter-frame CU. The added elements are shown in bold in Table 1.
[0130] In this embodiment, for each CU, the signaling flag `inter_ccp_mode_flag` is used to indicate the use of CCP mode to predict the chroma coded blocks of the considered CU. If this flag is on, further syntax elements can be signaled to allow the video decoder to identify the actual CCP mode used for the current CU. This takes the form of the `inter_ccp_mode()` syntax structure, as shown in Table 1 below.
[0131] Table 1: Proposed syntax modifications to support multiple CCP modes for inter-frame CU
[0132] The function CcpAllowed(x0, y0) is a function that determines whether cross-component prediction is allowed for the current block (x0, y0). For example, it checks whether certain conditions are met to allow cross-component prediction for the current block (x0, y0).
[0133] Table 2 below shows examples of possible syntax arrangements used to indicate the CCP mode and CCP parameters used to encode a given inter-frame block in known cross-component prediction methods (CCLM, MMLM, MDLM, GLM, GLCCCM).
[0134] Table 2: Proposed signaling for CCP mode and CCP parameters for chroma blocks
[0135] In Table 2 above, the semantics of the syntax elements are as follows: The cclm_flag syntax element is similar to the cclm_mode_flag syntax element in the VVC specification, but it is used here for inter-frame blocks. It indicates whether to use CCLM across prediction modes to encode and decode the current block.
[0136] ccp_mode_idx indicates the CCLM mode type used for the current inter-frame block, and the neighboring sample type used to derive the cross-component prediction linear model (between the reconstructed samples around the left, top, and top left of the current block).
[0137] `cccmAllowed(x0, y0)` is a function that determines whether CCCM mode is enabled for the current block. Typically, CCCM uses a selected set of neighboring samples to derive filter parameters. This neighborhood is the same as indicated by the `ccp_mode_idx` syntax element. Therefore, CCCM mode is generally allowed if there are enough neighboring samples available in the indicated neighborhood to derive the CCCM filter parameters. Note that this embodiment is not limited to this type of rule, and several other CCM enabling rules may be applied.
[0138] cccm_flag indicates whether CCCM mode is used for the current inter-frame block.
[0139] cccm_no_sub_sampling_flag indicates whether to use the CCCM mode without any luminance downsampling to predict chromaticity samples.
[0140] The gl_cccm_flag indicates whether GL-CCCM cross-prediction mode is used for the current inter-frame block.
[0141] GlmAllowed(x0, y0) is a function that determines whether GLM mode is allowed for the current inter-frame block. Generally, GLM is allowed if CCCM is not used and the color format of the video being encoded / decoded is 420.
[0142] Glm_flag indicates whether GLM cross-component prediction is used for the current inter-frame block.
[0143] Glm_idx indicates the gradient filter used for gradient computation during GLM cross-component prediction.
[0144] cclm_delta_flag indicates whether CCLM slope adjustment is used if CCCM and GLM are not used in the current inter-frame block.
[0145] Cclm_cb0_flag indicates whether slope adjustment is used in the linear model used for cross-component prediction of the Cb component.
[0146] cclm_cr0_flag indicates whether slope adjustment is used in the linear model used for cross-component prediction of the Cr component.
[0147] cclm_cb0_delta_idx indicates the slope adjustment parameter in the linear model used for cross-component prediction of the Cb component.
[0148] cclm_cr0_delta_idx indicates the slope adjustment parameter in the linear model used for cross-component prediction of the Cr component.
[0149] Then, in the case of multi-model CCLM across prediction modes, the following flags can exist in the bitstream.
[0150] Cclm_cb1_flag indicates whether slope adjustment is used in the second linear model used for cross-component prediction of the Cb component.
[0151] cclm_cr1_flag indicates whether slope adjustment is used in the second linear model used for cross-component prediction of the Cr component.
[0152] Cclm_cb1_delta_idx (if present) indicates the slope parameter in the second linear model used for cross-component prediction of the Cb component.
[0153] cclm_cr1_delta_idx (if present) indicates the slope parameter in the second linear model used for cross-component prediction of the Cr component.
[0154] Figure 16 An example of encoder method 1600 for selecting CCP modes for inter-frame prediction (CU) is given. It includes double nested loops (1601, 1602, 1604, 1605) over all inter-frame prediction modes (1601) supported by the codec (skip mode, merge, AMVP, merge affine, AMVP affine, etc.) and possible supported cross-component prediction modes (1602) (including CCLM, MMLM, CCCM, GLM, GLCCCM, and possibly others). For each pair of inter-frame prediction modes... and cross-component prediction mode The RD cost is evaluated (1603) for the CU under consideration. A set of inter-frame prediction and cross-component prediction modes that provide the minimum rate-distortion cost is selected (1606). In some embodiments, non-inter-frame modes may also be evaluated (1607) for the CU. At 1608, the CU uses the inter-frame prediction modes thus jointly selected. and cross-component prediction mode Perform compression and encoding.
[0155] Figure 17 An example of an inter-frame CU entropy coding method (1700) according to this embodiment is illustrated. The input to this method is the CU to be encoded, whose coding mode and coding parameters have been selected by the encoder's RD optimization stage.
[0156] At 1701, the CU prediction mode is encoded into the output bitstream. Then, at 1702, it is checked whether the CU is inter-coded or intra-coded. The case of intra-frame CUs is not shown here because it is not addressed in the embodiments provided herein. At 1703 (if the CU is inter-coded), the cross-component prediction information used to encode the current inter-frame CU is encoded. For example, the bitstream syntax proposed in Table 1 above can be used. At 1704 and 1705, the CU inter-frame mode information selected for the current CU and the associated motion data are encoded. At 1706, the CU residual block is encoded, and the process is complete.
[0157] and Figure 17 The encoding process is the inverse of the inter-frame CU bitstream parsing process. Figure 18 The following is given. In 1801, the prediction mode is resolved. Then, if inter-frame prediction is used (yes in 1802), the CCP mode is resolved (1803) for the CCP mode selected by the considered CU, providing the CCP mode. This mode can be equivalent to no CCP (meaning the Cb / Cr components are predicted according to the usual inter-frame prediction), CCLM, CCCM, GLM, GLCCM, and any other possible cross-component prediction mode. In one variant, the CCP mode... This is an index indicating the CCP mode in the set of possible CCP modes. In this variant, a specific value of the index indicates that a CCP mode is not used for a given CU.
[0158] In another variant, such as as shown in Table 1, a flag indicating whether CCP mode is used is first signaled, and if the flag indicates so, the index is signaled.
[0159] The decoded CCP information can also include parameters associated with the CCP mode under consideration. As an example, if the CCP mode is CCLM, some slope information of the linear model can also be decoded.
[0160] In 1804 and 1805, the inter-frame prediction mode and associated motion data, as well as the CU residual, are decoded for the current CU (1806).
[0161] Figure 19 The illustration shows an example of the CU decoding and reconstruction process 1900, which is used to decode and reconstruct data from... Figure 18The parsed data from the process is used to reconstruct the CU. In 1901, it is determined whether the current inter-frame CU is encoded in merge mode. If not, in 1902, an AMVP candidate list is constructed, and in 1903, the motion data of the current CU is reconstructed from the selected candidates in the AMVP list and the decoded motion residual. If the CU is encoded in merge mode, in 1904, a merge candidate list is constructed, and in 1905, the motion data of the current CU is reconstructed from the selected candidates in the merge list. These steps can be, for example, similar to the corresponding steps in a VVC or ECM codec implementation. In 1906, the motion data derived in the previous steps is used to predict the CU from the reference block.
[0162] In this embodiment, the cross-component prediction mode is used for inter-frame CU. The CCP mode is specified by the decoded index, which comes from the syntax element inter_ccp_mode() newly introduced in this disclosure (Tables 1 and 2). The CCP mode indicated in the bitstream may include CCLM mode, MMLM mode, CCM mode, GLM mode, GLCCM mode, etc. In 1907, cross-component mode parameters (e.g., taps for CCLM mode) are then determined between the post-predicted luma block and the post-predicted chroma block. The luma is then reconstructed in 1908 (by adding the residual to the prediction), and in 1909, the reconstructed luma and CCP mode are used. To perform CCP to predict chroma blocks. In 1910, the chroma Cb / Cr blocks are reconstructed by adding residuals and the predicted chroma blocks.
[0163] The following provides further embodiments regarding cross-component prediction mode inheritance for inter-frame CU.
[0164] According to another embodiment, the cross-component mode for inter-frame CUs can be inherited from some other CUs that have already been encoded or decoded.
[0165] To this end, a candidate CCP mode list can be constructed for the CU to be decoded based on the CUs that have already been decoded. The candidate list can include spatially adjacent candidates, non-adjacent candidates, history-based candidates from the current image, and CCP modes of possible temporal candidates (typically from the same co-located image used for temporal motion data inheritance). CCP mode derivation can therefore also be applied to temporal merge derivation. In this case, `merge_idx` will indicate the CCP modes of the co-located inter-frame blocks of the current block (if available in the co-located image) used to derive the CCP mode of the current inter-frame block. Note that co-located blocks can be determined in a manner similar to that already done for temporal merge candidate derivation (typically for TMVP or SbTMVP merge candidates for VVC or ECM).
[0166] To identify the CCP mode selected for the current CU, the dedicated ccp_merge_idx syntax element can be used to indicate the selected candidate CCP in the list.
[0167] According to one variant, this merged CCP signal notification mode can be used in the merged and affine merged modes of the video codecs under consideration.
[0168] On the other hand, in the case of AMVP inter-frame CU, explicit signaling is applied to the CCP mode of the current CU.
[0169] The advantage of this embodiment is that, compared with the above embodiment, it further improves compression efficiency by reducing the signal notification cost of CCP mode information.
[0170] Table 3 provides examples of the corresponding syntax for CU prediction data, with added elements shown in bold.
[0171] Table 3: Syntax Table for CU Prediction Data with CCP Data Inheritance
[0172] The corresponding CU prediction and CCP data bitstream parsing (2000) in this embodiment are in Figure 20 The following is given. In 2001, the CU prediction mode is decoded, and in 2002, it is checked whether the CU is encoded in inter-frame mode. Here, only inter-frame CUs are considered. As can be seen, different CCP information parsing is performed based on whether the CU is in merge mode (2003). In the case of AMVP mode, the application and reference are... Figure 18 The provided embodiments use the same parsing (2006, 2007, 2008, 2009). In the merge mode, the CCP merge index (2004) and the general merge index used to derive the CU motion data (2005) are parsed. In 2009, the CU residuals are decoded.
[0173] Based on Figure 20 An example of the CU decoding and reconstruction process (2100) of the motion and CCP data parsed in the solution is shown in Figure 21 The process is as follows: At 2101, it is determined whether the current inter-frame CU is encoded in merge mode. If not, at 2102, an AMVP candidate list is constructed based on the encoding standard or implementation, and at 2103, the motion data of the current CU is reconstructed from the selected candidates in the AMVP list and the decoded motion residuals. At 2104, the CCP mode of the inter-frame CU is derived from the parsed CCP mode index.
[0174] If the CU is encoded in merge mode, then at 2105, a merge candidate list is constructed. At 2106, the CU motion data is derived from the candidates selected in the merge candidate list using the merge index decoded from the CU. In this embodiment, at 2107, a CCP merge candidate list is constructed for inter-frame CUs, and at 2108, the CCP mode is derived from the candidates selected from the CCP merge candidate list based on the parsed CCP merge index.
[0175] The subsequent steps 2109-2113 for reconstructing the current CU are... Figure 19 The steps are similar.
[0176] Some variations of the above embodiments are provided below.
[0177] CCP mode derivation may only apply when the CU is in skip mode. A CU in skip mode is equivalent to a CU encoded in merge mode, except that its associated residual block is uncoded and equal to 0. In this case, the CU inherits the CCP mode from the merge candidate selected from the merge candidate list, similar to motion information.
[0178] In another variant, the CCP mode selected for the CU is the CCP mode that minimizes the distortion between the filtered predicted luma block and the temporally predicted chroma block. In this variant, it is not necessary to signal the CCP mode used for inter-frame CUs, and no new syntax for signaling the CCP mode and parameters is required.
[0179] According to another variant, a single candidate list is built on the encoder and decoder sides, and each candidate in the list includes motion data as well as CCP mode information.
[0180] According to this variant, instead of using the new merge index `ccp_merge_idx` syntax element, it still uses the single merge index already used for motion data to inherit CCP mode information. This variant has the advantage of improved coding efficiency. Figure 22 The image above shows the result. Figure 22 Steps and Figure 21 The steps are similar, except that... Figure 22 Instead of constructing a CCP merge candidate list, the same merge candidate list determined for motion data in 2205 and the parsed merge index are used to derive the CCP pattern in 2207.
[0181] In another variant, not only the CCP mode but also the CCP prediction parameters (e.g., CCCM filter parameters, CCLM slope parameters, etc.) are derived from past CUs indicated in the candidate list by the merge index. In this case, the computed CCP parameters are stored at the CU level and can be reused for decoding subsequent CUs.
[0182] This variant demonstrates the example CU decoding and reconstruction process 2300 in... Figure 23 The above diagram illustrates this. The advantage of this variant is that it saves computation by avoiding the calculation of CCP prediction parameters from post-predicted luma and chroma blocks for each CU. In this variant, CCP parameters are calculated only when the CU is coded in AMVP mode, and when the CU is in merge mode, CCP parameters are derived from CCP candidates.
[0183] Another advantage is that it makes it possible to apply CCPs in inter-frame blocks encoded in merge skip mode, which is not achieved in the prior art. Indeed, by inheriting the merge candidate CCP mode and parameters, CUs in skip mode (i.e., without any associated residual blocks) can still benefit from the potential improvement in chroma prediction by extending the proposed merge mechanism to CCP mode and parameters.
[0184] At 2301, it is determined whether the current inter-frame CU is encoded in merge mode. If not, at 2302, an AMVP candidate list is constructed based on the encoding standard or implementation, and at 2303, the motion data of the current CU is reconstructed from the selected candidates in the AMVP list and the decoded motion residuals. At 2304, the CU is predicted using the derived motion data and reference images. At 2305, the CCP mode is derived from the parsed CCP mode index, and the CCP parameters are determined from the predicted CU.
[0185] If the CU is encoded in a merged mode, then at 2306, a merge candidate list is constructed. At 2307, the CU motion data is derived from the candidate selected in the merge candidate list using the merge index decoded from the CU. In this embodiment, at 2308, the CCP mode is derived from the merge index, which is also used for merging motion. In another variant, the CCP mode can be derived from a CCP mode candidate selected from the CCP merge candidate list using a parsed CCP mode index.
[0186] In any of these variants, in 2309, the CCP prediction parameters for the CCP modes determined in 2308 are derived from candidates indicated by the merge index or the CCP merge index (depending on the variant used). In 2310, CUs are predicted using the derived motion data and reference images, as in 2304.
[0187] The subsequent steps 2311-2314 for reconstructing the current CU are... Figure 19 The corresponding steps for 22 are similar.
[0188] In one variant, these CCP parameters can be inherited only when derivation is performed using spatial adjacency merge candidates.
[0189] In another variant, the candidate CCP mode selected from the CCP merge candidate list can be used only when the merge candidate used to derive motion data corresponds to a spatially adjacent block of the current block. In this variant, non-adjacent merge candidates can be omitted from deriving the CCP mode of inter-frame blocks.
[0190] In another variant, CCP mode derivation may only apply when the CU is in skip mode.
[0191] In a variant of the above embodiment used for merging and deriving CCP modes, the proposed CCP mode signal notification and derivation may be performed only on inter-frame CUs (i.e., blocks encoded in inter-frame mode) and not on blocks or coding units encoded in IBC (intra-block copy) coding mode.
[0192] Alternatively, the proposed CCP mode signal notification and derivation in any of the merging modes described in the above embodiments can be used for both inter-frame and IBC cases.
[0193] Figure 24 The figure illustrates a system block diagram that can implement aspects of this embodiment according to another embodiment. Figure 24 An apparatus 2400 of one embodiment is shown for encoding or decoding video according to any embodiment described herein. The apparatus includes a processor 2410 and is interconnected to a memory 2420 via at least one port. The processor 2410 and the memory 2420 may also have one or more additional interconnects for external connection.
[0194] Processor 2410 is also configured to use any of the embodiments described herein. For example, processor 1220 is configured to use any of the embodiments described herein to: decode one or more syntax elements for a block of video encoded in an inter-frame coding mode; derive a cross-component prediction mode for the video block from the one or more syntax elements; reconstruct the luma component for the video block; and use the cross-component prediction mode to reconstruct one or more chroma components for the video block from the reconstructed luma component. For example, processor 2410 uses a computer program product that includes code instructions implementing any of the embodiments described herein.
[0195] In other embodiments, processor 2410 is configured to use any of the embodiments described herein to: reconstruct the luma component of a block of video encoded in an inter-frame coding mode; determine a cross-component prediction mode for the video block; encode one or more syntax elements for the video block, the one or more syntax elements being used to derive the cross-component prediction mode on the decoder side; and encode the chroma component of the video block based on the reconstructed luma component and the determined cross-component prediction mode. For example, processor 2410 is configured to use a computer program product including code instructions implementing any of the embodiments described herein.
[0196] In one embodiment, such as Figure 25 As shown, in the transmission context of a communication network NET between two remote devices A and B, device A includes a processor associated with RAM and ROM, configured as follows: Figure 1-24 The method for video encoding is described, and device B includes a processor associated with memory RAM and ROM, configured as follows: Figure 1-24 The implementation describes a method for video decoding. As an example, the network is a broadcast network adapted for broadcasting / transmitting encoded video from device A to decoding devices including device B.
[0197] Figure 26 A syntax example of a signal transmitted via a packet-based transport protocol is shown. Each transmitted packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include video data encoded according to any of the embodiments described above.
[0198] Various implementations involve decoding. As used herein, "decoding" can include, for example, performing on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of various embodiments described herein, such as entropy decoding of a sequence of binary symbols to reconstruct image or video data.
[0199] As a further example, in one embodiment "decoding" refers only to entropy decoding, in another embodiment "decoding" refers only to differential decoding, in yet another embodiment "decoding" refers to a combination of entropy decoding and differential decoding, and in yet another embodiment "decoding" refers to the entire image reconstruction process including entropy decoding. It will be clear, and believed to be readily understood, by those skilled in the art, that the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process, based on the specific context of the description.
[0200] Various implementations involve encoding. Similar to the discussion of "decoding" above, "encoding" as used in this application can include, for example, performing on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such processes also or alternatively include processes performed by encoders of various embodiments described in this application, such as determining resampling filter coefficients and resampling the decoded image.
[0201] As a further example, in one embodiment "encoding" refers only to entropy encoding, in another embodiment "encoding" refers only to differential encoding, and in yet another embodiment "encoding" refers to a combination of differential and entropy encoding. It will be clear, and believed that those skilled in the art, whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process, based on the specific context of the description.
[0202] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.
[0203] This disclosure describes various types of information, such as syntax, which may be transmitted or stored. This information can be packaged or arranged in various ways, including those common in video standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or slice headers), or SEI messages. Other methods are also available, including those common in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions, used for session announcement and session invitation purposes, such as described in RFCs and used with RTP (Real-Time Transport Protocol) transmission.
[0204] b. DASH MPD (Media Presentation Description) descriptors, such as those used in DASH and transmitted via HTTP, are associated with representations or sets of representations to provide additional features for the content representation.
[0205] c. RTP header extensions, for example, used during RTP streaming.
[0206] d. ISO-based media file formats, such as those used in OMAF, and using object-oriented building blocks defined by unique type identifiers and lengths, also known as "atoms" in some specifications.
[0207] e. An HLS (HTTP Live Streaming) manifest, which is transmitted over HTTP. A manifest can, for example, be associated with a version or set of versions of content to provide characteristics of that version or set of versions.
[0208] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0209] Some implementations involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually with constraints on computational complexity. Rate-distortion optimization is generally formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate-distortion optimization problem. For example, methods can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, where their encoding costs and associated distortion of the reconstructed signal after encoding and decoding are fully evaluated. Faster approaches can also be used to save encoding complexity, particularly by calculating approximate distortion based on prediction or prediction of the residual signal (rather than the reconstructed signal). A hybrid of these approaches can also be used, such as using approximate distortion only for some possible encoding options and full distortion for others. Other approaches evaluate only a subset of possible encoding options. More generally, many methods employ multiple techniques to perform optimization, but optimization is not necessarily a complete evaluation of encoding costs and associated distortion.
[0210] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatus or program). An apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor, which refers to a general processing device, such as including a computer, microprocessor, integrated circuit, or programmable logic device. A processor also includes communication devices, such as computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0211] The references to "an embodiment," "an embodiment," "an implementation," or "implementation," and other variations, mean that a particular feature, structure, characteristic, etc., described in the connected embodiments is included in at least one embodiment. Therefore, the appearance of the phrases "in an embodiment," "in an embodiment," "in an implementation," or "in an implementation," and any other variations in different locations within this application, does not necessarily all refer to the same embodiment.
[0212] Additionally, this application may involve "determining" various types of information. Determining information may include one or more of the following, such as estimation information, calculation information, prediction information, or information retrieved from memory.
[0213] Furthermore, this application may involve "accessing" various types of information. Accessed information may include one or more of the following, such as receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0214] Additionally, this application may involve "receiving" various types of information. "Receiving," like "accessing," is intended to be a broad term. Receiving information may include one or more of the following, for example, accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" typically involves some form of operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0215] It should be understood that any use of “ / ”, “and / or”, and “at least one” in the following contexts, such as “A / B”, “A and / or B”, and “at least one of A and B”, is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the contexts of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to any number of items listed.
[0216] Additionally, as used herein, the term "signal" refers to, among other things, instructing the corresponding decoder to do something. Thus, in embodiments, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can send (e.g., explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has that specific parameter as well as other parameters, it can simply inform the decoder and select that specific parameter without sending it (e.g., implicitly signal). Bit savings are achieved in various embodiments by avoiding the transmission of any actual function. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. While the foregoing refers to the verb form of the term "signal," the term "signal" may also be used herein as a noun.
[0217] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information. The information may, for example, include instructions for performing a method or data generated by one of the implementations. For example, the signal may be formatted to carry a bit stream of the embodiments described. Such signals may, for example, be formatted as electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may, for example, include encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may, for example, be analog or digital information. The signal can be transmitted via a variety of different wired or wireless links as known. The signal may be stored on a processor-readable medium.
[0218] Many embodiments have been described. The features of these embodiments may be provided individually or in any combination across a variety of claim classes and types.
Claims
1. A method comprising: Decode one or more syntax elements for a video block encoded in an inter-frame coding mode; Use the one or more syntax elements to determine the cross-component prediction mode for the video block from the cross-component prediction mode set; Reconstruct the luminance components used for the video block; The cross-component prediction mode is used to reconstruct one or more chroma components for the video block from the reconstructed luminance components.
2. An apparatus comprising one or more processors, wherein the one or more processors are operable to: Decode one or more syntax elements for a video block encoded in an inter-frame coding mode; Use the one or more syntax elements to determine the cross-component prediction mode for the video block from the cross-component prediction mode set; Reconstruct the luminance components used for the video block; The cross-component prediction mode is used to reconstruct one or more chroma components for the video block from the reconstructed luminance components.
3. A method comprising: Reconstruct the luminance component of a video block, which is encoded in an inter-frame coding mode; Determine the cross-component prediction mode for the video block; Encoding one or more syntax elements for the video block, the one or more syntax elements being used to derive cross-component prediction modes from a set of cross-component prediction modes on the decoder side; One or more chroma components of the video block are encoded based on the reconstructed luminance components and the determined cross-component prediction mode.
4. An apparatus comprising one or more processors, wherein the one or more processors are operable to: Reconstruct the luminance component of a video block, which is encoded in an inter-frame coding mode; Determine the cross-component prediction mode for the video block; Encoding one or more syntax elements for the video block, the one or more syntax elements being used to derive cross-component prediction modes from a set of cross-component prediction modes on the decoder side; One or more chroma components of the video block are encoded based on the reconstructed luminance components and the determined cross-component prediction mode.
5. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein the one or more syntax elements include an index indicating a cross-component prediction pattern in the set of cross-component prediction patterns.
6. The method of claim 1, 3 or 5, or the apparatus of claim 2, 4 or 5, wherein the set of cross-component prediction modes includes at least one of a convolutional cross-component model, a cross-component linear model, a multi-model linear model, and a gradient linear model.
7. The method or apparatus of claim 5 or 6, wherein the cross-component prediction mode set is constructed for the video block based on one or more reconstructed video blocks.
8. The method or apparatus of claim 7, wherein the one or more reconstructed blocks include at least one of spatially adjacent blocks, non-adjacent blocks, history-based blocks, and temporal blocks from reference images.
9. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein a template of the video block is used to determine a cross-component prediction mode for the video block.
10. The method or apparatus of any one of claims 5 to 9, wherein the one or more syntax elements include a flag indicating whether a cross-component prediction mode is used for the video block, and wherein the cross-component prediction mode is determined in response to determining that the flag indicates that a cross-component prediction mode is used for the video block.
11. The method or apparatus of any one of claims 5 to 10, wherein the one or more syntax elements include one or more parameters associated with the cross-component prediction mode.
12. The method or apparatus of claim 5 or 6, wherein the cross-component prediction mode set is a table known to the decoder.
13. The method or apparatus of claim 7 or 8, wherein in response to determining that the video block is in merge mode or intra-block copy mode, the one or more syntax elements include an index indicating a candidate block in the merge candidate block list for deriving motion data for the video block.
14. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein, in response to determining that the video block is in AMVP mode, the cross-component prediction mode is decoded from the one or more syntax elements together with motion data for the video block.
15. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein in response to determining that the video block is in a merge mode or a skip mode, motion data of the video block is derived from a selected candidate block in a merge candidate block list, and the cross-component prediction mode is inherited from the selected candidate block.
16. The method or apparatus of claim 13, wherein in response to determining that a selected candidate block is a spatially adjacent block of the video block, the cross-component prediction mode is inherited from the selected candidate block.
17. The method of any one of claims 1, 3, and 5 to 16, or the apparatus of any one of claims 2 and 4 to 16, wherein the set of cross-component prediction modes depends on the inter-frame prediction mode of the video block.
18. The method of any one of claims 1, 3, and 5 to 17, or the apparatus of any one of claims 2 and 4 to 17, wherein the cross-component prediction modes in the cross-component prediction mode set are reordered based on a cost determined on a template of the video block.
19. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein the derivation of the cross-component prediction mode for the video block is based on the distortion determined between the chroma component of a reference block of the video block and the chroma component predicted from the luma component of the reference block using the cross-component prediction mode.
20. A computer program product comprising instructions for causing one or more processors to perform the method as described in any one of claims 1, 3, and 5 to 19.
21. A non-transitory computer-readable medium storing executable program instructions for causing a computer executing the program instructions to perform the method according to any one of claims 1, 3, and 5 to 19.
22. A bitstream comprising data representing a video encoded using the method described in any one of claims 1, 3, and 5 to 19.
23. A non-transitory computer-readable medium storing the bit stream as described in claim 22.
24. An apparatus comprising: The apparatus as described in claim 2; and At least one of the following: (i) an antenna configured to receive or transmit a signal including data representing the video block; (ii) a band limiter configured to limit the signal to a frequency band including data representing the video block; and (iii) a display configured to display the video block.
25. The device of claim 24, wherein the device comprises at least one of a television, a mobile phone, a tablet computer, and a set-top box.