Cross-component residual prediction using residual templates

By using residual template cross-component residual model (RT-CCRM) and cross-component filtering in video encoding technology, the problem of insufficient redundancy utilization of video data in the prior art is solved, and more efficient video encoding and storage is achieved.

CN120077643APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004336.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2024-04-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing video encoding technologies transmit or store video data, it is difficult to effectively utilize the redundancy in the video data, resulting in waste of bandwidth and storage space.

Method used

Residual template cross-component residual model (RT-CCRM) is used to reconstruct the image frame of the video code stream by generating residuals of chroma samples, and apply cross-component filtering in the residual domain to achieve local lighting compensation and improve coding efficiency.

Benefits of technology

It improves the encoding efficiency of video content, reduces the bandwidth and storage space requirements, and enhances the fidelity of video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077643A_ABST
    Figure CN120077643A_ABST
Patent Text Reader

Abstract

Various embodiments described herein include methods and systems for encoding and decoding a video. In one aspect, a video bitstream includes a current image frame having a current coding block, and a first syntax element for a residual template cross-component residual model (RT-CCRM) mode is signaled. When the RT-CCRM mode is enabled, the computing system identifies, in the current coding block, a first chroma sample and one or more luma samples corresponding to the first chroma sample; determining one or more residuals of one or more luma samples in the current coding block; a residual filter corresponding to the RT-CCRM mode is applied to generate a first residual of the first chroma sample based on the residual of the one or more luma samples. The computing system reconstructs the first chroma sample to reconstruct the current image frame by compensating the predicted chroma sample with at least the first residual.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 461,581, filed on April 24, 2023, with the title "Cross-Component Residual Prediction by Using Residual Template", and is a continuation-in-part of, and claims priority to, U.S. Patent Application No. 18 / 643,953, filed on April 23, 2024, with the title "Cross-Component Residual Prediction by Using Residual Template". Technical Field

[0003] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for processing video data using cross-component residuals. Background Art

[0004] Various electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive digital video data over a communication network or otherwise transmit digital video data, and / or store digital video data on a storage device. Since the bandwidth capacity of the communication network is limited and the memory resources of the storage device are limited, video data can be video coded and compressed according to one or more video coding standards before being transmitted or stored. Video coding can be performed by hardware and / or software on an electronic device / client device or a server providing cloud services.

[0005] Video coding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) to exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended to be the successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). Alliance for Open Media (AOMedia) Video 1 (AV1) is an open video coding format designed as a replacement for HEVC. On January 8, 2019, a verified version 1.0.0 of this specification with Errata 1 was released. Summary of the invention

[0006] As mentioned above, encoding (compression) reduces bandwidth and / or storage space requirements. As described in detail below, both lossless compression and lossy compression can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully retained during encoding and cannot be fully restored during decoding. When using lossy compression, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. The amount of allowable distortion depends on the application. For example, users of certain consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher allowable distortions generally allow the use of encoding algorithms that produce higher losses and higher compression ratios.

[0007] The present disclosure describes a video compression method that uses a residual template cross-component residual model (RT-CCRM) to generate residuals of chrominance samples based on the residuals of associated luminance samples for reconstructing image frames of a video bitstream. Cross-component filtering is applied in the residual domain to predict the residuals of chrominance samples using the residuals of luminance samples associated with the chrominance samples, which provides local illumination compensation and enhances the coding efficiency of video content. Specifically, in some embodiments, a current coding block of a current image frame has a current template including a current luminance template and a current chrominance template. The current coding block corresponds to one or more reference coding blocks 436, such as located on the current image frame or different reference image frames. Each reference coding block 436 has a corresponding reference template that includes a reference luminance template and a reference chrominance template. The reference luminance template, the reference chrominance template, the current luminance template, and the current chrominance template are applied to determine the residual data of luminance samples and chrominance samples. The residual data of the luminance samples and chrominance samples of these templates is further applied to determine the filter coefficients of a residual filter applied in the RT-CCRM. The residual filter with the filter coefficients is used to determine the residual data of the chrominance samples of the current coding block based on the residual data of the luminance samples of the current coding block. The current coding block is reconstructed based on the determined residual data of the chrominance samples.

[0008] According to some embodiments, a method of video decoding is provided. The method includes: receiving a video bitstream including a current image frame and a first syntax element for a residual template cross-component residual model (RT-CCRM) mode. The method further includes: based on the first syntax element, determining to enable the RT-CCRM mode to generate a first residual of a first chrominance sample of a current coding block based on one or more residuals of one or more luminance samples in the current coding block of the current image frame. The method further includes: when the RT-CCRM mode is enabled, identifying the first chrominance sample and one or more luminance samples corresponding to the first chrominance sample in the current coding block; determining one or more residuals of one or more luminance samples in the current coding block; and applying a residual filter corresponding to the RT-CCRM mode to generate the first residual of the first chrominance sample based on the one or more residuals of the one or more luminance samples. The method further includes: reconstructing the first chrominance sample at least by compensating a predicted chrominance sample with the first residual to reconstruct the current image frame.

[0009] According to some embodiments, a method for video encoding is provided. The method includes: receiving video data including a current picture frame; encoding the current picture frame including a current coding block; and determining whether to enable a Residual Template Cross-Component Residual Model (RT-CCRM) to generate a first residual of a first chroma sample of the current coding block based on one or more residuals of one or more luma samples in the current coding block of the current picture frame. The method further includes: transmitting the encoded current picture frame via a video bitstream; and signaling a first syntax element via the video bitstream to indicate whether the RT-CCRM mode is enabled to generate a first residual of a first chroma sample of the current coding block based on one or more residuals of one or more luma samples in the current coding block of the current picture frame.

[0010] According to some embodiments, a method for bitstream conversion is provided. The method includes: obtaining a source video sequence including a current coding block of a current picture frame, and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current picture frame and a first syntax element for a Residual Template Cross-Component Residual Model (RT-CCRM) mode, the first syntax element indicating whether to generate a first residual of a first chroma sample of the current coding block based on one or more residuals of one or more luma samples in the current coding block of the current picture frame. According to determining that the first syntax element indicates enabling the RT-CCRM mode, applying a residual filter corresponding to the RT-CCRM mode to generate a first residual of the first chroma sample based on the one or more residuals of the one or more luma samples.

[0011] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes control circuitry and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any one of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).

[0012] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any one of the methods described herein.

[0013] Accordingly, apparatuses, systems, and methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all inclusive, and in particular, given the figures, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Further, it should be noted that the language used in the specification has been selected primarily for readability and guidance purposes and not necessarily to delineate or circumscribe the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] For a more detailed understanding of the present disclosure, reference may be made to the features of various embodiments, some of which are illustrated in the figures. However, the figures illustrate only the relevant features of the present disclosure and are therefore not necessarily considered limiting, as those skilled in the art will appreciate that other effective features may be present upon reading this disclosure.

[0015] Figure 1 is a block diagram illustrating an example communication system in accordance with some embodiments.

[0016] Figure 2A is a block diagram illustrating example elements of an encoder component in accordance with some embodiments.

[0017] Figure 2B is a block diagram illustrating example elements of a decoder component in accordance with some embodiments.

[0018] Figure 3 is a block diagram illustrating an example server system in accordance with some embodiments.

[0019] Figure 4A is a schematic diagram of an example RT-CCRM applied in a residual template cross-component residual model (RT-CCRM) mode in accordance with some embodiments.

[0020] Figure 4B is an example current image frame processed in an RT-CCRM mode in accordance with some embodiments.

[0021] Figure 5 is a diagram illustrating a current image frame 410 and two associated reference image frames in accordance with some embodiments.

[0022] Figure 6 is a diagram illustrating a template matching process performed within a search range around an initial motion vector in accordance with some embodiments.

[0023] Figure 7 is a schematic diagram of a CCRM applied in a cross-component residual model (CCRM) mode in accordance with some embodiments.

[0024] Figure 8 is a flowchart illustrating another example method of encoding a video according to some embodiments.

[0025] In accordance with conventional practice, the various features illustrated in the drawings need not be drawn to scale, and like reference numerals may be used throughout the specification and drawings to designate like features. Detailed Description

[0026] The present disclosure describes a video compression method that uses a residual template cross-component residual model (RT-CCRM) to generate a residual of chrominance samples based on a residual of associated luminance samples for reconstructing an image frame of a video bitstream. In some embodiments, a current coding block of a current image frame has a current template including a current luminance template and a current chrominance template. The current coding block corresponds to one or more reference coding blocks 436, such as located on the current image frame or a different reference image frame. Each reference coding block 436 has a corresponding reference template that includes a reference luminance template and a reference chrominance template. Residual data of luminance samples and chrominance samples are determined using the reference luminance template, the reference chrominance template, the current luminance template, and the current chrominance template. Further, the residual data of the luminance samples and the chrominance samples of these templates are applied to derive filter coefficients of a residual filter applied in the RT-CCRM. According to the RT-CCRM, the residual filter with the filter coefficients is used to determine residual data of chrominance samples of the current coding block based on the residual data of the luminance samples of the current coding block. Based on the determined residual data of the chrominance samples, the current coding block is reconstructed. Thus, cross-component filtering can be applied in the residual domain to predict the residual of chrominance samples using the residual of associated luminance samples, thereby facilitating local illumination compensation and enhancing the coding efficiency of video content.

[0027] Figure 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, such as used with video applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0028] The source device 102 includes a video source 104 (e.g., a camera component or a media storage space) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., for creating an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may have a larger data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth to transmit and less storage space to store compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to the network 110).

[0029] One or more networks 110 represent any number of networks for passing information between the source device 102, the server system 112, and / or the electronic device 120, including for example cable (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0030] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as the encoded video bitstreams from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode the encoded video bitstreams 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstreams 108.

[0031] In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstreams 108 in order to customize potentially different bitstreams for one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0032] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video bitstream that can be reproduced on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicate with an external display device and / or include media storage space). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0033] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user equipment". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0034] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may use the codec component 114 to decode and / or encode the encoded video bitstream 108. For example, the server system 112 may encode the video data in a manner more suitable for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0035] Figure 2AFIG. is a block diagram illustrating example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive video data (e.g., the source video sequence). In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a different device component from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously acquired / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that, when viewed in sequence, create a sense of motion. The pictures themselves may be organized as a spatial array of pixels, where each pixel may have one or more samples, depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples.

[0036] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a visual media data stream (e.g., a video stream). Implementing an appropriate encoding speed is one of the functions of the controller 204. In some embodiments, the controller 204 controls the other functional units described below and is functionally coupled to the other functional units. The parameters set by the controller 204 may include rate control-related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. A person of ordinary skill in the art can easily identify other functions of the controller 204, as they may involve the encoder component 106 optimized for a specific system design.

[0037] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. Thus, the prediction part of the encoder interprets the same sample values as the decoder uses during prediction in the decoding process as reference picture samples. The principle of this reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those of ordinary skill in the art.

[0038] The operation of the decoder 210 can be the same as that of a remote decoder (such as the decoder component 122), which is described in detail below in connection with Figure 2B the decoder component 122. Briefly referring to Figure 2B , however, since the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into an encoded video sequence losslessly, the entropy decoding part of the decoder component 122 (including the buffer memory 252 and the parser 254) may not be fully implemented in the local decoder 210.

[0039] The decoder techniques described in this application, except for parsing / entropy decoding, can exist in the corresponding encoder in substantially the same functional form. For this reason, this application focuses on decoder operations. The description of encoder techniques can be simplified because encoder techniques can be inverse to decoder techniques.

[0040] As part of its operation, the source encoder 202 can perform motion-compensated predictive coding, predicting and encoding an input frame by referring to one or more previously encoded frames designated as reference frames in the video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, and the reference frame can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0041] The decoder 210 can decode the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is in a video decoder ( Figure 2AWhen decoding at the (not shown in the figure), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process, which can be performed by a remote video decoder on the reference frames, and can cause the reconstructed reference frames to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores copies of the reconstructed reference frames, which have the same content (without transmission errors) as the reconstructed reference frames to be obtained by the remote video decoder.

[0042] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search in the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have a prediction reference extracted from multiple reference pictures stored in the reference picture memory 208.

[0043] The outputs of all the above functional units can undergo entropy encoding in the entropy encoder 214. The entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0044] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by the entropy encoder 214 to prepare for transmission via the communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted (e.g., encoded audio data and / or an auxiliary data stream (not shown in the source)). In some embodiments, the transmitter can send additional data and the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data (such as redundant pictures and slices), supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0045] The controller 204 can manage the operation of the encoder component 106. During encoding, the controller 204 can assign a certain encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow the use of different types of intra pictures, including for example independent decoder refresh (IDR) pictures. Those skilled in the art know those variants of I pictures and their respective applications and characteristics, and thus will not be repeated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0046] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded on a block-by-block basis. These blocks can be prediction-encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be encoded non-predictively, or can be predictively encoded with reference to encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded non-predictively by spatial prediction or by temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture can be encoded non-predictively by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0047] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra picture prediction (often abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded with a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0048] The encoder component 106 may perform encoding operations according to a predetermined video encoding technique or standard (such as any technique or standard described in this disclosure). In its operation, the encoder component 106 may perform various compression operations, including predictive encoding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0049] Figure 2B FIG. is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in FIG. is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0050] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data and other data (e.g., encoded audio data and / or auxiliary data streams), and these other data may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data and the encoded video. The additional data may be included as part of one or more encoded video sequences. The decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The decoder component 122 may be implemented at least partially in software.

[0052] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., the buffer memory 252 is configured to handle playout timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, the buffer memory 252 may not be needed, or the buffer memory 252 can be small. When used on a best-effort packet network such as the Internet, the buffer memory 252 may be needed, the buffer memory 252 can be relatively large, and / or have an adaptive size, and can be implemented at least partially in an operating system or a similar element outside the decoder component 122.

[0053] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols can include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device can be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set segment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard, and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup can include a group of pictures (GOP), a picture, a tile, a stripe, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser 254 can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0054] The reconstruction of the symbols 270 can involve multiple different units, depending on the type of the encoded video picture or its part (such as: inter-picture and intra-picture, inter-block and intra-block) and other factors. Which specific units are involved and how they are involved can be controlled by subgroup control information, which is parsed by the parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between the parser 254 and the multiple units below is not described.

[0055] The decoder component 122 can be conceptually divided into multiple functional units, and in some embodiments, many of these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual division of the functional units is still retained.

[0056] The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as which transform mode to use, block size, quantization factor, and / or quantization scaling matrix) as one or more symbols 270 from the parser 254. The scaler / inverse transform unit 258 can output a block including sample values, and these sample values can be input into the aggregator 268.

[0057] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. This predictive information can be provided by the intra prediction unit 262. The intra prediction unit 262 can use the surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 can add the predictive information already generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.

[0058] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion compensating the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (which is called the residual sample or residual signal in this case), thereby generating output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 obtains the prediction samples can be controlled by a motion vector. The motion vector can be provided to the motion compensation prediction unit 260 in the form of symbols 270, and these symbols 270 can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values obtained from the reference picture memory 266, a motion vector prediction mechanism, etc. when using sub-sampled accurate motion vectors.

[0059] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and provided to loop filter unit 256 as symbols 270 from parser 254, but can also respond to meta-information obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0060] The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124 and stored in reference picture memory 266 for subsequent inter-picture prediction.

[0061] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture has been fully reconstructed and the encoded picture has been identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266 and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0062] Decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be specified in a standard such as any of the standards described in this disclosure. The encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that it follows the syntax of the video compression technique or standard as specified in the video compression technique documentation or standard and specifically in the profile documentation therein. Moreover, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within bounds defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the assumptions of the hypothetical reference decoder (HRD) buffer management and metadata signaled in the encoded video sequence can further limit the limits set by the level.

[0063] Figure 3FIG. is a block diagram of a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0064] One or more network interfaces 304 may be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication network may be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay tolerant network, and so on. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, and so on. Such communication may be unidirectional receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using local area digital networks or wide area digital networks). Such communication may include communication to one or more cloud computing networks.

[0065] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. One or more input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. One or more output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.

[0066] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices located remotely from the control circuitry 302. The memory 314 or optionally one or more non-volatile solid-state memory devices within the memory 314 includes non-volatile computer-readable storage media. In some embodiments, the memory 314 or the non-volatile computer-readable storage media of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof:

[0067] ● An operating system 316, including processes for handling various basic system services and performing hardware-related tasks;

[0068] ● A network communication module 318, for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections);

[0069] ● A codec module 320, for performing various functions related to encoding and / or decoding of data (such as video data). In some embodiments, the codec module 320 is an example of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following:

[0070] ○ A decoding module 322, for performing various functions related to decoding of encoded data, such as the functions related to the decoder component 122 described previously; and

[0071] ○ An encoding module 340, for performing various functions related to encoding of data, such as the functions related to the encoder component 106 described previously; and

[0072] ● A picture memory 352, for storing pictures and picture data, e.g., for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0073] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions related to the parser 254 described previously), a transformation module 326 (e.g., configured to perform various functions related to the scaler / inverse transformation unit 258 described previously), a prediction module 328 (e.g., configured to perform various functions related to the motion compensation prediction unit 260 and / or the intra prediction unit 262 described previously), and a filter module 330 (e.g., configured to perform various functions related to the loop filter 256 described previously).

[0074] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions related to the source encoder 202 and / or the encoding engine 212 described previously) and a prediction module 344 (e.g., configured to perform various functions related to the predictor 206 described previously). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0075] Each of the above-identified modules stored in the memory 314 corresponds to a set of instructions for performing the functions described in this disclosure. The above-identified modules (e.g., sets of instructions) need not be implemented as separate software programs, routines, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.

[0076] Although Figure 3 illustrates a server system 112 according to some embodiments, Figure 3 it is more intended as a functional description of the various features that may be present in one or more server systems than as a structural schematic of the embodiments described in this disclosure. In practice, and as will be appreciated by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112, and how the features are distributed among them, will vary from one implementation to another and will optionally depend in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods.

[0077] Figure 4A is a schematic diagram of an example RT-CCRM 400 applied in the RT-CCRM mode according to some embodiments. Figure 4B is an example current image frame 410 processed in the RT-CCRM mode according to some embodiments. The GOP includes a sequence of image frames, which further includes the current image frame 410. The current image frame 410 may include a color image, i.e., a non-monochrome image frame, having a plurality of color samples (e.g., chrominance samples 402 and luminance samples 404) that are co-located with each other. In some embodiments, in the RT-CCRM mode, the RT-CCRM 400 is applied to encode the current coding block 406C of the current image frame 410. The video decoder 122 ( Figure 2B)Receives video bitstream 116, which includes current picture frame 410 and first syntax element 408 for RT-CCRM mode. Based on the first syntax element 408, video decoder 122 determines to enable RT-CCRM mode to generate first residual 412A of first chroma sample 402A of current coded block 406C based on one or more residuals 414L of one or more luma samples 404L (e.g., Figure 4B L0 to L6 in

[0078] corresponding to first chroma sample 402A). When RT-CCRM mode is enabled, video decoder 122 identifies first chroma sample 402A and one or more luma samples 404L corresponding to the first chroma sample 402A in current coded block 406C. Video decoder 122 determines one or more residuals 414L of one or more luma samples 404 in current coded block 406C, and applies residual filter 416 corresponding to RT-CCRM mode to generate first residual 412A of first chroma sample 402A based on one or more residuals 414L of one or more luma samples 404L (e.g., Figure 4B L0 to L6 in

[0079] ). In some embodiments, first residual 412A of first chroma sample 402A is clipped within a dynamic range defined between a first residual value and a second residual value, and the second residual value is greater than the first residual value. Video decoder 122 reconstructs first chroma sample 402A by compensating predicted chroma sample 402P with at least first residual 412A to reconstruct current picture frame 410. Figure 2A ) to generate video bitstream 116. Video bitstream 116 includes current picture frame 410 and first syntax element 408 for RT-CCRM mode, and the first syntax element 408 indicates whether to generate first residual 412A of first chroma sample 402A of current coded block 406C based on one or more residuals 414L of one or more luma samples 404L in current coded block 406C of current picture frame 410. During decoding, according to determining that the first syntax element 408 indicates enabling RT-CCRM mode, residual filter 416 corresponding to RT-CCRM mode is applied to generate first residual 412A of first chroma sample 402A based on one or more residuals 414L of one or more luma samples 404L.

[0080] In some embodiments, on the video encoding side, video encoder 106 (Figure 2A )Generate a video bitstream 116 based on video data including a current image frame 410, and provide the video bitstream 116 to a video decoder 122 Figure 2B ). After obtaining the video data, a video encoder 106 encodes a current image frame 410 including a current coding block 406C, and determines whether to enable RT-CCRM to generate a first residual 412A of a first chrominance sample 402A of the current coding block based on one or more residuals 414L of one or more luma samples 404L in the current coding block 406C of the current image frame. The encoded current image frame 410 is transmitted via the video bitstream 116, wherein a first syntax element 408 is signaled to indicate whether to enable the RT-CCRM mode to generate a residual of the first chrominance sample of the current coding block 406C based on one or more residuals 414L of one or more luma samples 404L in the current coding block 406C of the current image frame.

[0081] In addition, in some embodiments, the current coding block 406C is encoded in one of an inter prediction mode, an intra block copy (IBC) mode, and an intra template matching prediction (IntraTMP) mode. When the current coding block 406C is encoded in one of the inter prediction mode, the IBC mode, and the IntraTMP mode, the first syntax element 408 is signaled. When the video decoder 122 detects the first syntax element 408, the video decoder 122 determines that the current coding block 406C is encoded in one of the inter prediction mode, the IBC mode, and the IntraTMP mode. In addition, in some embodiments, it is determined to disable the cross-component residual model (CCRM) mode to abort predicting a first chrominance sample based on a reconstructed luma sample corresponding to one or more luma samples. When the current coding block 406C is encoded in one of the inter prediction mode, the IBC mode, and the IntraTMP mode and when the CCRM mode is disabled, the first syntax element 408 is signaled. In other words, in some embodiments, when the video decoder 122 detects the first syntax element 408, the video decoder 122 determines to disable the CCRM mode.

[0082] In some embodiments, based on template availability and the position of the coding block in a picture, sub-picture, tile, or strip of the current picture frame, the first syntax element 408 is reset to a value indicating that the RT-CCRM mode is disabled. In an example, for the current coding block 406C, the first syntax element 408 is reset. In another example, for a coding block different from the current coding block 406C, the first syntax element 408 is reset. In some cases, the reference coding block 436 is located at the upper left corner of the current picture frame 410 without any template. The first syntax element 408 indicates that the RT-CCRM mode is disabled for the current coding block 406C encoded based on the reference coding block 436, and the current coding block 406C disables the RT-CCRM mode. In some embodiments, when the current coding block 406C is located adjacent to the top boundary of the current picture frame 410, the first syntax element 408 indicates that the RT-CCRM mode is disabled for the current coding block 406C, and the current coding block 406C disables the RT-CCRM mode.

[0083] In some embodiments, the first syntax element 408 for the RT-CCRM mode includes a first flag 408A and a second flag 408B. The first flag 408A indicates whether to enable the RT-CCRM mode to generate the blue chrominance residual of the blue differential (Cb) chrominance sample 402Cb, and the second flag 408B indicates whether to enable the RT-CCRM mode to generate the red chrominance residual of the red differential (Cr) chrominance sample 402Cr. The first chrominance sample 402A includes the Cb chrominance sample 402Cb and the Cr chrominance sample 402Cr.

[0084] In some embodiments, the video decoder 122 determines to predict the current coding block 406C using an alternative filter 418, and the alternative filter 418 corresponds to one of a multi-model linear model (MMLM) mode, a cross-component linear model (CCLM) mode, a convolutional cross-component intra prediction model (CCCM), and a gradient linear model (GLM) mode. The alternative filter 418 has a plurality of first filter coefficients 420. The video decoder 122 (e.g., the cross-component residual filter coefficient derivation module 422) determines a plurality of filter coefficients 424 of the residual filter 416 based on the plurality of first filter coefficients 420 of the alternative filter 418.

[0085] In some embodiments, the residual filter 416 includes a first residual filter 416A for the first chrominance sample among the blue difference (Cb) chrominance samples 402Cb and the red difference (Cr) chrominance samples 402Cr. The video decoder 122 determines to predict the current encoded block 406C using a first alternative filter 418A and a second alternative filter 418B, where the first alternative filter 418A and the second alternative filter 418B correspond to two different modes among the MMLM mode, the CCLM mode, the CCCM, and the GLM mode. The video decoder 122 (e.g., module 422) determines that the first alternative filter 418A has a plurality of first filter coefficients 420A, and the second alternative filter 418B has a plurality of second filter coefficients 420B. Based on the plurality of first filter coefficients 420A of the first alternative filter 418A, a plurality of filtering coefficients 424 of the first residual filter 416A are determined. Based on the plurality of second filter coefficients 420B of the second alternative filter 418B, a plurality of filtering coefficients 424 of the second residual filter 416B are determined. The second residual filter 416B is applied to generate the residual of the second different chrominance sample among the Cb chrominance samples 402Cb and the Cr chrominance samples 402Cr.

[0086] Reference Figure 4A , in some embodiments, the video decoder 122 includes a luminance predictor (also referred to as the predicted luminance sample 404P), which provides the predicted luminance sample 404P that is predicted based on one or more reference encoded blocks 436 located in the current image frame 410 or in one or more other image frames in the GOP including the current image frame 410. In some embodiments, the video decoder 122 includes a chrominance predictor (also referred to as the predicted chrominance sample 402P), which provides the predicted chrominance sample 402P that is predicted based on one or more reference encoded blocks 436. The predicted chrominance sample 402P corresponding to the first chrominance sample 402A is compensated with a first residual 412A determined based on one or more luminance residuals 414L to generate a compensated chrominance sample 402AC. In some embodiments, the compensated chrominance sample 402AC is further adjusted based on a second residual 430A to reconstruct the first chrominance sample 402A.

[0087] In some embodiments, the current coding block 406C corresponds to the current template 434, which includes the reconstructed portion of the current image frame 410, including the reconstructed sample block adjacent to the current coding block 406C. In some embodiments, the current template 434 is selected from a top current template 434T located at the top of the current coding block 406C, a left current template 434L located to the left of the current coding block 406C, and an L-shaped current template (not shown) that shares the top edge and the left edge of the current coding block 406C. Further, in some embodiments, each reference coding block 436 of the current coding block 406C also has a reference template 438, which has the same shape and the same size as the current template 434. Template matching is a decoder-side block vector derivation method for finding the closest match between the current template 434 of the current coding block 406C and the reference template 438, where the reference template 438 is in a reconstructed image frame different from the current image frame 410 or in the reconstructed portion of the current image frame 410. A template matching cost is applied to determine whether the reference template 438 is one of the closest matches of the current template 434. In an example, the sum of absolute differences (SAD) of corresponding samples of the reference template 434 and the current template 434 is used to determine the template matching cost.

[0088] In some embodiments, the current template 434 includes a filter template applied to determine a first filter (e.g., Figure 4A an alternative filter 418 in ) that corresponds to one of an inter prediction mode, an intraTMP mode, an MMLM mode, a CCLM mode, a CCCM mode, and a GLM mode. For example, applying the same top current template 434T determines the filter coefficients 420 of one or more alternative filters 418 and the filter coefficients 424 of the residual filter 416.

[0089] In some embodiments, for the current template 434, the top current template 434T has at least one row of corresponding color samples (e.g., corresponding to the number of rows), and the left current template 434L has at least one column of corresponding color samples (e.g., corresponding to the number of columns). The reference template 438 has the same shape and the same size as the current template 434. For the reference template 438, the top reference template 438T has at least one row of corresponding color samples, and the left reference template 438L has at least one column of corresponding color samples. In some embodiments, the number of rows is equal to the number of columns. Conversely, in some embodiments, the number of rows is not equal to the number of columns.

[0090] In some embodiments, the reference template 438 and the current template 434 correspond to one of a plurality of predefined template types. Additionally, in some embodiments, the plurality of predefined template types at least includes: a first type of reconstructed top adjacent region (e.g., templates 438T and 434T), a second type of reconstructed left adjacent region (e.g., templates 438L and 434L), and a third type of reconstructed L-shaped adjacent region.

[0091] Additionally, in some embodiments, the current template 434 has a current luminance template 434A and a current chrominance template 434B, and the reference template 438 has a reference luminance template 438A and a reference chrominance template 438B. The video encoder 122 includes a luminance residual template calculator 432A and a chrominance residual template calculator 432B. The luminance residual template calculator 432A determines the residual of the luminance samples 404 of the current luminance template 434A based on the current luminance template 434A and the reference luminance template 438A. The chrominance residual template calculator 432B determines the residual of the chrominance samples 402 of the current chrominance template 434B based on the current chrominance template 434B and the reference chrominance template 438B. The residuals of the luminance samples 404 and the chrominance samples 402 of the current template 434 are processed by the filter coefficient derivation module 422 to determine the filter coefficients 424 of the residual filter 416. After determining the filter coefficients 424 of the residual filter 416 based on the current template 434 and the reference template 438, the residual filter 416 is applied to determine a first residual 412A of a first chrominance sample 402A in the current coding block 406C based on one or more residuals 414L of one or more luminance samples 404L.

[0092] In some embodiments, in the RT-CCRM mode, the residual filter 416 corresponds to one of the blue difference (Cb) chrominance samples 402Cb and the red difference (Cr) chrominance samples 402Cr. When the plurality of filter coefficients 424 of the residual filter 416 are not derived based on a plurality of color templates (e.g., the current template 434 and the reference template 438), the residual of one of the Cb chrominance samples 402Cb and the Cr chrominance samples 402Cr is set to 0. In an example, the filter coefficients 424 of the first residual filter 416A corresponding to the Cr chrominance samples 402Cr cannot be derived from the template 434 and the template 438, and the residual of the Cr chrominance samples 402Cr is set to 0. In another example, the filter coefficients 424 of the second residual filter 416B corresponding to the Cb chrominance samples 402Cb cannot be derived from the template 434 and the template 438, and the residual of the Cb chrominance samples 402Cb is set to 0.

[0093] Figure 5FIG. is a diagram of a GOP 500 according to some embodiments, the GOP 500 including a current image frame 410 and two associated reference image frames 502A and 502B. The current image frame 410 includes a current coding block 406C, and the current coding block 406C is encoded according to RT-CCRM 400 in the RT-CCRM mode. The video decoder 122( Figure 2B ) receives a video bitstream 116, the video bitstream 116 including the current image frame 410 and a first syntax element 408 for the RT-CCRM mode. Based on the first syntax element 408, the video decoder 122 determines to enable the RT-CCRM mode to generate a first residual 412A of a first chrominance sample 402A of the current coding block 406C based on one or more residuals 414L of one or more luma samples 404L (e.g., Figure 4B L0 to L6 in) corresponding to the first chrominance sample 402A in the current coding block 406C of the current image frame 410.

[0094] In some embodiments, the current coding block 406C corresponds to a current template 434, the current template 434 including a reconstructed portion of the current image frame 410 and including a block of reconstructed samples adjacent to the current coding block 406C. In some embodiments, the current template 434 is selected from a top current template 434T located at the top of the current coding block 406C, a left current template 434L located to the left of the current coding block 406C, and an L-shaped current template (not shown) sharing the top and left edges of the current coding block 406C. In some embodiments, the current coding block 406C is predicted based on two reference coding blocks 436-1 and 436-2, the reference coding blocks 436-1 and 436-2 being located on two reference image frames 502A and 502B in the same GOP as the current image frame 410, respectively. Each of the reference coding blocks 436-1 or 436-2 also has a corresponding reference template 438-1 or 438-2, the reference templates 438-1 or 438-2 having the same shape and the same size as the current template 434. Template matching is a decoder-side block vector derivation method for finding two closest matches between the current template 434 of the current coding block 406C and the reference templates 438-1 and 438-2 in two different reconstructed image frames 502A and 502B, respectively.

[0095] In some embodiments, the video decoder 122 determines that the current encoded block 406C is predicted based on a merge mode (e.g., an inter-frame merge mode, an affine merge mode, an intra block copy (IBC) mode). In some embodiments, the merge mode is specified such that motion parameters (e.g., motion vectors MV1 and MV2) of the current encoded block 406C are obtained from adjacent encoded blocks (including spatial candidates and temporal candidates), as well as additional scheduling introduced in VVC. The merge mode can be applied to inter-frame predicted encoded blocks, not just for the skip mode. Alternatively, in some embodiments, the merge mode is implemented on each encoded block by explicitly transmitting motion parameters (e.g., motion vectors, reference picture indices for each reference picture list, reference picture list usage flags, other required information). According to the merge mode, the video decoder 122 determines, for example, the motion vector MV1 or MV2 based on the motion vectors of one or more adjacent encoded blocks of the current encoded block 406C, and further identifies a merge candidate 436 (e.g., reference encoded block 436-1 or 436-2) based on the motion vector MV1 or MV2. A reference template 438 (e.g., 438-1 or 438-2) associated with the merge candidate 436 is also identified based on the motion vector MV1 or MV2. The video decoder 122 identifies a current template 434 associated with the current encoded block 406C, and determines a plurality of filter coefficients 424 of the residual filter 416 based on the reference template 438 and the current template 434. More details regarding the determination of the filter coefficients 424 of the residual filter 416 are discussed above with reference to Figure 4A and Figure 4B more details are discussed regarding the determination of the filter coefficients 424 of the residual filter 416.

[0096] In some embodiments, the merge mode corresponds to bi - directional prediction. The merge candidates 436 include a first candidate 436 - 1 selected from a first reference list and a second candidate 436 - 2 selected from a second reference list. The reference template 438 is a weighted average of a first reference template 438 - 1 and a second reference template 438 - 2, where the first reference template 438 - 1 corresponds to the first candidate (e.g., reference coded block 436 - 1) and the second reference template 438 - 2 corresponds to the second candidate (e.g., reference coded block 436 - 2). Additionally, in some embodiments, the video decoder 122 determines at least one CU - level weight bi - directional prediction (BCW) for the first reference template 438 - 1 and the second reference template 438 - 2, and applies the at least one BCW to determine the reference template 438 as a weighted average of the first reference template 438 - 1 and the second reference template 438 - 2. In some embodiments, the BCW is an enhanced version of the bi - directional prediction blend in HEVC, and performs a weighted average of two prediction signals based on weights selected from a predefined set of weights. VVC also supports a geometric partition mode (GPM) that divides the coded block into non - rectangular sub - partitions, each non - rectangular sub - partition associated with a translational motion vector.

[0097] In some embodiments, when a flag for local intensity compensation (LIC) is enabled for a merge candidate 436 (e.g., reference coded block 436 - 1 or 436 - 2), the video decoder 122 applies local intensity compensation to the reference template 438 (e.g., the reference template 438 is a weighted combination of reference template 438 - 1 and reference template 438 - 2).

[0098] In some embodiments, the advanced motion vector prediction (AMVP) mode is enabled. The merge candidates 436 are located in a reference image frame 502A or 502B different from the current image frame 410. The video decoder 122 determines the motion vector of the merge candidate 436 based on the motion vector predictor (MVP) and the motion vector difference (MVD) of the merge candidate 436 (e.g., reference coded block 436 - 1 or 436 - 2). More details regarding the search for the MVD of the merge candidate 436 are discussed below with reference to Figure 6 Discuss more details about the search for the MVD of the merge candidate 436.

[0099] Alternatively, in some embodiments, when the AMVP mode is enabled, the merge candidates 436 are located on the current image frame 410. The video decoder 122 determines the block vector (BV) of the merge candidate 436 based on the block vector predictor (BVP) and the block vector difference (BVD) of the merge candidate.

[0100] In some embodiments, the current coding block 406C is predicted in the IntraTMP mode. The video decoder 122 determines merge candidates 506 based on the block vector BV, and based on the block vector BV, identifies a reference template 438 associated with the merge candidates 506. The reference template 438 includes a reference luminance template 438A and a reference chrominance template 438B. The current template 434 is associated with the current coding block 406C. Reference Figure 4A , a plurality of filter coefficients 424 of the residual filter are based on the reference template 438 and the current template 434.

[0101] Figure 6 FIG. is a diagram illustrating a template matching process 600 performed within a search range 610 around an initial motion vector 602 according to some embodiments. In some embodiments, the video decoder 122 determines to predict the current coding block 406C based on a merge mode (e.g., an inter-frame merge mode, an affine merge mode, an intra block copy (IBC) mode). According to the merge mode, the video decoder 122 determines an initial motion vector 602 based on, for example, motion vectors of one or more neighboring coding blocks of the current coding block 406C, and further based on the initial motion vector MV, identifies merge candidates 436 (e.g., including Figure 4B reference coding blocks 436 in). Also based on the initial motion vector 602, a reference template 438 (e.g., 438-1 or 438-2) associated with the merge candidates 436 is identified. The video decoder 122 identifies a current template 434 associated with the current coding block 406C, and based on the reference template 438 and the current template 434, determines a plurality of filter coefficients 424 of the residual filter 416. More details regarding the determination of the filter coefficients 424 of the residual filter 416 are discussed above with reference to Figure 4A and Figure 4B .

[0102] Template matching is a decoder-side block vector derivation method for finding the closest match between the current template 434 of the current coding block 406C and a reference template 438 that is in a reconstructed picture frame different from the current picture frame 410 or in a reconstructed portion of the current picture frame 410. A template matching cost is applied to determine whether the reference template 438 is one of the closest matches to the current template 434. In an example, the sum of absolute differences (SAD) of corresponding samples of the reference template 434 and the current template 434 is used to determine the template matching cost. More specifically, the template matching cost is monitored when searching for a motion vector around the initial motion vector 602 of the current coding block 406C within a search range 610 (e.g., [-8, +8] pixel search range). In some embodiments associated with JVET-J0021, the template matching is implemented using a search step determined based on the AMVP mode and can be cascaded with bilateral matching in the merge mode. The closest match of the reference template 438 corresponds to a candidate match 436 of the current coding block 406C (also referred to as the reference coding block 436 in Figure 4B ), and after identifying the closest match within the search range 610, a plurality of filter coefficients 424 of the residual filter 416 are determined. Thus, in some embodiments (e.g., associated with the AMVP mode), the motion vector of the merge candidate 436 can be determined based on the motion vector predictor of the merge candidate 436 (e.g., selecting the initial motion vector 602) and the motion vector difference (e.g., identifying the search result within the search range 610).

[0103] Figure 7Schematic diagram of CCRM 700 applied in cross-component residual model (CCRM) mode according to some embodiments. In some embodiments, when the first syntax signal indicates disabling the RT-CCRM mode, the second syntax element 440 for the CCRM mode is signaled to indicate whether the first chroma sample 402A is predicted based on one or more luma samples 404L. In some embodiments (e.g., associated with JVET-AD0108), when the current coding block 406C uses inter prediction or intra block copy (IBC), CCRM 700 can be applied to predict the first chroma sample 402A (e.g., including Cr chroma sample 402Cr and Cb chroma sample 402Cb) based on the reconstructed luma samples 404L. The luma predictor corresponds to the predicted luma samples 404P determined based on one or more reference coding blocks 436, and these predicted luma samples 404P are compensated by the luma residual 414L to reconstruct the luma samples 404L. The filter coefficients of the CCRM filter 702 are derived by the filter coefficient module 704 based on the predicted luma samples 404P and the predicted chroma samples 402B and 402R. The predicted luma samples 404P are provided by the luma predictor, and the predicted chroma samples 402B and 402R are provided by the Cb predictor and the Cr predictor. These predicted chroma samples include the predicted Cb chroma sample 402B and the predicted Cr chroma sample 402R. Based on the derived filter coefficients, the CCRM filter 702 is applied to predict the chroma samples 402BP and 404RP corresponding to the first chroma sample 402A. The predicted chroma samples 402BP and 404RP are compensated by the first residuals 412Cb and 412Cr to reconstruct the chroma samples 402Cr and 404Cb corresponding to the first chroma sample 404A.

[0104] In some embodiments, an eight-tap CCRM filter 702 is applied to combine six spatial luma samples (e.g., L0 to L5), a non-linear term, and a bias term to predict the first chroma sample 402A. Refer to Figure 7 , in some embodiments, the spatial luma samples (e.g., L0 to L5) are obtained from the luma grid by selecting the six luma samples closest to the chroma position C without downsampling. The predicted chroma value 402BR or 402BP (predChromaVal) is determined according to the following:

[0105]

[0106] where, c 0 to c 7Are filter coefficients corresponding to six spatial luminance samples L0 to L5, a non-linear term nonlinear((L0 + L3 + 1) >> 1), and a bias term B. In some embodiments, a division-free Gaussian elimination method associated with an enhanced compression model (ECM) can be used to derive the filter coefficients. In some embodiments, one or more offsets are applied to the luminance samples before deriving the filter. In some embodiments, when the current coding block 406C has less than 64 chrominance samples, intra reference samples are used as additional input samples used in filter derivation. Design using at most six rows and at most six columns of intra reference samples of the CCCM. Coding blocks having 256 chrominance samples or more are divided into sub-blocks having at most 256 chrominance samples. Sub-blocks containing zero luminance residuals are skipped.

[0107] In some embodiments, when the CCRM 700 is disabled, a first syntax element 408 is signaled for the current coding block 406C. Further, in some embodiments, when the current coding block 406C is coded in one of an inter prediction mode, an IBC mode, and an IntraTMP mode and when the CCRM 700 is disabled, the first syntax element 408 is signaled. In some embodiments, when the first syntax element 408 indicates that the RT-CCRM mode is disabled, a second syntax element 440 for the CCRM mode is signaled to indicate whether to predict the first chrominance sample 402A ( Figure 4B ) according to the reconstructed luminance samples corresponding to one or more luminance samples 404L.

[0108] Figure 8 Is a flowchart illustrating a method 800 for decoding video according to some embodiments. The method 800 can be performed at a computing system (e.g., the server system 112, the source device 102, or the electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, the method 800 is performed by executing instructions stored in the memory of the computing system (e.g., the memory 314). In some embodiments, the method 800 is applied in conjunction with one or more video codecs including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, when an encoded block is coded by inter prediction or intra block copy (IBC), the CCRM 700 can be used ( Figure 7) Predict the chrominance sample 402 according to the reconstructed luminance sample 404, wherein, based on two reference coded blocks 436 located in the same image frame, determine the prediction sample of the image frame coded block. When the reconstructed luminance sample 404 is used to predict the chrominance sample 402, the CCRM 700 acts as an inter-frame CCCM and is performed in a different manner from the residual cross-component filtering, which in the residual domain uses the residual data from the luminance sample 404 to predict the residual data of the chrominance sample 402. By these means, the residual cross-component filtering improves the local illumination compensation.

[0109] In some embodiments, apply CU-level weighted bi-directional prediction (BCW) to combine two prediction samples from two reference coded blocks 436-1 and 436-2. In some embodiments, Intra-frame template matching prediction (IntraTMP) is a special intra-frame prediction mode that copies the best prediction block 506 from the reconstructed part of the current image frame 410, and the template of the best prediction block 506 (e.g., having an L shape) matches the current template 434. For a predefined search range, the encoder searches the reconstructed part of the current image frame 410 for the template most similar to the current template 434 and uses the corresponding block 506 as the prediction block. Then, the encoder signals to use this mode, and the same prediction operation is performed on the decoder side.

[0110] For inter-frame prediction modes including but not limited to inter-frame prediction, IBC, IntraTMP, use the residual data of the current template 434 and the reference template from the luminance component and the chrominance component to derive the filter coefficients 424 of the residual filter 416. Based on these filter coefficients 424, apply the residual filter 416 to process the luminance reference block residual data 414L to predict the current chrominance block residual data (e.g., the first residual 412A). Refer to Figure 4A , based on the chrominance sample 402P from inter-frame prediction, IBC, or IntraTMP, and the above-predicted current chrominance block residual data (e.g., the first residual 412A), derive the compensated chrominance sample 402AC.

[0111] In some embodiments associated with the bitstream signaling, signal a flag (e.g., in the Figure 4B first syntax element 408) to indicate whether to apply RT-CCRM when the coded block 406C is coded in the inter-frame mode, IBC mode, or IntraTMP mode. In some embodiments, when the coded block 406C is coded in the inter-frame mode, IBC mode, or IntraTMP mode, and when the CCRM flag indicating whether to apply the CCRM 700 is false, signal a flag (e.g., in the Figure 4Bin the first syntax element 408 in), to indicate whether to apply RT-CCRM. In some embodiments, when the encoded block 406C is encoded in an inter mode, an IBC mode, or an IntraTMP mode, a flag is signaled (e.g., in Figure 4B in the first syntax element 408 in), to indicate whether to apply RT-CCRM. If the flag is false, when CCRM 700 is available for the encoded block 406C, a CCRM flag is signaled to indicate whether to use CCRM 700 ( Figure 7 ). In some embodiments, two flags 408A and 408B are signaled ( Figure 4B ) to indicate whether to apply RT-CCRM to the Cb chrominance samples and the Cr chrominance samples respectively. In some embodiments, the flag (e.g., in Figure 4B in the first syntax element 408 in) is inferred to be false based on the template availability or the position of the encoded block 406C in the current image frame 410, the associated sub-picture, the associated tile, or the associated image strip.

[0112] In some embodiments associated with template matching, the block vector or motion vector of the merge candidate 436 (e.g., Figure 4B the reference encoded block 436 in) is merged for pointing to the reference template 438 for both the luminance samples 404 and the chrominance samples 402 in the merge mode. The merge mode here includes but is not limited to inter-frame merge, affine merge, and IBC merge. In some embodiments, if the merge candidate 436 is bi-predictive, the reference template 438 is the weighted average of the first reference template 438-1 ( Figure 5 ) from the first reference list and the second reference template 438-2 ( Figure 5 ) from the second reference list. In some embodiments, when the LIC flag of the merge candidate 436 is true, LIC can be applied to the reference template 438. In some embodiments, when the merge candidate 436 does not use equal weights, the BCW weights can be applied to the weighted average of the first reference template 438-1 and the second reference template 438-2. In some embodiments, the block vector of intraTMP is used to point to the reference template 438 for the luminance component and the chrominance component.

[0113] In some embodiments, the Advanced Motion Vector Prediction (AMVP) mode is enabled. Based on the Motion Vector Predictor (MVP) and Motion Vector Difference (MVD) of merge candidate 436, the motion vector of merge candidate 436 is determined. Alternatively, in some embodiments, when the AMVP mode is enabled, merge candidate 436 is located on the current picture frame 410. Based on the Block Vector Predictor (BVP) and Block Vector Difference (BVD) of merge candidate 436, the block vector (BV) of merge candidate 436 is determined. This motion vector or block vector is used to point to the reference template 438 for the luminance component and chrominance component in the AMVP mode.

[0114] In some embodiments, the current template 434 or the reference template 438 is an adjacent reconstructed region above or to the left of the current coding block 406C, or an L-shaped adjacent reconstructed region. The current template 434 is also used for inter prediction, intraTMP or CCLM, MMLM, CCCM, GLM filter derivation template matching. In some embodiments, the number of rows of the current template 434 located above the current coding block and the number of columns of the current template 434 located to the left of the current coding block are greater than or equal to 1. Additionally, in the example, the number of rows of the current template 434 is not equal to the number of columns. In some embodiments, multiple template types are selected to derive the filter coefficients 424 ( Figure 4A ) of the residual filter 416. For example, based on three template types (e.g., the adjacent reconstructed region above the current coding block 406C, the adjacent reconstructed region to the left of the current coding block 406C, and the L-shaped adjacent reconstructed region), these filter coefficients are derived. Note that the reference template 438 has the same size and the same shape as the current template 434.

[0115] In some embodiments associated with filter coefficient derivation, any cross-component filter coefficient derivation method, including but not limited to CCLM, MMLM, CCCM, GLM, can be used to derive one or more filter coefficients. In some embodiments, two sets of RT-CCRM filter coefficients 424 ( Figure 4A ) are applied to generate the residuals 412A for the Cb chrominance samples 402Cb and the Cr chrominance samples 402Cr, respectively.

[0116] In some embodiments associated with RT-CCRM filtering, filter coefficients 424 cannot be derived for one of the Cb component and the Cr component when the RT-CCRM flag (e.g., in Figure 4BWhen it is true in the first syntax element 408 in), for the chrominance component for which the filter coefficients 424 cannot be derived, the filter output of the residual filter 416 (e.g., the first residual 412A) is zero. In addition, in some embodiments, for the Cb color component for which the filter coefficients 424 of the Cb component cannot be derived, the filter output (e.g., the first residual 412A) is zero. In some embodiments, for the Cr color component for which the filter coefficients 424 of the Cr component cannot be derived, the filter output (e.g., the first residual 412A) is zero. In some embodiments, the minimum value and / or the maximum value will be applied to limit the filter output (e.g., the first residual 412A of the first chrominance sample 402A) within the desired dynamic range.

[0117] In various embodiments of the present application, the luminance color component is replaced by the first color component, and the chrominance color component is replaced by the second color component. For example, the video bitstream includes the current image frame and the first syntax element 408 for the RT-CCRM mode. Based on this first syntax element 408, it is determined to enable the RT-CCRM mode to generate the first residual of the first sample of the second color component in the current coding block based on one or more residuals of one or more samples of the first color component corresponding to the first sample of the second color component in the current coding block of the current image frame. When the RT-CCRM mode is enabled, the first sample of the second color component and one or more samples of the first color component are identified. One or more residuals of the first color component are identified in the current coding block and processed by the residual filter corresponding to the RT-CCRM mode to generate the first residual of the first sample of the second color component. The predicted sample of the second color component is compensated with the first residual to reconstruct the first sample of the second color component, and the first sample of the second color component is applied to reconstruct the current image frame 410.

[0118] Although Figure 8 Although multiple logical stages are illustrated in a specific order, the stages that do not depend on the order can be reordered, and other stages can be combined or split. A certain reordering or other grouping not specifically mentioned is obvious to those of ordinary skill in the art, so the ordering and grouping presented herein are not exhaustive. In addition, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.

[0119] Now turning to some example embodiments.

[0120] (A1)In some embodiments, method 800 is implemented for decoding video data. Method 800 includes receiving (operation 802) a video bitstream including a current picture frame and a first syntax element for a residual template cross-component residual model (RT-CCRM) mode; based on the first syntax element, determining (operation 804) to enable the RT-CCRM mode to generate a first residual of a first chrominance sample of a current coding block based on one or more residuals of one or more luma samples corresponding to the first chrominance sample in the current coding block of the current picture frame; when the RT-CCRM mode is enabled (operation 806): in the current coding block, identifying (operation 808) the first chrominance sample and one or more luma samples corresponding to the first chrominance sample; determining (operation 810) one or more residuals of one or more luma samples in the current coding block; and applying (operation 812) a residual filter corresponding to the RT-CCRM mode to generate a first residual of the first chrominance sample based on the one or more residuals of the one or more luma samples; and reconstructing the first chrominance sample by compensating a predicted chrominance sample with at least the first residual to reconstruct (operation 814) the current picture frame.

[0121] (A2)In some embodiments of A1, the first syntax element for the RT-CCRM mode includes a first flag and a second flag, the first flag indicating whether to enable the RT-CCRM mode to generate a blue chrominance residual of a blue difference (Cb) chrominance sample, the second flag indicating whether to enable the RT-CCRM mode to generate a red chrominance residual of a red difference (Cr) chrominance sample, and the first chrominance sample including a Cb chrominance sample and a Cr chrominance sample.

[0122] (A3)In some embodiments of A1 or A2, method 800 further includes: determining to predict the current coding block based on a merge mode; determining a merge candidate based on a vector according to the merge mode; identifying a reference template associated with the merge candidate based on the vector; identifying a current template associated with the current coding block; and determining a plurality of filter coefficients of a residual filter based on the reference template and the current template.

[0123] (A4)In some embodiments of A3, the merge mode is one of an inter-frame merge mode, an affine merge mode, and an intra-block copy (IBC) mode.

[0124] (A5)In some embodiments of A3 or A4, the merge mode corresponds to bidirectional prediction; the merge candidate includes a first candidate selected from a first reference list and a second candidate selected from a second reference list. The reference template is a weighted average of a first reference template and a second reference template, the first reference template corresponding to the first candidate and the second reference template corresponding to the second candidate.

[0125] (A6) In some embodiments of A5, method 800 includes determining at least one CU-level weighted bi-prediction (BCW) for a first reference template and a second reference template; and applying the at least one BCW to determine the reference template as a weighted average of the first reference template and the second reference template.

[0126] (A7) In some embodiments of any one of A3 to A6, method 800 further includes applying local intensity compensation to the reference template when a flag for local intensity compensation (LIC) is enabled for a merge candidate.

[0127] (A8) In some embodiments of any one of A3 to A7, method 800 further includes, in advanced motion vector prediction (AMVP) mode, implementing one of the following: determining a vector of a merge candidate based on a motion vector predictor (MVP) and a motion vector difference (MVD) of the merge candidate; determining a vector of the merge candidate based on a block vector predictor (BVP) and a block vector difference (BVD) of the merge candidate.

[0128] (A9) In some embodiments of any one of A3 to A8, the current template is selected from: a reconstructed top adjacent region, a reconstructed left adjacent region, a reconstructed L-shaped adjacent region.

[0129] (A10) In some embodiments of any one of A3 to A9, the current template includes a filter template that is applied to determine a first filter corresponding to one of an inter prediction mode, an intraTMP mode, a multi-model linear model (MMLM) mode, a cross-component linear model (CCLM) mode, a convolutional cross-component intra prediction model (CCCM), and a gradient linear model (GLM) mode.

[0130] (A11) In some embodiments of any one of A3 to A10, in the reference template and the current template, each template includes a corresponding top template and a corresponding left template, the corresponding top template having at least one row, and the corresponding left template having at least one column.

[0131] (A12) In some embodiments of any one of A3 to A11, the reference template and the current template correspond to one of a plurality of predefined template types.

[0132] (A13) In some embodiments of A12, the plurality of predefined template types at least includes: a first type of reconstructed top adjacent region, a second type of reconstructed left adjacent region, and a third type of reconstructed L-shaped adjacent region.

[0133] In some embodiments of any one of A1 to A13, method 800 further includes: determining to predict a current coding block in an Intra Temporal Motion Vector Prediction (IntraTMP) mode; determining merge candidates based on a block vector; identifying, based on the block vector, a reference template associated with the merge candidates, the reference template including a reference luminance template and a reference chrominance template; identifying a current template associated with the current coding block; and determining a plurality of filter coefficients of a residual filter based on the reference template and the current template.

[0134] (A15) In some embodiments of any one of A1 to A14, method 800 further includes: determining to predict the current coding block using an alternative filter corresponding to one of a Multi-Model Linear Model (MMLM) mode, a Cross-Component Linear Model (CCLM) mode, a Convolutional Cross-Component Intra Prediction Model (CCCM), and a Gradient Linear Model (GLM) mode; determining that the alternative filter has a plurality of first filter coefficients; and determining a plurality of filtering coefficients of the residual filter based on the plurality of first filter coefficients of the alternative filter.

[0135] (A16) In some embodiments of any one of A1 to A14, the residual filter includes a first residual filter for a first chrominance sample among Chrominance Blue (Cb) chrominance samples and Chrominance Red (Cr) chrominance samples. Method 800 further includes: determining to predict the current coding block using a first alternative filter and a second alternative filter corresponding to two different ones of the MMLM mode, the CCLM mode, the CCCM, and the GLM mode; determining that the first alternative filter has a plurality of first filter coefficients and that the second alternative filter has a plurality of second filter coefficients; determining a plurality of filtering coefficients of the first residual filter based on the plurality of first filter coefficients of the first alternative filter; and determining a plurality of filtering coefficients of a second residual filter based on the plurality of second filter coefficients of the second alternative filter, wherein the second residual filter is applied to generate a residual of a second different chrominance sample among the Cb chrominance samples and the Cr chrominance samples.

[0136] (A17) In some embodiments of any one of A1 to A16, the residual filter corresponds to one of Chrominance Blue (Cb) chrominance samples and Chrominance Red (Cr) chrominance samples. Method 800 further includes: setting the residual of one of the Cb chrominance samples and the Cr chrominance samples to 0 when the plurality of filter coefficients of the residual filter are not derived based on a plurality of color templates.

[0137] (A18) In some embodiments of any one of A1 to A17, method 800 further includes: clipping the residual of the first chrominance sample within a dynamic range defined between a first residual value and a second residual value, the second residual value being greater than the first residual value.

[0138] (A19)In some embodiments of any one of A1 to A18, the video bitstream further includes a second residual of the first chrominance samples, and reconstructing the current image frame includes: compensating the predicted chrominance samples with both the first residual and the second residual to reconstruct the first chrominance samples.

[0139] (A20)In some embodiments, a method includes: receiving video data including a current image frame; encoding the current image frame including a current coding block; determining whether to enable a Residual Template Cross-Component Residual Model (RT-CCRM) to generate a residual of a first chrominance sample of the current coding block based on one or more residuals of one or more luma samples in the current coding block of the current image frame; transmitting the encoded current image frame via the video bitstream; and signaling a first syntax element via the video bitstream to indicate whether to enable the RT-CCRM mode to generate a residual of a first chrominance sample of the current coding block based on one or more residuals of one or more luma samples in the current coding block of the current image frame.

[0140] (A21)In some embodiments of A20, the method further includes: determining to encode the current coding block in one of an inter prediction mode, an Intra Block Copy (IBC) mode, and an Intra Template Matching Prediction (IntraTMP) mode; wherein, when encoding the current coding block in one of the inter prediction mode, the IBC mode, and the IntraTMP mode, signaling the first syntax element.

[0141] (A22)In some embodiments of A21, the method further includes: determining to disable the Cross-Component Residual Model (CCRM) mode to abort predicting the first chrominance sample based on the reconstructed luma samples corresponding to one or more luma samples. When encoding the current coding block in one of the inter prediction mode, the IBC mode, and the IntraTMP mode and when CCRM is disabled, signaling the first syntax element.

[0142] (A23)In some embodiments of A21, the method further includes: when the first syntax signal indicates disabling the RT-CCRM mode, signaling a second syntax element for the CCRM mode, the second syntax element indicating whether to predict the first chrominance sample based on the reconstructed luma samples corresponding to the one or more luma samples.

[0143] (A24)In some embodiments of any one of A20 to A23, the method further includes resetting the first syntax element to a value indicating disabling the RT-CCRM mode based on template availability and the position of the coding block in a picture, sub-picture, tile, or stripe of the current image frame.

[0144] (A25) In some embodiments, a method includes: obtaining a source video sequence including a current coded block of a current picture frame, and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes the current picture frame and a first syntax element for a residual template cross-component residual model (RT-CCRM) mode, the first syntax element indicating whether to generate a residual of a first chroma sample of the current coded block based on one or more residuals of one or more luma samples in the current coded block of the current picture frame. According to determining that the first syntax element indicates enabling the RT-CCRM mode, applying a residual filter corresponding to the RT-CCRM mode to generate a residual of the first chroma sample based on the one or more residuals of the one or more luma samples.

[0145] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A25 above).

[0146] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A25 above).

[0147] Unless otherwise specified, any of the syntax elements described herein may be high-level syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to the sequence level, picture level, slice level, or tile level. As another example, HLS elements may be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, picture header, tile header, and / or CTU header.

[0148] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should further be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0149] As used herein, depending on the context, the term "if" can be interpreted to mean "when the stated precondition is true" or "once the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "in response to detecting that the stated precondition is true". Similarly, depending on the context, the phrase "if it is determined that [the stated precondition is true]" or "if [the stated precondition is true]" or "when [the stated precondition is true]" can be interpreted to mean "upon determining that the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "upon detecting that the stated precondition is true" or "in response to detecting that the stated precondition is true".

[0150] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and practical applications, thereby enabling others skilled in the art to implement.

Claims

1. A method for decoding video data, characterized in that: include: Receiving a video bitstream including a current image frame and a first syntax element, wherein the first syntax element is used in a residual template cross-component residual model (RT-CCRM) mode; Based on the first syntax element, determine to enable the RT-CCRM mode to generate a first residual of a first chroma sample of a current coding block of the current image frame based on one or more residuals of one or more luma samples corresponding to a first chroma sample in the current coding block of the current image frame; When the RT-CCRM mode is enabled: In the current coding block, identifying the first chroma sample and the one or more luma samples corresponding to the first chroma sample; Determining one or more residuals of the one or more luma samples in the current coding block; as well as applying a residual filter corresponding to the RT-CCRM mode to generate a first residual of the first chroma sample based on the one or more residuals of the one or more luma samples; as well as The first chroma samples are reconstructed by compensating the predicted chroma samples with at least the first residual to reconstruct the current image frame.

2. The method according to claim 1, characterized in that The first syntax element for the RT-CCRM mode includes a first flag and a second flag, the first flag indicating whether the RT-CCRM mode is enabled to generate a blue chroma residual of a blue difference (Cb) chroma sample, and the second flag indicating whether the RT-CCRM mode is enabled to generate a red chroma residual of a red difference (Cr) chroma sample, the first chroma sample including a Cb chroma sample and a Cr chroma sample.

3. The method according to claim 1, characterized in that Further including: Determine, based on a merge mode, to predict the current coding block; According to the merging mode, based on the vector, a merging candidate is determined; Based on the vector, identifying a reference template associated with the merge candidate; identifying a current template associated with the current coding block; as well as Based on the reference template and the current template, a plurality of filter coefficients of the residual filter are determined.

4. The method according to claim 3, characterized in that The merge mode is one of an inter-frame merge mode, an affine merge mode, and an intra-frame block copy (IBC) mode.

5. The method according to claim 3, characterized in that: The merge mode corresponds to bidirectional prediction; The merge candidates include a first candidate selected from the first reference list and a second candidate selected from the second reference list; The reference template is a weighted average of a first reference template and a second reference template, the first reference template corresponds to the first candidate, and the second reference template corresponds to the second candidate.

6. The method according to claim 3, characterized in that Further including: When a flag for local intensity compensation (LIC) is enabled for the merge candidate, local intensity compensation is applied to the reference template.

7. The method according to claim 3, characterized in that Further including: In Advanced Motion Vector Prediction (AMVP) mode, implement one of the following: Determining a vector of the merge candidate based on a motion vector predictor (MVP) and a motion vector difference (MVD) of the merge candidate; A vector of the merge candidate is determined based on a block vector predictor (BVP) and a block vector difference (BVD) of the merge candidate.

8. The method according to claim 3, characterized in that The current template is selected from: a reconstructed top adjacent region, a reconstructed left adjacent region, and a reconstructed L-shaped adjacent region.

9. The method according to claim 3, characterized in that: The current template includes a filter template, which is used to determine a first filter, which corresponds to one of the inter-frame prediction mode, intra-frame template matching prediction (IntraTMP) mode, multi-model linear model (MMLM) mode, cross-component linear model (CCLM) mode, convolution cross-component intra-frame prediction model (CCCM), and gradient linear model (GLM) mode.

10. The method according to claim 3, characterized in that In the reference template and the current template, each template includes a corresponding top template and a corresponding left template, the corresponding top template has at least one row, and the corresponding left template has at least one column.

11. The method according to claim 1, characterized in that: Further including: Determine to predict the current coding block in an intra-frame template matching prediction mode; Based on the block vector, determining a merge candidate; Based on the block vector, identifying a reference template associated with the merge candidate, the reference template comprising a reference luma template and a reference chroma template; identifying a current template associated with the current coding block; as well as Based on the reference template and the current template, a plurality of filter coefficients of the residual filter are determined.

12. The method according to claim 1, characterized in that Further including: Determining to use a substitute filter to predict the current coding block, the substitute filter corresponds to one of a multi-model linear model (MMLM) mode, a cross-component linear model (CCLM) mode, a convolutional cross-component intra prediction model (CCCM), and a gradient linear model (GLM) mode; determining that the replacement filter has a plurality of first filter coefficients; as well as A plurality of filter coefficients of the residual filter are determined based on the plurality of first filter coefficients of the substitute filter.

13. The method according to claim 1, characterized in that The residual filter includes a first residual filter for a first chroma sample of a blue difference (Cb) chroma sample and a red difference (Cr) chroma sample, and the method further includes: Determine to use a first alternative filter and a second alternative filter to predict the current coding block, wherein the first alternative filter and the second alternative filter correspond to two different modes among an MMLM mode, a CCLM mode, a CCCM mode, and a GLM mode; determining that the first alternative filter has a plurality of first filter coefficients and the second alternative filter has a plurality of second filter coefficients; determining a plurality of filter coefficients of the first residual filter based on the plurality of first filter coefficients of the first substitute filter; and A plurality of filter coefficients of a second residual filter are determined based on the plurality of second filter coefficients of the second replacement filter, wherein the second residual filter is applied to generate a residual of a different second chroma sample among the Cb chroma sample and the Cr chroma sample.

14. The method according to claim 1, characterized in that Further including: The residual of the first chroma sample is clipped within a dynamic range defined between a first residual value and a second residual value, wherein the second residual value is greater than the first residual value.

15. The method according to claim 1, characterized in that The video code stream further includes a second residual of the first chroma sample, and reconstructing the current image frame includes: compensating the predicted chroma sample with both the first residual and the second residual to reconstruct the first chroma sample.

16. A computing system, characterized in that: include: Control circuit system; as well as a memory storing one or more programs configured to be executed by the control circuit system, the one or more programs further comprising instructions for: Receiving video data including a current image frame; Encoding the current image frame including the current coding block; Determining whether to enable a residual template cross-component residual model (RT-CCRM) to generate a first residual of a first chrominance sample of the current coding block based on one or more residuals of one or more luma samples in the current coding block of the current image frame; Transmitting the encoded current image frame via a video code stream; as well as A first syntax element is signaled via the video code stream to indicate whether an RT-CCRM mode is enabled to generate a first residual of the first chrominance sample of the current coding block based on one or more residuals of the one or more luma samples in the current coding block of the current image frame.

17. The computer system according to claim 16, characterized in that: The one or more programs further include instructions for: Determine to encode the current coding block in one of an inter-frame prediction mode, an intra-frame block copy (IBC) mode, and an intra-frame template matching prediction (IntraTMP) mode; Wherein, when the current coding block is encoded in one of the inter-frame prediction mode, the IBC mode and the IntraTMP mode, the first syntax element is signaled.

18. The computer system according to claim 17, characterized in that: The one or more programs further include instructions for: determining to disable a cross-component residual model (CCRM) mode to discontinue predicting the first chroma samples from reconstructed luma samples corresponding to the one or more luma samples; Wherein, when the current coding block is encoded in one of the inter-frame prediction mode, the IBC mode, and the IntraTMP mode, and when the CCRM is disabled, the first syntax element is signaled.

19. The computer system according to claim 17, wherein: The one or more programs further include instructions for: When the first syntax signal indicates that the RT-CCRM mode is disabled, a second syntax element for the CCRM mode is signaled, the second syntax element indicating whether the first chroma samples are predicted from reconstructed luma samples corresponding to the one or more luma samples.

20. A non-volatile computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores one or more programs for execution by control circuitry of a computing system, the one or more programs including instructions for: Obtaining a source video sequence including a current coding block of a current image frame; as well as Perform conversion between the source video sequence and a video code stream, wherein the video code stream includes: the current image frame; and a first syntax element for a residual template cross-component residual model (RT-CCRM) mode, the first syntax element indicating whether to generate a first residual of a first chrominance sample of a current coding block of the current image frame based on one or more residuals of one or more luma samples in the current coding block of the current image frame; Wherein, according to determining that the first syntax element indicates that the RT-CCRM mode is enabled, a residual filter corresponding to the RT-CCRM mode is applied to generate a first residual of the first chrominance sample based on one or more residuals of the one or more luma samples.