Block adaptive multi-hypothesis cross-component prediction

Through the block adaptive multi-assumption cross-component prediction method, chroma samples are generated using multiple luminance samples, which solves the problem of insufficient accuracy and compression efficiency of cross-component intra prediction in the prior art, and achieves more efficient video encoding.

CN120188481APending Publication Date: 2025-06-20TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078585.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2023-10-31
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are difficult to effectively utilize multiple luminance samples in cross-component intra prediction, resulting in a lack of accuracy and compression efficiency of chromaticity sample prediction.

Method used

The block adaptive multi-assumption cross-component prediction method is adopted, and the hypothesis tap index is represented by the signal, one of the multiple hypothesis tap combinations is selected, and the weighting factor is determined based on the reference area of ​​the corresponding coded block, and a linear or nonlinear weighted sum is generated as a chroma sample.

Benefits of technology

The cross-component intra prediction accuracy and compression efficiency of video data are improved, and the encoder's adaptability to different encoding blocks is enhanced, reducing the code rate without damaging the video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120188481A_ABST
    Figure CN120188481A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems of video coding. In one aspect, a video bitstream includes a current coded block of an image frame and signals a cross-component intra prediction mode and a hypothetical tap index. The computing system identifies a first luma sample and a first chroma sample in the current coding block co-located with the first luma sample, and selects one of a plurality of hypothetical tap combinations based on a hypothetical tap index. The computing system identifies neighboring luminance samples of the first luminance sample based on the selected hypothesis tap combination and generates a hypothesis value based on the identified neighboring luminance samples of the first luminance sample. The computing system further generates a first chroma sample based at least on the first luma sample and the hypothesis value, and reconstructs a current coded block comprising the first chroma sample.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 535,472, filed Aug. 30, 2023, entitled “Block Adaptive Multi-Hypothesis Cross Component Prediction,” and this application is a continuation of, and claims priority to, U.S. Patent Application No. 18 / 497,915, filed Oct. 30, 2023, entitled “Block Adaptive Multi-Hypothesis Cross Component Prediction.” Technical Field

[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for cross-component intra prediction of video data. Background Art

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise convey digital video data over a communication network and / or store the digital video data on a storage device. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding may be used to compress video data according to one or more video coding standards before the video data is conveyed or stored.

[0004] Multiple video coding and decoding standards have been developed. For example, video coding standards include Alliance for Open Media (AOMedia) Video 1 (AV1), Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Experts Group (MPEG) coding. Video coding typically utilizes prediction methods (e.g., inter prediction, intra prediction, etc.), such prediction methods exploiting the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality.

[0005] HEVC (also known as H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC, also known as H.266) is a video compression standard designed to be a successor to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AV1 is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the validation version 1.0.0 (with errata 1) of the specification was released. Summary of the Invention

[0006] As mentioned above, encoding (compression) reduces the bandwidth and / or storage space requirements. As described in detail later, lossless compression and lossy compression can be employed. Lossless compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal through a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not completely retained during encoding and not completely restored during decoding. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications can tolerate higher distortion compared to users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows for coding algorithms that produce higher losses and higher compression ratios.

[0007] This disclosure describes applying multiple parameters to achieve cross-component intra prediction of video data in the cross-component intra prediction (CCIP) mode, where, in the CCIP mode, each chroma sample among multiple chroma samples of a current coding block is determined based on one or more luma samples. For example, in multi-hypothesis cross-component prediction (MH-CCP), a linear or non-linear weighted sum of multiple versions of luma samples is used to predict the chroma sample. The multiple versions of luma samples include the luma sample C co-located with the chroma sample, and based on adjacent luma samples (e.g., Figure 4AFiltered luminance samples determined by the W, N, E, S, NW, NE, SW, and SE in [the reference region] and applied as filter inputs. Each filter input of the weighted sum is referred to as a hypothesis. In MH-CCP, each hypothesis is associated with a weighting factor. In one aspect of the present application, the hypothesis tap index is signaled in the video bitstream passed from the encoder to the decoder, and the hypothesis tap index is applied to select at least one hypothesis tap combination from a plurality of hypothesis tap combinations for the current coding block. In another aspect of the present application, for each coding block, the weighting factor applied to generate the linear or non-linear weighted sum is determined based on the reference region of the corresponding coding block. In some embodiments, these weighting factors are determined by applying a least mean square calculation kernel to process the reconstructed samples of the reference blocks of each coding block.

[0008] According to some embodiments, a video decoding method is provided. The method includes: receiving a video bitstream including a current coding block of a current image frame. The video bitstream includes a video bitstream that signals: (i) a first syntax element for a cross-component intra prediction (CCIP) mode, where the CCIP mode indicates that each chrominance sample of the current coding block is determined based on one or more luminance samples; and (ii) a second syntax element for a hypothesis tap index that selects at least one hypothesis tap combination from a plurality of hypothesis tap combinations for the current coding block. The method further includes: selecting one hypothesis tap combination from a plurality of hypothesis tap combinations based on the hypothesis tap index. The method further includes: generating a plurality of hypothesis values based on a plurality of neighboring luminance samples of the first luminance sample based on the selected one hypothesis tap combination from the plurality of hypothesis tap combinations; and generating a first chrominance sample co-located with the first luminance sample based on at least the first luminance sample and the plurality of hypothesis values. The method further includes: reconstructing the current coding block including the first chrominance sample.

[0009] According to some embodiments, a video decoding method is provided. The method includes: receiving a video bitstream including a current coding block of a current image frame. The video bitstream includes: (i) a first syntax element for a cross-component intra prediction (CCIP) mode that indicates whether each chrominance sample of the current coding block is determined based on one or more luminance samples; and (ii) a second syntax element for a hypothesis tap index that selects at least one hypothesis tap combination from a plurality of hypothesis tap combinations for the current coding block. The method includes: selecting one hypothesis tap combination from a plurality of hypothesis tap combinations based on the hypothesis tap index; and generating a plurality of hypothesis values based on a plurality of neighboring luminance samples of the first luminance sample based on the selected one hypothesis tap combination from the plurality of hypothesis tap combinations. The method includes: generating a first chrominance sample co-located with the first luminance sample based on at least the first luminance sample and the plurality of hypothesis values; and reconstructing the current coding block including the first chrominance sample.

[0010] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0011] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets executable by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0012] Thus, devices and systems using methods for encoding and decoding video are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for video encoding and decoding.

[0013] The features and advantages described in the specification do not necessarily include all, and in particular, considering the accompanying drawings, the specification, and the claims provided in the present disclosure, some additional features and advantages will be obvious to those of ordinary skill in the art. In addition, it should be noted that the language used in the specification is mainly selected for readability and guiding purposes, and not necessarily for depicting or defining the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To be able to understand the present disclosure in more detail, a more specific description may be made with reference to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand when reading the present disclosure that the specification may include other effective features.

[0015] Figure 1 is a block diagram showing an exemplary communication system according to some embodiments.

[0016] Figure 2A is a block diagram showing exemplary elements of an encoder component according to some embodiments.

[0017] Figure 2B is a block diagram showing exemplary elements of a decoder component according to some embodiments.

[0018] Figure 3 is a block diagram showing an exemplary server system according to some embodiments.

[0019] Figure 4A shows an exemplary scheme for generating a first chrominance sample based on a plurality of luminance samples according to some embodiments, Figure 4Band Figure 4C is a schematic diagram of two hypothesized tap combinations of four adjacent luminance samples including a first luminance sample according to some embodiments.

[0020] Figures 5A to 5D is a schematic diagram of four exemplary hypothesized tap combinations according to some embodiments, wherein each of the four exemplary hypothesized tap combinations is generated based on two adjacent luminance samples directly adjacent to the first luminance sample.

[0021] Figures 6A to 6D is a schematic diagram of a set of four exemplary hypothesized tap combinations according to some embodiments, wherein a target hypothesized tap combination is selected from each set of the four exemplary hypothesized tap combinations based on a hypothesized tap index.

[0022] Figure 7A is a diagram showing an exemplary current image frame according to some embodiments, wherein, in the exemplary current image frame, a first coding block has a reference region.

[0023] Figure 7B is a diagram showing two exemplary coding blocks and associated reference regions according to some embodiments.

[0024] Figure 8 is a flowchart showing an exemplary method for encoding video according to some embodiments.

[0025] Figure 9 is a flowchart showing another exemplary method for encoding video according to some embodiments.

[0026] By convention, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and drawings to represent the same features. Detailed Description

[0027] The present disclosure describes cross-component intra prediction of video data in a cross-component intra prediction (CCIP) mode, wherein, in the CCIP mode, each chrominance sample among a plurality of chrominance samples of a current coding block is determined based on one or more associated luma samples. For example, the CCIP mode includes a multi-hypothesis cross-component prediction (MH-CCP) mode, in which multiple versions of luma samples are combined to generate a linear or non-linear weighted sum as the chrominance sample. The multiple versions of luma samples include the luma sample C co-located with the chrominance sample and filtered luma samples, which are also referred to as hypotheses and are equal to a weighted combination of two or more adjacent luma samples of the luma sample C. The luma sample C and the multiple hypothesis values are combined based on multiple weighting factors to generate a chrominance sample co-located with the luma sample C. In one aspect of the present application, a hypothesis tap index is signaled in a video bitstream transmitted from an encoder to a decoder, and the hypothesis tap index is applied to select at least one hypothesis tap combination from multiple hypothesis tap combinations for a current coding block. In another aspect of the present application, for each coding block, the weighting factors applied to generate a linear or non-linear weighted sum are determined based on a reference region of the corresponding coding block.

[0028] In some embodiments, the multiple weighting factors are applied in combination with one or two additional weighting factors to combine the luma sample C and the hypothesis values with non-linear terms and bias terms. For example, a cross-shaped 5-tap filter has five inputs, which include a central (C) luma sample co-located with the chrominance sample to be predicted and four hypothesis values. For example, each hypothesis value includes a combination of two or more of the upper / north (N) adjacent sample, lower / south (S) adjacent sample, left / west (W) adjacent sample, and right / east (E) adjacent sample. The non-linear term P represents the square of the central luma sample C scaled to the sample value range. The bias term B represents a scalar offset between the input and the output, for example, set to the middle chrominance value (set to 512 for 10-bit content). In some embodiments, the output of the filter is determined as the convolution of the weighting factors ci (also referred to as filter coefficients ci) with the input luma sample C and the hypothesis values, and is clipped to the range of valid chrominance samples.

[0029] Figure 1 is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), and the source device 102 and the plurality of electronic devices 120 are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, which is used with video-enabled applications such as video conferencing applications, digital television applications, media storage, and / or distribution applications.

[0030] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 may be of high data volume compared to the encoded video bitstreams generated by the encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video data to the network 110).

[0031] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0032] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content, such as the encoded video stream from the source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination of hardware and software. In some embodiments, the encoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108.

[0033] In some embodiments, the server system 112 serves as a media-aware network element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0034] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0035] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smartphone, a tablet, or a laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0036] In an exemplary operation of the communication system 100, the source device 102 transmits the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may use the encoder component 114 to decode and / or encode the encoded video stream 108. For example, the server system 112 may apply encoding to the video data, which is more optimal for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to recover and optionally display the video pictures.

[0037] In some embodiments, the transmission discussed above is a one-way data transmission. One-way data transmission is sometimes used in applications such as media services. In some embodiments, the transmission discussed above is a two-way data transmission. Two-way data transmission is sometimes used in applications such as video conferencing. In some embodiments, the encoded video stream 108 and / or the encoded video data 116 are encoded and / or decoded according to any video coding / compression standard described herein (e.g., HEVC, VVC, and / or AV1).

[0038] Figure 2Ais a block diagram showing exemplary elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb, or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0, or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously acquired / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be organized into a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0039] The encoder component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can easily identify other functions of the controller 204, as these functions may relate to the encoder component 106 optimized for a particular system design.

[0040] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to the way a (remote) decoder creates sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. Thus, the prediction part of the encoder interprets the reference picture samples as the same sample values as the decoder will interpret when using the prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is known to those of ordinary skill in the art.

[0041] The operation of the decoder 210 can be the same as that of the remote decoder of the decoder component 122 described in detail below. However, briefly referring to Figure 2B Since the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into the encoded video sequence losslessly, the entropy decoding part of the decoder component 122 including the buffer memory 252 and the parser 254 may not be fully implemented in the local decoder 210. Figure 2B

[0042] At this point, it can be observed that any decoder technology present in the decoder (other than parsing / entropy decoding) must also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are needed and provided only in certain areas.

[0043] As part of its operation, the source encoder 202 may perform motion compensated predictive coding, which predictively encodes an input frame by referring to one or more previously encoded frames designated as reference image frames in the video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference image frame, which can be selected as the prediction reference for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.

[0044] The decoder 210 decodes the encoded video data of a frame that can be designated as a reference picture frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown), the reconstructed video sequence can be a replica of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on the reference picture frame and can cause the reconstructed reference picture frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference picture frame that has the same content (in the absence of transmission errors) as the reconstructed reference picture frame that will be obtained by the remote video decoder.

[0045] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as a candidate reference pixel block) or some metadata, such as a reference picture motion vector, block shape, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel block basis of sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.

[0046] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.

[0047] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission over a communication channel 218, which may be a hardware / software link to a storage device that can store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / signal noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, and the like.

[0048] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are familiar with these variations of I pictures and their corresponding applications and characteristics, and thus will not be described in detail herein. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0049] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded with reference to a previously encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures either through spatial prediction or through temporal prediction.

[0050] The video captured can be multiple source pictures (video pictures) in a time series. Intra-picture prediction (commonly abbreviated as intra-frame prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded with a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0051] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard such as those described herein. In operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.

[0052] Figure 2B is a block diagram showing exemplary elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0053] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link leading to a storage device storing the encoded video data. The receiver may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0054] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as one integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 is implemented at least partially in software.

[0055] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 located inside decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (e.g., to prevent network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be needed, or buffer memory 252 may be smaller. For use on a service packet network such as the Internet, buffer memory 252 may be needed, buffer memory 252 may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not depicted) external to decoder component 122.

[0056] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include, for example, information for managing the operation of the decoder components 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard and may follow various principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0057] Depending on the type of the encoded video picture or a portion of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the way they are involved may be controlled by the parser 254 through subgroup control information parsed from the encoded video sequence. For clarity, such subgroup control information flows between the parser 254 and the multiple units below are not depicted.

[0058] In addition to the functional blocks already mentioned, the decoder component 122 may conceptually be subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is maintained as conceptually subdivided into the multiple functional units below.

[0059] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as symbols 270 and control information (e.g., which transform to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 may output a block including sample values, and the sample values may be input into the aggregator 268.

[0060] In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may generate a block having the same size and shape as the block being reconstructed using the surrounding reconstructed information extracted from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 may add the predictive information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a per-sample basis.

[0061] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After motion compensating the extracted samples according to the symbol 270 belonging to the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (in this case, referred to as residual samples or a residual signal), thereby generating output sample information. The extraction of the prediction samples by the motion compensation prediction unit 260 from an address within the reference picture memory 266 may be controlled by a motion vector. The motion vector may be provided to the motion compensation prediction unit 260 in the form of the symbol 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of the sample values extracted from the reference picture memory 266 when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0062] The output samples of the aggregator 268 may be subjected to various loop filtering techniques in the loop filter unit 256. The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and available to the loop filter unit 256 as the symbol 270 from the parser 254, and the video compression technique may also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0063] The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device (e.g., the display 124) and stored in the reference picture memory 266 for future inter-picture prediction.

[0064] Once fully reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture has been fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent coded pictures.

[0065] The decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard (e.g., any standard described herein). The coded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the coded video sequence follows the syntax of the video compression technique or standard (as specified in the video compression technique document or standard, particularly as specified in the profile of the video compression technique or standard). Additionally, to conform to some video compression techniques or standards, the complexity of the coded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0066] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., a CPU (central processing unit), a GPU (graphics processing unit), and / or a DPU (data processing units)). In some embodiments, the control circuitry includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0067] The network interface 304 can be configured to connect to one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network can be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Such communications can be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CANBus connected to certain CANBus devices), or two-way (e.g., connecting to other computer systems using a local area network or wide area digital network). Such communications can include communications to one or more cloud computing networks.

[0068] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of a keyboard, a mouse, a touchpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output device 308 can include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a display or a monitor), etc.

[0069] The memory 314 can include high-speed random access memory (e.g., DRAM (dynamic random access memory), SRAM (static random access memory), DDR RAM (double data rate random access memory), and / or other solid-state random access memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disc storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Optionally, the memory 314 includes one or more storage devices arranged remotely from the control circuit 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; ● A network communication module 318, which is used to connect the server system 112 to other computing devices through one or more network interfaces 304 (e.g., through wired and / or wireless connections); ● An encoding / decoding module 320, which is used to perform various functions related to encoding and / or decoding data (e.g., video data). In some embodiments, the encoding / decoding module 320 is an instance of the encoder component 114. The encoding / decoding module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322, which is used to perform various functions related to decoding encoded data, such as the functions previously described for the decoder component 122; and ○ An encoding module 340, which is used to perform various functions related to encoding data, such as the functions previously described for the encoder component 106; and ● A picture memory 352, which is used to store pictures and picture data, e.g., for use with the encoding / decoding module 320. In some embodiments, the picture memory 352 includes one or more of the reference picture memory 208, the buffer memory 252, the current picture memory 264, and the reference picture memory 266.

[0070] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described for the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described for the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described for the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described for the loop filter 256).

[0071] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described for the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described for the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0072] Each of the above-identified modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, optionally, the codec module 320 does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores subsets of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0073] In some embodiments, the server system 112 includes a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and web pages and applications that are implemented using Common Gateway Interface (CGI) scripts, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), Hypertext Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), and the like.

[0074] Although Figure 3 FIG. shows a server system 112 according to some embodiments, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems rather than a structural schematic of the embodiments described herein. In practice, as will be appreciated by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112, and how the features are distributed among the servers, will vary depending on the implementation and, optionally, will depend in part on the volume of data traffic processed by the server system during peak usage periods as well as during average usage periods.

[0075] Figure 4A FIG. shows an exemplary scheme 400 for generating a first chrominance sample 402C in the MH-CCP mode based on a plurality of luminance samples 404C and 404X, Figure 4B and Figure 4CSchematic diagrams of two exemplary hypothesis tap combinations 420A and 420B that each include four adjacent luminance samples 404X including the first luminance sample 404C, according to some embodiments. In some embodiments, the current encoded block 406A of the current image frame 408 is encoded in a cross-component intra prediction (CCIP) mode. In the CCIP mode, the decoder 122 ( Figure 2B ) determines each chrominance sample 402 of the current encoded block 406A based on one or more reconstructed luminance samples 404. In some cases, the CCIP mode includes a cross-component linear model mode (CCLM), in which, in the CCLM, the first chrominance sample 402C is transformed based on a linear model according to the reconstructed luminance sample 404C co-located with the chrominance sample. Alternatively, in some cases, the CCIP mode includes a convolutional cross-component mode (CCCM), in which, in the CCCM, the first chrominance sample 402C is predicted directly based on a plurality of reconstructed luminance samples 404X located adjacent to the first luminance sample 404C according to the filter shape of the filter. Alternatively and additionally, in some cases, the CCIP mode includes an MH-CCP mode, in which, in the MH-CCP mode, the first chrominance sample 402C is generated by combining at least the first luminance sample 404C co-located with the first chrominance sample 402C with a plurality of hypothesis values 410 using a plurality of weighting factors (e.g., c0 to c6). A plurality of adjacent luminance samples 404X of the first luminance sample 404C are combined using a plurality of coefficients to generate a plurality of hypothesis values 410.

[0076] In some embodiments, the plurality of adjacent luminance samples 404X include a north adjacent luminance sample (also referred to as an upper luminance sample) 400N, a south adjacent luminance sample (also referred to as a lower luminance sample) 400S, a west adjacent luminance sample (also referred to as a left luminance sample) 400W, and an east adjacent luminance sample (also referred to as a right luminance sample) 400E. Additionally, referring Figure 4B to, in some embodiments, the north adjacent luminance sample 404N and the south adjacent luminance sample 404S are combined to generate a first subset of one or more hypothesis values 410A and 410C. The west adjacent luminance sample 404W and the east adjacent luminance sample 404E are combined to generate a second subset of one or more hypothesis values 410B and 410D. For example, the north adjacent luminance sample 404N and the south adjacent luminance sample 404S are combined to generate a first hypothesis value 410A(a) and a third hypothesis value 410C(c), and the west adjacent luminance sample 404W and the east adjacent luminance sample 404E are combined to generate a second hypothesis value 410B(b) and a fourth hypothesis value 410D(d). Specifically, the four hypothesis values 410 are represented as follows: a = w1 * N + w1’ * S (1) b = w2 * W + w2’ * E (2) c = w3 * N + w3’ * S (3) d = w4 * W + w4’ * E (4) where a, b, c, d, and e are assumed values 410A, 410B, 410C, and 410D, respectively; N, W, S, and E are the luminance values of adjacent luminance samples 404N, 404W, 404S, and 404E, respectively; and w1, w1’, w2, w2’, w3, w3’, w4, and w4’ are coefficients used to combine adjacent luminance samples 404X to generate the assumed values 410. In one example, w1 and w1’ are equal to 1; w2 and w2’ are equal to 1; w3 and w3’ are equal to 1 and -1, respectively; and w4 and w4’ are equal to 1 and -1, respectively. In some embodiments, each of equations (1) to (4) is normalized. The sum of the absolute values of w1 and w1’, the sum of the absolute values of w2 and w2’, the sum of the absolute values of w3 and w3’, and the sum of the absolute values of w4 and w4’ is equal to 1.

[0077] Based on the determination of applying the MH-CCP mode, the first chroma sample 402 is predicted according to one of the following equations: predChromaVal = c0e + c1a + c2b + c3c + c4d (5.1) predChromaVal = c0e + c1a + c2b + c3c + c4d + c5P (5.2) predChromaVal = c0e + c1a + c2b + c3c + c4d + c6B (5.3) predChromaVal = c0e + c1a + c2b + c3c + c4d + c5P + c6B (5.4) where predChromaVal is the predicted chroma value of the first chroma sample 402C; e is the luma value of the first luma sample 404C co-located with the first chroma sample 402C; P is a non-linear term, e.g., equal to (C*C + median luma) >> bitdepth; B is an offset; and c0 to c6 are weighting factors. In some embodiments (e.g., in Equation (5.1)), the non-linear term P and the offset B are not applied to predict the first chroma sample 402C. Instead, in some embodiments (e.g., in Equation (5.2) or (5.3)), only one of the non-linear term P and the offset B is applied to predict the first chroma sample 402C. Instead, in some embodiments (e.g., in Equation (5.4)), both the non-linear term P and the offset B are applied to predict the first chroma sample 402C. In some embodiments, B is the median luma value or the average luma value of the luma samples 404 of the current coding block 406A.

[0078] In some embodiments, the plurality of weighting factors c0 to c6 are determined based on a set of one or more reference luma samples 404R and a set of one or more co-located reference chroma samples 402R within a reference region 412 of the current coding block 406A. The reference region 412 is located within the current image frame 408. Further, in some embodiments, the reference luma samples 404R of the reference region 412 are used to generate corresponding reference hypothesis values based on Equations (1) to (4), and these reference hypothesis values are further combined to re-generate one or more chroma samples based on any of Equations (5.1) to (5.4). In some embodiments, a set of one or more co-located reference chroma samples 402R is compared with a set of one or more re-generated chroma samples to generate a least mean square (LMS) value. The plurality of weighting factors c0 to c6 are iteratively adjusted to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., in the predetermined criterion, the LMS value is less than an LMS threshold, or the LMS value is minimized).

[0079] In some embodiments, the current image frame 408 further includes a second coding block 406B different from the current coding block 406A (also referred to as the first coding block). For each of the coding blocks 406A and 406B, a reference region of the corresponding coding block is identified. The reference region 412 of the second coding block 406B is different from the reference region 412 of the first coding block 406A. In some cases, the size of the reference region 412 of the current coding block 406A is different from the size of the reference region 412 of the second coding block 406B.

[0080] In some embodiments, multiple adjacent luminance samples 404X include a northwest adjacent luminance sample (also referred to as an upper left luminance sample) 400NW, a southeast adjacent luminance sample (also referred to as a lower right luminance sample) 400SE, a southwest adjacent luminance sample (also referred to as a lower left luminance sample) 400SW, and a northeast adjacent luminance sample (also referred to as an upper right luminance sample) 400NE. Additionally, referring to Figure 4C , in some embodiments, the northwest adjacent luminance sample 404NW and the southeast adjacent luminance sample 404SE are combined to generate a first subset of one or more hypothesized values. The southwest adjacent luminance sample 404SW and the northeast adjacent luminance sample 404NE are combined to generate a second subset of one or more hypothesized values. For example, the northwest adjacent luminance sample 404NW and the southeast adjacent luminance sample 404SE are combined to generate a first hypothesized value 410A(a) and a third hypothesized value 410C(c), and the southwest adjacent luminance sample 404SW and the northeast adjacent luminance sample 404NE are combined to generate a second hypothesized value 410B(b) and a fourth hypothesized value 410D(d). Specifically, the four hypothesized values 410 (a to d) are represented as follows: a = w1*NW + w1’*SE (6) b = w2*SW + w2’*NE (7) c = w3*NW + w3’*SE (8) d = w4*SW + w4’*NE (9) where a, b, c, d, and e are the four hypothesized values 410A, 410B, 410C, and 410D respectively; NW, SW, SE, and NE are the luminance values of the adjacent luminance samples 404NW, 404SW, 404SE, and 404NE respectively; w1, w1’, w2, w2’, w3, w3’, w4, and w4’ are coefficients for combining the adjacent luminance samples 404X to generate the hypothesized values 410. In one example, w1 and w1’ are equal to 1; w2 and w2’ are equal to 1; w3 and w3’ are equal to 1 and -1 respectively; w4 and w4’ are equal to 1 and -1 respectively. In some embodiments, equations (6) to (9) are normalized. The sum of the absolute values of w1 and w1’, the sum of the absolute values of w2 and w2’, the sum of the absolute values of w3 and w3’, and the sum of the absolute values of w4 and w4’ are equal to 1. According to the determination of the applied MH-CCP mode, the first chrominance sample 402C is predicted according to any one of equations (5.1) to (5.4).

[0081] In some embodiments, the plurality of weighting factors c0 to c6 are determined based on a set of one or more luminance samples 404 and a set of one or more co-located chrominance samples 402 within a reference region 412 of the current coding block 406A. The reference region 412 is located within the current image frame 408. Further, in some embodiments, the set of one or more luminance samples 404 of the reference region 412 are used to generate corresponding reference hypothesis values based on equations (1) to (4) or equations (6) to (9). The corresponding reference hypothesis values are further combined to generate one or more reference chrominance samples based on any of equations (5.1) to (5.4). The set of one or more co-located chrominance samples 402 are compared with the one or more reference chrominance samples to generate an LMS value. The plurality of weighting factors c0 to c6 are iteratively adjusted to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., in the predetermined criterion, the LMS value is less than an LMS threshold, or the LMS value is minimized).

[0082] In some embodiments, the first luminance sample 404C is a downsampled luminance sample using a downsampling filter (when luminance and chrominance have different dimensions (e.g., 4:2:2 or 4:2:0)), and the same applies to each adjacent sample (e.g., N, W, E, S, NW, NE, SW, SE) used to derive the corresponding hypothesis value 410. Alternatively, in some embodiments, the first luminance sample 404C is an original luminance sample that is co-located with the first chrominance sample 402C and not downsampled at all. Each adjacent sample (e.g., N, W, E, S, NW, NE, SW, SE) used to derive the corresponding hypothesis value 410 includes an original luminance sample that is adjacent to the co-located luminance sample and not downsampled at all.

[0083] In some embodiments, at least one of the weighting factors c0 to c6 is derived based on chrominance samples and luminance samples within a reference region 412 of the current coding block 406A, and the reference region 412 includes one or more coding blocks decoded before the current coding block 406A (e.g., Figure 4A 8 coding blocks in). In some embodiments, a subset of the one or more coding blocks is directly adjacent to the current coding block 406A. In some embodiments, a subset of the one or more coding blocks is separated from the current coding block 406A by one or more coding blocks. In some embodiments, the reference region 412 includes at least a part of one or more rows above the current coding block 406A and / or a part of one or more columns to the left of the current coding block 406A. For example, reference Figure 4A , the reference region 412 includes 7 rows of luminance samples 404 above the current coding block 406A and 9 columns of luminance samples 404 to the left of the current coding block 406A.

[0084] In some embodiments, at least one of the weighting factors c0 to c6 is determined by minimizing the mean square error (MSE) between the predicted chrominance samples 402 in the reference region 412 and the reconstructed chrominance samples 402. The MSE minimization is performed by calculating the autocorrelation matrix of the luminance samples 404 and the cross-correlation vector between the luminance samples 404R of the reference region 412 and the chrominance samples 402R. The autocorrelation matrix is processed by LDL decomposition, and back substitution is used to calculate multiple weighting factors. This process generally follows the calculation of the filter coefficients of the adaptive loop filter (ALF) in the enhanced compression model (ECM) video coding. The LDL decomposition does not use square root operations but only integer arithmetic operations.

[0085] Reference Figure 4B and Figure 4C , each of the two exemplary hypothesis tap combinations 420A and 420B includes four hypothesis values 410A to 410D. The video bitstream 116 ( Figure 1 ) received by the decoder 122 includes a hypothesis tap index 414 that selects at least one hypothesis tap combination 420 (e.g., including Figure 4A and Figure 4B two exemplary combinations 420A and 420B therein) for the current encoded block 406A. In one example, the hypothesis tap index 414 selects the hypothesis tap combination 420A, and a plurality of adjacent luminance samples 404X (e.g., 404N, 404W, 404S, and 404E) are identified based on the selected hypothesis tap combination 420A. The hypothesis values 410A to 410D are generated based on the identified adjacent luminance samples 404X (e.g., 404N, 404W, 404S, and 404E). In another example, the hypothesis tap index 414 selects the hypothesis tap combination 420B, and a plurality of adjacent luminance samples 404X (e.g., 404NW, 404SW, 404SE, and 404NE) are identified based on the selected hypothesis tap combination 420B. The hypothesis values 410A to 410D are generated based on the identified adjacent luminance samples 404X (e.g., 404NW, 404SW, 404SE, and 404NE). In some embodiments, in the video bitstream 116, the hypothesis tap index 414 is signaled at one of the block level, superblock level, picture frame level, key frame level, and picture sequence level of the current encoded block 406A.

[0086] In some embodiments, the plurality of hypothesized tap combinations 420 do not include one or both of combinations 420A and 420B. In some embodiments, the plurality of hypothesized tap combinations 420 include one or more additional hypothesized tap combinations different from combinations 420A and 420B. Further, in some embodiments, each of combinations 420A or 420B includes four hypothesized taps. In contrast, the additional hypothesized tap combinations include the same number (i.e., 4) or a different number (e.g., 2, 8) of hypothesized values. More details regarding hypothesized values, taps, and associated combinations are explained below with reference to Figures 5A to 5D and Figures 6A to 6D explain more details regarding hypothesized values, taps, and associated combinations.

[0087] Figures 5A to 5D FIG. Figures 5A to 5D is a schematic diagram of four exemplary hypothesized tap combinations according to some embodiments, wherein each of the four exemplary hypothesized tap combinations is generated based on two adjacent luminance samples that are directly adjacent to a first luminance sample. Figure 5A FIG. Figures 6A to 6D is a schematic diagram of a hypothesized tap combination 500 including a north adjacent luminance sample 404N and a south adjacent luminance sample 404S according to some embodiments; Figure 5B FIG. is a schematic diagram of another hypothesized tap combination 520 including a west adjacent luminance sample 404W and an east adjacent luminance sample 404E according to some embodiments; Figure 5C FIG.

[0087] is a schematic diagram of another hypothesized tap combination 540 including a northwest adjacent luminance sample 404NW and a southeast adjacent luminance sample 404SE according to some embodiments; Figure 5D FIG. Figures 5A to 5D is a schematic diagram of another hypothesized tap combination 560 including a southwest adjacent luminance sample 404SW and a northeast adjacent luminance sample 404NE according to some embodiments. The current coded block 406A of the current image frame 408 includes a first chrominance sample 402C, a first luminance sample 404C co-located with the first chrominance sample 402C, and a plurality of adjacent luminance samples 404X (e.g., 404N, 404S, 404W, 404S, 404NW, 404NE, 404SW, and 404SE) of the first luminance sample 404C. The plurality of adjacent luminance samples 404X of the first luminance sample 404C are combined using a plurality of coefficients to generate a plurality of hypothesized values 410, and the plurality of hypothesized values 410 are further combined to generate the first chrominance sample 402C co-located with the first luminance sample 404C. In some embodiments (e.g., in Figure 4A ), the plurality of adjacent luminance samples 404X include four adjacent luminance samples 404X, and the four adjacent luminance samples 404X are combined to generate four hypothesized values 410, and the four hypothesized values 410 are applied to generate the first chrominance sample 402C. Referring to Figures 5A to 5D , each hypothesized tap combination includes two hypothesized values 410, and the two hypothesized values 410 are generated based on two adjacent luminance samples 404X.

[0088] Specifically, in some embodiments, a plurality of adjacent luminance samples 404X includes a first adjacent luminance sample (e.g., 404N) and a second adjacent luminance sample (e.g., 404S), and a first position of the first adjacent luminance sample (e.g., 404N) and a second position of the second adjacent luminance sample (e.g., 404S) are symmetric with respect to a position of the first luminance sample 404C. Further, in some embodiments, the first adjacent luminance sample (e.g., 404N) and the second adjacent luminance sample (e.g., 404S) are combined to generate a first hypothesized value (e.g., 410A in Figure 5A and a second hypothesized value (e.g., 410C in Figure 5A ). Further, in some embodiments, a first coefficient (e.g., w1) and a second coefficient (e.g., w1’) are used to combine the first adjacent luminance sample (e.g., 404N) and the second adjacent luminance sample (e.g., 404S) in a weighted manner to generate a first hypothesized value (e.g., 410A in Figure 5A ). A third coefficient (e.g., w3) and a fourth coefficient (e.g., w3’) are used to combine the first adjacent luminance sample (e.g., 404N) and the second adjacent luminance sample (e.g., 404S) in a weighted manner to generate a second hypothesized value (e.g., 410C in Figure 5A ). The first coefficient is equal to the third coefficient, and the second coefficient is opposite to the fourth coefficient. Further, in some embodiments, the first coefficient and the second coefficient are normalized (e.g., the sum of the associated magnitudes of the first coefficient and the second coefficient is 1), and the third coefficient and the fourth coefficient are normalized (e.g., the sum of the associated magnitudes of the third coefficient and the fourth coefficient is 1).

[0089] Referring to Figure 5A , in some embodiments, the first adjacent luminance sample and the second adjacent luminance sample include a north adjacent luminance sample 404N and a south adjacent luminance sample 404S. The luminance samples 404N and 404S are used to generate two hypothesized values 410A and 410C (i.e., a and c), and further, the two hypothesized values 410A and 410C are combined with the first luminance sample 404C in a weighted manner to generate a first chroma sample 402C according to one of the following equations: predChromaVal = c0e + c1a + c3c (10.1) predChromaVal = c0e + c1a + c3c + c5P (10.2) predChromaVal = c0e + c1a + c3c + c6B (10.3) predChromaVal = c0e + c1a + c3c + c5P + c6B(10.4) Thus, in some embodiments, the hypothesized tap index 414 (Figure 4A ) Select the hypothesized tap combination 500. The selected hypothesized tap combination 500 corresponds to the first luminance sample 404C, the north luminance sample 404N directly above the first luminance sample 404C, and the south luminance sample 404S directly below the first luminance sample 404C.

[0090] Reference Figure 5B , in some embodiments, the west adjacent luminance sample 404W and the east adjacent luminance sample 404E are used to generate two hypothesized values 410B and 410D (i.e., b and d), and further these two hypothesized values 410B and 410D are combined with the first luminance sample 404C in a weighted manner to generate the first chrominance sample 402C according to one of the following equations: predChromaVal = c0e + c2b + c4d (11.1) predChromaVal = c0e + c2b + c4d + c5P (11.2) predChromaVal = c0e + c2b + c4d + c6B (11.3) predChromaVal = c0e + c2b + c4d + c5P + c6B(11.4) Thus, in some embodiments, the hypothesized tap index 414 ( Figure 4A ) Select the hypothesized tap combination 520. The selected hypothesized tap combination corresponds to the first luminance sample 404C, the west luminance sample 404W directly to the left of the first luminance sample 404C, and the east luminance sample 404E directly to the right of the first luminance sample 404C.

[0091] Reference Figure 5C , in some embodiments, the northwest adjacent luminance sample 404NW and the southeast adjacent luminance sample 404SE are used to generate two hypothesized values 410A and 410C, and further these two hypothesized values 410A and 410C are combined with the first luminance sample 404C in a weighted manner to generate the first chrominance sample 402C according to any of the equations (10.1) to (10.4). Reference Figure 5D , in some embodiments, the southwest adjacent luminance sample 404SW and the northeast adjacent luminance sample 404NE are used to generate two hypothesized values 410B and 410D, and further these two hypothesized values 410B and 410D are combined with the first luminance sample 404C in a weighted manner to generate the first chrominance sample 402C according to any of the equations (11.1) to (11.4).

[0092] In some embodiments, based on a plurality of weighting factors (e.g., c0 to c6), the first luminance sample 404C and a plurality of hypothesis values 410 are combined with at least one of (1) a non - linear term P of the first luminance sample 404C and a subset of a plurality of adjacent luminance samples 404X and (2) a bias term B. Further, in some embodiments, the subset of the first luminance sample 404C and the plurality of adjacent luminance samples 404X includes only the first luminance sample 404C. The non - linear term P is determined based on the first luminance sample 404C. In one example, the non - linear term P is equal to the square of the luminance value of the first luminance sample 404C. Further, in some embodiments, the bias term B is determined based on at least one of (i) the median of the set of luminance samples 404 of the current coding block 406A and (ii) the average of the set of luminance samples 404 of the current coding block 406A. Optionally, the set of luminance samples 404 includes all the luminance samples that have been reconstructed for the current coding block 406A. Optionally, the set of luminance samples 404 includes fewer luminance samples than all the luminance samples that have been reconstructed for the current coding block 406A.

[0093] In some embodiments, the north - adjacent luminance sample 404N is located directly above the first luminance sample 404C, and the south - adjacent luminance sample 404S is located directly below the first luminance sample 404C. In some embodiments, the west - adjacent luminance sample 404W is located directly to the left of the first luminance sample 404C, and the east - adjacent luminance sample 404E is located directly to the right of the first luminance sample 404C. In some embodiments, the pixel block corresponding to the northwest - adjacent luminance sample 404NW is connected to the upper - left corner of the pixel block corresponding to the first luminance sample 404C, and the pixel block corresponding to the southeast - adjacent luminance sample 404SE is connected to the lower - right corner of the pixel block corresponding to the first luminance sample 404C. In some embodiments, the pixel block corresponding to the southwest - adjacent luminance sample 404SW is connected to the lower - left corner of the pixel block corresponding to the first luminance sample, and the pixel block corresponding to the northeast - adjacent luminance sample 404NE is connected to the upper - right corner of the pixel block corresponding to the first luminance sample 404C.

[0094] Figures 6A to 6D is a schematic diagram of four exemplary hypothesis - tap combination sets 600, 620, 640, and 660 according to some embodiments, wherein a target hypothesis - tap combination is selected from each of the four exemplary hypothesis - tap combination sets 600, 620, 640, and 660 based on a hypothesis - tap index 414. The video bitstream 116 received by the decoder 122 ( Figure 1) includes a hypothesized tap index 414 that selects, for at least the current coding block 406A, one of a plurality of hypothesized tap combinations 420. In other words, the plurality of hypothesized tap combinations 420 form a set of hypothesized tap combinations, and a target hypothesized tap combination is selected from this set of hypothesized tap combinations based on the hypothesized tap index 414. The target hypothesized tap combination is used to generate a plurality of hypothesized values 410 based on neighboring luma samples 404X of the first luma sample 404C, and to generate a first chroma sample 402C co-located with the first luma sample 404 based on the first luma sample 404C and the hypothesized values 410. In some embodiments, in the video bitstream 116, the hypothesized tap index 414 is signaled at one of the block level, superblock level, picture frame level, key frame level, and picture sequence level of the current coding block 406A. When the decoder 122 receives the hypothesized tap index 414, the set of hypothesized tap combinations is stored in a memory associated with and known to the decoder 122. In some embodiments, the set of hypothesized tap combinations is passed for a GOP. Alternatively, in some embodiments, the set of hypothesized tap combinations is predefined for the decoder 122 and the encoder 106.

[0095] Reference Figure 6A , in some embodiments, the first luma sample 404C is surrounded by eight nearest luma sample candidates 404N, 404W, 404S, 404E, 404NW, 404NE, 404SW, and 404SE. The plurality of hypothesized tap combinations 420 includes a first hypothesized tap combination 602 and a second hypothesized tap combination 604. The first hypothesized tap combination 602 corresponds to the first luma sample 404C and two neighboring luma samples 404X selected from the eight nearest luma sample candidates. The second hypothesized tap combination 604 corresponds to the first luma sample 404C and four neighboring luma samples 404X selected from the eight nearest luma sample candidates. For example, the first hypothesized tap combination 602 includes one of hypothesized tap combinations 500, 520, 540, and 560. Two hypothesized values 410 are generated and combined with the first luma sample 404C in a weighted manner to generate the first chroma sample 402C according to any of equations (10.1) to (10.4) and equations (11.1) to (11.4). The second hypothesized tap combination 604 includes one of hypothesized tap combinations 420A and 420B. Four hypothesized values 410 are generated and combined with the first luma sample 404C in a weighted manner to generate the first chroma sample 402C according to any of equations (5.1) to (5.4).

[0096] Reference Figure 6B , in some embodiments, the plurality of hypothesized tap combinations 420 includes a first hypothesized tap combination 420A( Figure 4B) and the second hypothesized tap combination 420B( Figure 4C ), the first hypothesized tap combination 420A corresponds to the first luminance sample 404C, the north luminance sample 404N, the south luminance sample 404S, the west luminance sample 404W, and the east luminance sample 404S, and the second hypothesized tap combination 420B corresponds to the first luminance sample 404C, the northwest luminance sample 404NW, the northeast luminance sample 404NE, the southwest luminance sample 404SW, and the southeast luminance sample 404SE. Both the hypothesized tap combinations 420A and 420B include four adjacent luminance samples 404X located at different positions.

[0097] Reference Figure 6C , a plurality of hypothesized tap combinations 420 includes at least four hypothesized tap combinations 500, 520, 540, and 560. The first hypothesized tap combination 500 corresponds to the first luminance sample 404C, the north luminance sample 404N directly above the first luminance sample 404C, and the south luminance sample 404S directly below the first luminance sample 404C. The second hypothesized tap combination 520 corresponds to the first luminance sample 404C, the west luminance sample 404W directly to the left of the first luminance sample 404C, and the east luminance sample 404E directly to the right of the first luminance sample 404C. The third hypothesized tap combination 540 corresponds to the first luminance sample 404C, the northwest (NW) luminance sample 404NW, and the southeast (SE) luminance sample 404SE. The positions of the northwest (NW) luminance sample 404NW and the southeast (SE) luminance sample 404SE are symmetric with respect to the position of the first luminance sample 404C. The fourth hypothesized tap combination 560 corresponds to the first luminance sample 404C, the southwest (SW) luminance sample 404SW, and the northeast (NE) luminance sample 404NE. The positions of the southwest (SW) luminance sample 404SW and the northeast (NE) luminance sample 404NE are symmetric with respect to the position of the first luminance sample 404C. For each of the combinations 500, 520, 540, and 560, four hypothesized values 410 are generated, and these four hypothesized values 410 are combined with the first luminance sample 404C in a weighted manner to generate the first chrominance sample 402C according to any one of equations (10.1) to (10.4) and equations (11.1) to (11.4).

[0098] Reference Figure 6D, a plurality of hypothesis tap combinations 420 includes at least two hypothesis tap combinations 500 and 606. The first hypothesis tap combination 500 corresponds to the first luminance sample 404C, the north luminance sample 404N directly above the first luminance sample 404C, and the south luminance sample 404S directly below the first luminance sample 404C. The second hypothesis tap combination 606 corresponds to the first luminance sample 404C, two or more west luminance samples (e.g., 404W, 404W1) directly to the left of the first luminance sample 404C, and two or more east luminance samples (e.g., 404E, 404E1) directly to the right of the first luminance sample 404C. One of the combinations 500 and 606 is selected to generate two or four hypothesis values 410, and the two or four hypothesis values 410 are combined with the first luminance sample 404C in a weighted manner to generate the first chrominance sample 402C.

[0099] Figure 7A is a diagram showing a current image frame 408 according to some embodiments, in which, in the current image frame 408, the first coding block 406A has a reference region 412. Figure 7B is a diagram showing two coding blocks 406A and 406B and associated reference regions 412-1 and 412-2 according to some embodiments. The reference region 312 includes reference samples 402R and 404R for determining the weighting factors (e.g., c0 to c6) of the first coding block 406A. Specifically, the decoder 122 receives a video bitstream 116 including the first coding block 406A and the second coding block 406B of the current image frame 408. The video bitstream 116 includes syntax elements for cross-component intra prediction (CCIP) mode, and the CCIP mode indicates whether each chrominance sample 402C of the first coding block 406A and the second coding block 406B is determined based on one or more luminance samples 404C and 404X. For each of the first coding block 406A and the second coding block 406B, the decoder 122 identifies the first luminance sample 404C, the first chrominance sample 402C co-located with the first luminance sample 404C in the corresponding coding block 406A or 406B, and a plurality of adjacent luminance samples 404X of the first luminance sample 404C. A plurality of hypothesis values 410 are generated based on the plurality of adjacent luminance samples 404X of the first luminance sample 402C. The decoder 122 identifies the reference region 412 of the corresponding coding block 406A or 406B, determines a plurality of weighting factors based on a set of one or more reference samples 402R and 404R in the reference region 412, and combines the first luminance sample 404C with the plurality of hypothesis values 410 based on the plurality of weighting factors to generate the first chrominance sample 402C. The decoder 122 reconstructs the current image frame 408 including the first chrominance sample 402C of each of the first coding block 406A and the second coding block 406B.

[0100] The size of the reference region 412 of the first coding block 406A is different from the size of the reference region 412 of the second coding block 406B. For example, the reference region 412 of the second coding block 406B includes the reference region 412B, while the reference region 412 of the first coding block 406A does not include the reference region 412B.

[0101] In some embodiments, for each of the first coding block 406A and the second coding block 406B, the reference region 412 includes one or more coding blocks that are adjacent to and decoded before the corresponding coding block. In some embodiments, the one or more coding blocks are directly adjacent to the current coding block 406. In some embodiments, the one or more coding blocks are separated from the current coding block 406 by one or more coding blocks.

[0102] In addition, in some embodiments, the reference region 412 of the first coding block 406A includes one or more of the following: the upper left reference region 412TL, the upper reference region 412T, the upper right reference region 412TR, the lower left reference region 412BL, and the left reference region 412L. In one example, the reference region 412 includes the upper reference region 412T and the left reference region 412L. Each reference region includes one or more coding blocks. In other words, in some embodiments, the reference region 412 includes at least a part of multiple rows above the current coding block 406 and / or a part of multiple columns to the left of the current coding block 406. For example, Figure 7A , the reference region 412 includes a first part of 7 chrominance samples above the current coding block 406 and a second part of 9 chrominance samples to the left of the current coding block 406. The first part is determined by the length of the current coding block 406, and the second part is determined by the width of the current coding block 406. In some embodiments, the reference region 412 extends one coding block width to the right of the right boundary of the current coding block 406 and one coding block height below the bottom boundary of the current coding block 406. In some embodiments, the reference region 412 is adjusted to include only available samples. The extension 412E of the reference region 412 is required to support the side samples of the cross-shaped spatial filter, and the extension 412E of the reference region 412 is filled with unavailable regions.

[0103] Reference Figure 7B, in some embodiments, a set of one or more reference samples in the reference region 412-1 of the first coding block 406A includes a first number 702 of rows (e.g., 3 rows) of reference samples, and a set of one or more reference samples in the reference region 412-2 of the second coding block 406B includes a second number 704 of rows (e.g., 1 row) of reference samples. In one example, the first number is not equal to the second number. In one example, the first number is equal to the second number. The rows of reference samples correspond to one of the rows and columns of the reference samples in the reference region 412. In some embodiments, the rows of reference samples extend to cover an entire row or column of reference samples in the current image frame 408. Alternatively, the rows of reference samples extend to cover an entire row or column of reference samples in a reference region (e.g., 412T), where the reference region (e.g., 412T) includes a portion (less than all of the reference samples in the entire row or column of reference samples) of an entire row or column of reference samples in the current image frame 408. In some embodiments, the first number 702 and the second number are signaled in the video bitstream 116. In some embodiments, the first number 702 and the second number are not signaled in the video bitstream 116.

[0104] The decoder 122 determines the first number based on a subset of the block partition, block size, block shape, and block resolution of the first coding block 406A, and determines the second number based on a subset of the block partition, block size, block shape, and block resolution of the second coding block 406B.

[0105] In some embodiments, the block resolution of the first coding block 406A is greater than a resolution threshold. The block resolution of the second coding block 406B is less than the resolution threshold. The first number 702 is greater than the second number 704. Conversely, in some embodiments, the block resolution of the first coding block 406A is greater than a resolution threshold. The block resolution of the second coding block 406B is less than the resolution threshold. The first number 702 is less than the second number 704.

[0106] In some embodiments, the block size of the first coding block 406A is greater than a block size threshold. The block size of the second coding block 406B is less than the block size threshold. The first number 702 is greater than the second number 704. Conversely, in some embodiments, the block size of the first coding block 406A is greater than a block size threshold, and the block size of the second coding block 406B is less than the block size threshold. The first number 702 is less than the second number 704.

[0107] The rows of the reference samples correspond to one of the rows and columns of the reference samples in the reference region 412. In other words, in some embodiments, each of the first quantity 702 rows of reference samples and the second quantity 704 rows of reference samples includes the reference samples of the corresponding row in the corresponding reference region 412. Alternatively, in some embodiments, each of the first quantity 702 rows of reference samples and the second quantity 704 rows of reference samples includes the reference samples of the corresponding column in the corresponding reference region 412.

[0108] In some embodiments, the video bitstream 116 includes a syntax element of a reference region index 706 for each of the first coding block 406A and the second coding block 406B. The decoder 122 determines the first quantity 702 based on the reference region index 706 of the first coding block 406A, and determines the second quantity based on the reference region index of the second coding block. Additionally, in some embodiments, a look-up table is used to determine the first quantity 702 based on the reference region index 706 of the first coding block 406A, and a look-up table is used to determine the second quantity 704 based on the reference region index 706 of the second coding block 406B. In some embodiments, for each of the first coding block 406A and the second coding block 406B, in the video bitstream, the reference region index 706 is signaled at one of the block level, super-block level, picture frame level, key frame level, and picture sequence level of the corresponding coding block.

[0109] In some embodiments, the decoder 122 determines a plurality of weighting factors as follows: determining a least mean square (LMS) value based on a set of one or more reference samples; and iteratively adjusting the plurality of weighting factors to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., the LMS value is minimized, or is less than an LMS threshold).

[0110] Figure 8is a flowchart showing an exemplary method 800 for encoding a video according to some embodiments. Method 800 may be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120), which has a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, method 800 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, method 800 is executed by executing instructions stored in the memory of the computing system (e.g., the codec module 320 of memory 314). In some embodiments, multiple combinations of unfiltered / filtered cross-component samples are supported. When decoder 120 receives a video bitstream including a current encoded block 406A of current image frame 408, a selection of which combination of unfiltered / filtered samples to use is signaled (804) in the video bitstream 116 and parsed at decoder 120. In one example, a combination 420 of 3 filtered samples is supported, with a first combination being {N, S}, a second combination being {W, E}, and a third combination being {N, S, W, E}. Signaling occurs at one or more of block level, superblock level, frame level, keyframe level, and sequence level.

[0111] In some embodiments, a selection between an M-tap linear model and an N-tap linear model is signaled in bitstream 116 ( Figure 6A ). In one example, M equals 3 and N equals 5. In some embodiments, a selection between different sets of luminance samples (or sets of filter taps) is signaled in the bitstream. In some embodiments, a selection between a first hypothesized tap combination 420A (e.g., including {N, W, S, E, C}) ( Figure 4B ) and a second hypothesized tap combination 420B (e.g., including {NW, NE, SW, SE, C}) ( Figure 4C ) is signaled in bitstream 116. In some embodiments, referring to Figure 6C , a selection among four hypothesized tap combinations 500, 520, 540, and 560 (e.g., including {N, S, C}, {W, E, C}, {NW, SE, C}, {SW, NE, C}) is signaled in the bitstream.

[0112] In some embodiments ( Figure 5A and Figure 5C ), a, c, e, P, and B are used to derive a first chrominance sample 402C according to equations (10.1) to (10.4). Signaling occurs at one or more of block level, superblock level, frame level, keyframe level, and sequence level. In some embodiments ( Figure 5B andFigure 5D ) wherein b, d, e, P, and B are used to derive the chrominance prediction value. The signaling occurs at one or more of the block level, superblock level, frame level, key frame level, and sequence level.

[0113] In some embodiments (e.g., in Figure 6D ), the filter taps / sizes can be further extended. The filter shape can also be adjusted. For example, a five-tap filter is involved in the horizontal direction to improve performance, while a three-tap filter is maintained in the vertical direction to avoid an additional line buffer. This can also be signaled in the bitstream. The signaling occurs at one or more of the block level, superblock level, frame level, key frame level, and sequence level.

[0114] Although Figure 8 a number of logical stages are shown in a particular order, the stages that are not order-dependent can be reordered, and other stages can be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.

[0115] Now, turning to some exemplary embodiments.

[0116] (A1) In some implementations, a method 800 for decoding video data is implemented. Method 800 includes: receiving (802) a video bitstream that includes a current encoded block of a current image frame, the video bitstream including (804): (i) a first syntax element for cross-component intra prediction (CCIP) mode, the CCIP mode indicating whether each chrominance sample of the current encoded block is determined based on one or more luminance samples; and (ii) a second syntax element for a hypothesized tap index that selects at least one hypothesized tap combination from a plurality of hypothesized tap combinations for the current encoded block. Method 800 further includes: selecting (808) one hypothesized tap combination from the plurality of hypothesized tap combinations based on the hypothesized tap index. Method 800 includes: generating (812) a plurality of hypothesized values based on a plurality of neighboring luminance samples of a first luminance sample based on the one hypothesized tap combination selected from the plurality of hypothesized tap combinations; generating (814) a first chrominance sample co-located with the first luminance sample based at least on the first luminance sample and the plurality of hypothesized values; and reconstructing (816) the current encoded block that includes the first chrominance sample.

[0117] (A2) In some embodiments of A1, in the video bitstream, the hypothesized tap index is signaled at one of the block level, superblock level, image frame level, key frame level, and image sequence level of the current encoded block.

[0118] (A3)In some embodiments of A1 or A2, the first luminance sample is surrounded by eight nearest luminance sample candidates, and the multiple hypothesis tap combinations include: a first hypothesis tap combination and a second hypothesis tap combination. The first hypothesis tap combination corresponds to the first luminance sample and two adjacent luminance samples selected from the eight nearest luminance sample candidates; the second hypothesis tap combination corresponds to the first luminance sample and four adjacent luminance samples selected from the eight nearest luminance sample candidates.

[0119] (A4)In some embodiments of A1 or A2, the multiple hypothesis tap combinations include: a first hypothesis tap combination and a second hypothesis tap combination. The first hypothesis tap combination corresponds to the first luminance sample, the north luminance sample, the south luminance sample, the west luminance sample, and the east luminance sample; the second hypothesis tap combination corresponds to the first luminance sample, the northwest luminance sample, the northeast luminance sample, the southwest luminance sample, and the southeast luminance sample.

[0120] (A5)In some embodiments of A1 or A2, the multiple hypothesis tap combinations include: a first hypothesis tap combination, a second hypothesis tap combination, a third hypothesis tap combination, and a fourth hypothesis tap combination. The first hypothesis tap combination corresponds to the first luminance sample, the north luminance sample directly above the first luminance sample, and the south luminance sample directly below the first luminance sample; the second hypothesis tap combination corresponds to the first luminance sample, the west luminance sample directly to the left of the first luminance sample, and the east luminance sample directly to the right of the first luminance sample; the third hypothesis tap combination corresponds to the first luminance sample, the northwest (NW) luminance sample, and the southeast (SE) luminance sample, wherein the positions of the northwest (NW) luminance sample and the southeast (SE) luminance sample are symmetric with respect to the position of the first luminance sample; the fourth hypothesis tap combination corresponds to the first luminance sample, the southwest (SW) luminance sample, and the northeast (NE) luminance sample, wherein the positions of the southwest (SW) luminance sample and the northeast (NE) luminance sample are symmetric with respect to the position of the first luminance sample.

[0121] (A6)In some embodiments of A1 or A2, according to the hypothesis tap index, one of the multiple hypothesis tap combinations corresponds to the first luminance sample, the north luminance sample directly above the first luminance sample, and the south luminance sample directly below the first luminance sample.

[0122] (A7)In some embodiments of A1 or A2, according to the hypothesis tap index, one of the multiple hypothesis tap combinations corresponds to the first luminance sample, the west luminance sample directly to the left of the first luminance sample, and the east luminance sample directly to the right of the first luminance sample.

[0123] (A8) In some embodiments of A1 or A2, the multiple hypothesized tap combinations include: a first hypothesized tap combination and a second hypothesized tap combination. The first hypothesized tap combination corresponds to a first luminance sample, a north luminance sample directly above the first luminance sample, and a south luminance sample directly below the first luminance sample. The second hypothesized tap combination corresponds to the first luminance sample, two or more west luminance samples directly to the left of the first luminance sample, and two or more east luminance samples directly to the right of the first luminance sample.

[0124] (A9) In some embodiments of any one of A1 to A8, generating a first chrominance sample based at least on the first luminance sample and the multiple hypothesized values further includes: using multiple weighting factors to combine the first luminance sample, the multiple hypothesized values with at least one of (1) a non - linear term of the first luminance sample and a subset of multiple adjacent luminance samples and (2) a bias term.

[0125] (A10) In some embodiments of A9, method 800 includes: determining multiple weighting factors based on a set of one or more luminance samples and a set of one or more co - located chrominance samples within a reference region of a current coding block, wherein the reference region is within a current image frame.

[0126] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit. The memory stores one or more instruction sets configured to be executed by the control circuit, and the one or more instruction sets include instructions for performing any method described herein (e.g., A1 to A10 above).

[0127] In yet another aspect, some embodiments include a non - transitory computer - readable storage medium storing one or more instruction sets executed by a control circuit of a computing system, and the one or more instruction sets include instructions for performing any method described herein (e.g., A1 to A10 above).

[0128] Figure 9FIG. 900 is a flowchart illustrating another exemplary method for encoding video according to some embodiments. Method 900 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and a memory storing instructions executable by the control circuitry. In some embodiments, method 900 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, method 900 is performed by executing instructions stored in a memory (e.g., codec module 320 of memory 314) of the computing system. Block-level adaptive reference samples for applying multi-hypothesis cross-component prediction (MH-CCP) are used. For different coding blocks 406, different numbers of reference rows or columns are employed in MH-CCP. For example, for the first coding block 406A, M rows / columns of reference samples are employed in MH-CCP, while for the second coding block 406B, N rows / columns of reference samples are employed in MH-CCP. M and N are different positive numbers.

[0129] In some embodiments, the selection of the number of reference rows is inferred by some coding information (e.g., partition, block size, block shape, resolution).

[0130] In some embodiments, if the resolution / block size is greater than a threshold, more columns / rows are used.

[0131] In another embodiment, if the resolution / block size is less than a threshold, more columns / rows are used.

[0132] In some embodiments, the selection of the number of reference rows is signaled. A plurality of predetermined numbers of columns / rows are stored in a lookup table, and an index of the selected number is signaled. For example, there are two selections 3 and 6, and the index 0 or 1 signaled represents the selection. Signaling occurs at one or more of block level, super-block level, frame level, key-frame level, and sequence level.

[0133] Although Figure 9 multiple logical stages are shown in a particular order, stages that are not order-dependent may be reordered, and other stages may be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0134] Now, turning to some exemplary embodiments.

[0135] (B1)In some implementations, a method 900 for decoding video data is implemented. Method 900 includes: receiving (902) a video bitstream including a first coded block and a second coded block of a current picture frame, wherein the video bitstream includes (904) a syntax element for cross-component intra prediction (CCIP) mode, and the CCIP mode indicates whether each chrominance sample of the first coded block and the second coded block is determined based on one or more luminance samples. The method further includes: for each of the first coded block and the second coded block (906): generating (912) a plurality of hypothesis values based on a plurality of neighboring luminance samples of a first luminance sample; identifying (914) a reference region of the corresponding coded block, wherein the size of the reference region of the first coded block is different from the size of the reference region of the second coded block; determining (916) a plurality of weighting factors based on a set of one or more reference samples in the reference region; and combining (918) the first luminance sample and the plurality of hypothesis values based on the plurality of weighting factors to generate a first chrominance sample co-located with the first luminance sample. The method further includes: (920) reconstructing the current picture frame including the first chrominance sample of each of the first coded block and the second coded block.

[0136] (B2)In some embodiments of B1, for each of the first coded block and the second coded block, the reference region includes one or more coded blocks adjacent to and decoded before the corresponding coded block.

[0137] (B3)In some embodiments of B1 or B2, a set of one or more reference samples in the reference region of the first coded block includes a first number of rows of reference samples, and a set of one or more reference samples in the reference region of the second coded block includes a second number of rows of reference samples, and the first number is not equal to the second number.

[0138] (B4)In some embodiments of B3, method 900 further includes: determining the first number based on a subset of block partitioning, block size, block shape, and block resolution of the first coded block; and determining the second number based on a subset of block partitioning, block size, block shape, and block resolution of the second coded block.

[0139] (B5)In some embodiments of B3 or B4, the block resolution of the first coded block is greater than a resolution threshold; the block resolution of the second coded block is less than the resolution threshold; and the first number is greater than the second number.

[0140] (B6)In some embodiments of B3 or B4, the block resolution of the first coded block is greater than a resolution threshold; the block resolution of the second coded block is less than the resolution threshold; and the first number is less than the second number.

[0141] In some embodiments of any one of B3 to B6, the block size of the first coding block is greater than a block size threshold; the block size of the second coding block is less than the block size threshold; and the first quantity is greater than the second quantity.

[0142] (B8) In some embodiments of any one of B3 to B6, the block size of the first coding block is greater than a block size threshold; the block size of the second coding block is less than the block size threshold; and the first quantity is less than the second quantity.

[0143] (B9) In some embodiments of any one of B3 to B8, each of the reference samples of the first quantity of rows and the reference samples of the second quantity of rows includes the reference samples of the corresponding rows in the corresponding reference region.

[0144] (B10) In some embodiments of any one of B3 to B8, each of the reference samples of the first quantity of rows and the reference samples of the second quantity of rows includes the reference samples of the corresponding columns in the corresponding reference region.

[0145] (B11) In some embodiments of any one of B3 to B10, the video bitstream signals a reference region index for each of the first coding block and the second coding block, and method 900 further includes: determining a first quantity based on the reference region index of the first coding block; and determining a second quantity based on the reference region index of the second coding block.

[0146] (B12) In some embodiments of B11, a lookup table is used to determine the first quantity based on the reference region index of the first coding block, and a lookup table is used to determine the second quantity based on the reference region index of the second coding block.

[0147] (B13) In some embodiments of B11 or B12, for each of the first coding block and the second coding block, in the video bitstream, the reference region index is signaled at one of a block level, a superblock level, an image frame level, a key frame level, and an image sequence level of the corresponding coding block.

[0148] (B14) In some embodiments of any one of B1 to B13, for each of the first coding block and the second coding block, the reference region of the corresponding coding block includes one or more of the following: the upper left reference region, the upper reference region, the upper right reference region, the lower left reference region, and the left reference region of the corresponding coding block.

[0149] (B15) In some embodiments of any one of B1 to B14, combining the first luminance sample and a plurality of hypothesis values further includes: combining the first luminance sample and the plurality of hypothesis values with at least one of (1) a non - linear term of the first luminance sample and a subset of a plurality of adjacent luminance samples and (2) a bias term based on a plurality of weighting factors.

[0150] (B16)In some embodiments of any one of B1 to B15, to determine a plurality of weighting factors, method 900 further includes: determining a least mean square (LMS) value based on a set of one or more reference samples; and iteratively adjusting the plurality of weighting factors to reduce the LMS value until the LMS value meets a predetermined criterion.

[0151] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more instruction sets configured to be executed by the control circuitry, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., B1 to B16 above).

[0152] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets executed by a control circuit of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., B1 to B16 above).

[0153] The proposed methods can be used individually or combined in any order. Additionally, each method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute a program stored in a non-transitory computer-readable medium. Hereinafter, the term "block" may be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU).

[0154] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0155] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to any and all possible combinations of one or more of the listed related items and includes any and all possible combinations of one or more of the listed related items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0156] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the stated precondition is true, depending on the context. Similarly, the phrase "if it is determined [that the stated precondition is true]" or "if [the stated precondition is true]" or "when [the stated precondition is true]" may be interpreted to mean "upon determining" or "in response to determining" or "in accordance with determining" or "upon detecting" or "in response to detecting" that the stated precondition is true, depending on the context.

[0157] For purposes of explanation, the foregoing description has been presented with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to utilize the invention.

Claims

1. A video data decoding method, comprising: Receive a video bitstream including a current coded block of a current image frame, where the video bitstream includes: (i) a first syntax element for a cross-component intra prediction (CCIP) mode, the CCIP mode indicating whether each chrominance sample of the current coded block is determined based on one or more luma samples; and (ii) a second syntax element for a hypothesized tap index, the hypothesized tap index selecting at least one hypothesized tap combination from a plurality of hypothesized tap combinations for the current coded block; Select the one hypothesized tap combination from the plurality of hypothesized tap combinations based on the hypothesized tap index; Generate a plurality of hypothesized values based on a plurality of neighboring luma samples of a first luma sample based on the one hypothesized tap combination selected from the plurality of hypothesized tap combinations; Generate a first chrominance sample collocated with the first luma sample based at least on the first luma sample and the plurality of hypothesized values; and Reconstruct the current coded block including the first chrominance sample.

2. The method according to claim 1, wherein, In the video bitstream, the hypothesized tap index is signaled at one of a block level, a superblock level, an image frame level, a key frame level, and an image sequence level of the current coded block.

3. The method according to claim 1, wherein, The first luma sample is surrounded by eight nearest luma sample candidates, and the plurality of hypothesized tap combinations include: A first hypothesized tap combination corresponding to the first luma sample and two neighboring luma samples selected from the eight nearest luma sample candidates; and A second hypothesized tap combination corresponding to the first luma sample and four neighboring luma samples selected from the eight nearest luma sample candidates.

4. The method according to claim 1, wherein, The plurality of hypothesized tap combinations include: A first hypothesized tap combination corresponding to the first luma sample, a north luma sample, a south luma sample, a west luma sample, and an east luma sample; and A second hypothesized tap combination corresponding to the first luma sample, a northwest luma sample, a northeast luma sample, a southwest luma sample, and a southeast luma sample.

5. The method according to claim 1, wherein, The plurality of hypothesized tap combinations include: A first hypothesized tap combination corresponding to the first luma sample, a north luma sample directly above the first luma sample, and a south luma sample directly below the first luma sample; A second hypothesized tap combination corresponding to the first luma sample, a west luma sample directly to the left of the first luma sample, and an east luma sample directly to the right of the first luma sample; A third hypothesized tap combination corresponding to the first luma sample, a northwest (NW) luma sample, and a southeast (SE) luma sample, where positions of the northwest (NW) luma sample and the southeast (SE) luma sample are symmetric with respect to a position of the first luma sample; and A fourth hypothesized tap combination corresponding to the first luminance sample, a south-west (SW) luminance sample, and a north-east (NE) luminance sample, wherein the positions of the south-west (SW) luminance sample and the north-east (NE) luminance sample are symmetric about the position of the first luminance sample.

6. The method according to claim 1, wherein, According to the hypothesized tap index, one of the plurality of hypothesized tap combinations corresponds to the first luminance sample, a north luminance sample directly above the first luminance sample, and a south luminance sample directly below the first luminance sample.

7. The method according to claim 1, wherein, According to the hypothesized tap index, one of the plurality of hypothesized tap combinations corresponds to the first luminance sample, a west luminance sample directly to the left of the first luminance sample, and an east luminance sample directly to the right of the first luminance sample.

8. The method according to claim 1, wherein, The plurality of hypothesized tap combinations includes: A first hypothesized tap combination corresponding to the first luminance sample, a north luminance sample directly above the first luminance sample, and a south luminance sample directly below the first luminance sample; and A second hypothesized tap combination corresponding to the first luminance sample, two or more west luminance samples directly to the left of the first luminance sample, and two or more east luminance samples directly to the right of the first luminance sample.

9. The method according to claim 1, wherein, Generating the first chrominance sample based at least on the first luminance sample and the plurality of hypothesized values further includes: Using a plurality of weighting factors to combine the first luminance sample, the plurality of hypothesized values with at least one of (1) a non-linear term of the first luminance sample and a subset of the plurality of adjacent luminance samples and (2) a bias term.

10. The method according to claim 9, further comprising: Determining the plurality of weighting factors based on a set of one or more luminance samples and a set of one or more co-located chrominance samples within a reference region of the current coded block, wherein the reference region is located within the current image frame.

11. A video data decoding method, comprising: Receiving a video bitstream including a first coded block and a second coded block of a current image frame, wherein the video bitstream includes syntax elements for a cross-component intra prediction (CCIP) mode that indicates whether each chrominance sample of the first coded block and the second coded block is determined based on one or more luminance samples; For each of the first coded block and the second coded block: Generating a plurality of hypothesized values based on a plurality of adjacent luminance samples of the first luminance sample; Identifying a reference region of the corresponding coded block, wherein the size of the reference region of the first coded block is different from the size of the reference region of the second coded block; Determining a plurality of weighting factors based on a set of one or more reference samples within the reference region; Combining the first luminance sample and the plurality of hypothesized values based on the plurality of weighting factors to generate a first chrominance sample co-located with the first luminance sample; and Reconstructing the current image frame including the first chrominance sample of each of the first coded block and the second coded block.

12. According to the method of claim 11, wherein, For each of the first coded block and the second coded block, the reference region includes one or more coded blocks that are adjacent to and decoded before the corresponding coded block.

13. According to the method of claim 11, wherein, The set of one or more reference samples in the reference region of the first coded block includes a first number (M) of rows of reference samples, and the set of one or more reference samples in the reference region of the second coded block includes a second number (N) of rows of reference samples, where the first number is not equal to the second number.

14. According to the method of claim 13, further comprising: Determine the first number based on a subset of the block partition, block size, block shape, and block resolution of the first coded block; and Determine the second number based on a subset of the block partition, block size, block shape, and block resolution of the second coded block.

15. According to the method of claim 13, wherein: The block resolution of the first coded block is greater than a resolution threshold; The block resolution of the second coded block is less than the resolution threshold; and The first number is greater than the second number.

16. According to the method of claim 13, wherein: The block resolution of the first coded block is greater than a resolution threshold; The block resolution of the second coded block is less than the resolution threshold; and The first number is less than the second number.

17. According to the method of claim 13, wherein: The block size of the first coded block is greater than a block size threshold; The block size of the second coded block is less than the block size threshold; and The first number is greater than the second number.

18. According to the method of claim 13, wherein: The block size of the first coded block is greater than a block size threshold; The block size of the second coded block is less than the block size threshold; and The first number is less than the second number.

19. According to the method of claim 13, wherein, Each of the first number of rows of reference samples and the second number of rows of reference samples includes the reference samples of the corresponding row in the corresponding reference region.

20. According to the method of any one of claims 13, wherein, Each of the first number of rows of reference samples and the second number of rows of reference samples includes the reference samples of the corresponding column in the corresponding reference region.

21. According to the method of any one of claims 13, wherein, The video bitstream signals a reference region index for each of the first coded block and the second coded block, and the method further includes: Determine the first number based on the reference region index of the first coded block; and Determine the second number based on the reference region index of the second coded block.

22. A computing system, comprising: Control circuit; A memory that stores one or more programs, the one or more programs being configured to be executed by the control circuit, the one or more programs further including instructions for: Receiving a video bitstream including a first coded block and a second coded block of a current image frame, where the video bitstream includes syntax elements for cross-component intra prediction (CCIP) mode, and the CCIP mode indicates whether each chrominance sample of the first coded block and the second coded block is determined based on one or more luma samples; For each of the first coded block and the second coded block: Generating a plurality of hypothesis values based on a plurality of adjacent luma samples of a first luma sample; Identifying the reference region of the corresponding coded block, where the size of the reference region of the first coded block is different from the size of the reference region of the second coded block; Determining a plurality of weighting factors based on a set of one or more reference samples in the reference region; Combining the first luminance sample and the plurality of hypothesized values based on the plurality of weighting factors to generate a first chrominance sample collocated with the first luminance sample; and Reconstructing the current picture frame with the first chrominance samples of each of the first coded block and the second coded block.

23. The computing system according to claim 22, wherein, Using a lookup table, determining the first quantity based on a reference region index of the first coded block, and using the lookup table, determining the second quantity based on a reference region index of the second coded block.

24. The computing system according to claim 22, wherein, For each of the first coded block and the second coded block, in the video bitstream, the reference region index is signaled at one of a block level, a superblock level, a picture frame level, a keyframe level, and a picture sequence level of the corresponding coded block.

25. The computing system according to claim 22, wherein, For each of the first coded block and the second coded block, the reference region of the corresponding coded block includes one or more of the following: a top-left reference region, an upper reference region, a top-right reference region, a bottom-left reference region, and a left reference region of the corresponding coded block.

26. The computing system according to claim 22, wherein, The combining of the first luminance sample and the plurality of hypothesized values further comprises: Combining the first luminance sample, the plurality of hypothesized values with at least one of (1) a non-linear term of the first luminance sample and a subset of the plurality of neighboring luminance samples and (2) a bias term, based on the plurality of weighting factors.

27. The computing system according to claim 22, determining that the plurality of weighting factors further includes: Determining a least mean square (LMS) value based on a set of the one or more reference samples; and Iteratively adjusting the plurality of weighting factors to reduce the LMS value until the LMS value meets a predetermined criterion.