Multi-hypothesis cross-component prediction model
Through the multi-assumption cross-component prediction (MH-CCP) method, the relationship between brightness and chromaticity samples is used to predict, and the video encoding process is optimized, which solves the problem of difficult to balance encoding efficiency and quality in the prior art, and achieves more efficient video data compression.
Patent Information
- Application Number
- CN202480004814.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2024-05-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing video encoding technology is difficult to effectively utilize the redundant information of video data during the compression process, which makes it difficult to optimize the balance between encoding efficiency and quality loss.
Multi-assumption Cross-component prediction (MH-CCP) method is used to predict using the relationship between the luminance sample and the chromaticity sample through linear or nonlinear weighting and filtering technology, and a multi-tap model is used to generate chromaticity samples to optimize the encoding process.
Improve the efficiency and quality of video encoding, reduce the amount of data, while maintaining the availability of video and the decoded quality.
Smart Images

Figure CN120283405A_ABST
Abstract
Description
Cross - reference
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 597,329, titled "Multi - Hypothesis Cross - Component Prediction Model", filed on November 8, 2023, and U.S. Provisional Patent Application No. 63 / 604,095, titled "Multi - Hypothesis Cross - Component Prediction Model", filed on November 29, 2023, and this application is a continuation - in - part of and claims priority to U.S. Patent Application No. 18 / 622,837, titled "Multi - Hypothesis Cross - Component Prediction Model", filed on March 29, 2024. Technical Field
[0002] Embodiments of the present disclosure generally relate to video coding and decoding, including but not limited to systems and methods for processing video data using multi - hypothesis cross - component prediction (MH - CCP). Background Art
[0003] Various electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. These electronic devices send and receive digital video data via a communication network or otherwise transmit digital video data, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video data can be compressed using video coding according to one or more video coding standards before being transmitted or stored. Video coding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.
[0004] Video coding is typically performed using prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the Moving Picture Experts Group-High-Efficiency (MPEG-H) project. The International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4), respectively. Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2), respectively. The Alliance for Open Media (AOMedia) Video (AV1) format is an open video coding format designed as an alternative to HEVC. A verified version 1.0.0 containing Errata 1 was released on January 8, 2019. Summary of the Invention
[0005] As described above, encoding (compression) reduces the demand for bandwidth and / or storage space. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal through a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not completely retained during the encoding process and cannot be fully restored during the decoding process. When using lossy compression, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achieved by a specific encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows encoding and decoding algorithms with higher losses and higher compression ratios.
[0006] The present disclosure describes a video compression method using intra prediction. A linear or non-linear weighted sum of multiple types of luma samples is used to predict chroma samples, for example, in multi-hypothesis cross-component prediction (MH-CCP). The multiple types of luma samples include luma sample C co-located with the chroma sample and filtered luma samples determined based on neighboring luma samples and used as filter inputs. Each filter input of the weighted sum is referred to as a hypothesis. Each hypothesis is associated with a weighting factor in MH-CCP. In one aspect of the present application, the weighting factors are applied to generate a linear or non-linear weighted sum of different types of luma samples. These weighting factors are determined for each coding block based on the reference region of the coding block. In some embodiments, these weighting factors are determined by applying a least mean square calculation kernel to process the reconstructed samples of the reference block of each coding block.
[0007] In other words, in some embodiments, samples of a second color component are predicted as a linear or non-linear weighted sum of samples of the second color component co-located with the samples of the second color component and one or more associated adjacent luminance samples according to a multi-tap model associated with MH-CCP. The multi-tap model includes multi (N) taps, which are selected from co-located samples of the second color component, one or more associated adjacent luminance samples of the second color component, non-linear terms, and offset terms. The multi-tap model corresponds to the same number (N) of selected items, which are combined to determine samples of the second color component. In some embodiments, samples of the first color component are one of a chrominance sample and a luminance sample, and samples of the second color component are luminance samples. For example, the chrominance sample is a weighted combination of terms selected from corresponding co-located luminance samples, one or more adjacent luminance samples, non-linear terms, and offset terms. Alternatively, in some embodiments, the first color component is one of red, green, and blue, and the second color component is another one of red, green, and blue. Alternatively, in some embodiments, the first color component and the second component correspond to a color format different from the YCbCr color format and the RGB color format.
[0008] According to some embodiments, a method of video decoding is provided. The method includes: receiving a video bitstream including a current encoded block of a current picture frame. The video bitstream includes a first syntax element for the MH-CCP mode. The method further includes: determining, based on the first syntax element in the video bitstream, to enable the MH-CCP mode to reconstruct each chrominance sample by using at least a corresponding luminance sample co-located with each of a plurality of chrominance samples of the current encoded block and one or more adjacent luminance samples corresponding to the corresponding luminance sample. The method further includes: determining, at least for the current encoded block, a number (N) of model parameters to be used in the MH-CCP mode; identifying one or more adjacent luminance samples of a first luminance sample based on the number (N) of the model parameters; generating a first chrominance sample co-located with the first luminance sample based on the first luminance sample and the one or more adjacent luminance samples; and reconstructing the current encoded block including the first chrominance sample.
[0009] According to some embodiments, a video encoding method is provided. The method includes: receiving video data including a current coded block of a current picture frame; encoding the current picture frame according to intra prediction; determining to enable the MH-CCP mode to determine each chroma sample based on a corresponding luma sample co-located with each chroma sample of the current coded block and one or more neighboring luma samples corresponding to the corresponding luma sample, wherein the MH-CCP mode is associated with a number (N) of model parameters for identifying one or more neighboring luma samples of a first luma sample. The method further includes: transmitting the encoded current picture frame via a video bitstream; and signaling, via the video bitstream, a first syntax element to indicate applying the MH-CCP mode to reconstruct a first chroma sample co-located with the first luma sample based at least on the first luma sample and the one or more neighboring luma samples.
[0010] According to some embodiments, a bitstream conversion method is provided. The method includes: obtaining a source video sequence including a current coded block of a current picture frame; and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes: the current coded block of the current picture frame; and a first syntax element for the MH-CCP mode, the first syntax element indicating whether to reconstruct each chroma sample based at least on a corresponding luma sample co-located with each chroma sample of the current coded block and one or more neighboring luma samples corresponding to the corresponding luma sample. The MH-CCP mode is associated with a number (N) of model parameters for identifying one or more neighboring luma samples of a first luma sample, and the number (N) of model parameters is applied to reconstruct a first chroma sample co-located with the first luma sample based at least on the first luma sample and the one or more neighboring luma samples.
[0011] According to some embodiments, a method for video decoding is provided. The method includes: receiving a video bitstream that includes a current coded block of a current picture frame. The video bitstream includes a first syntax element for the MH-CCP mode. The method further includes: determining, based on the first syntax element in the video bitstream, to enable the MH-CCP mode to reconstruct each sample of the second color component of the current coded block by using at least the corresponding sample of the first color component co-located with each sample of the second color component of the current coded block and one or more neighboring samples of the second color component corresponding to the corresponding sample of the first color component. The method further includes: determining, at least for the current coded block, a number (N) of model parameters to be used in the MH-CCP mode; identifying one or more neighboring samples of a first sample of the first color component based on the number (N) of model parameters; generating a first sample of the second color component co-located with the first sample of the first color component based on the first sample of the first color component and the one or more neighboring samples of the first color component; and reconstructing the current coded block including the first sample of the second color component.
[0012] According to some embodiments, a computing system is provided, which is, for example, a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).
[0013] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets executable by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0014] Thus, devices and systems are disclosed that use methods for encoding and decoding video. Such methods, devices, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding. The features and advantages described in this specification need not be all-inclusive, and in particular, considering the drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Further, it should be noted that the language used in the specification is primarily selected for readability and guidance purposes and not necessarily to define or circumscribe the subject matter described herein. Description of the Drawings
[0015] To understand the present disclosure in more detail, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure, and these features need not be considered restrictive, because those skilled in the art will understand when reading the present disclosure that these descriptions can accommodate other valid features.
[0016] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0017] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0018] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0019] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0020] Figure 4 shows an example scheme 400 for generating a first chrominance sample from multiple luminance samples in the MH-CCP mode according to some embodiments.
[0021] Figure 5 shows an example scheme for generating a first chrominance sample in the MH-CCP mode corresponding to a two-parameter model having two model parameters according to some embodiments.
[0022] Figure 6 shows an example scheme for generating a first chrominance sample in the MH-CCP mode corresponding to a three-parameter model having three model parameters according to some embodiments.
[0023] Figure 7 shows an example scheme for generating a first chrominance sample in the MH-CCP mode corresponding to a four-parameter model having four model parameters according to some embodiments.
[0024] Figure 8 shows an example position of a current coding block relative to a current super block according to some embodiments.
[0025] Figure 9 is a flowchart showing a video decoding method according to some embodiments.
[0026] By convention, the various features shown in the drawings need not be drawn to scale, and throughout the specification and drawings, like reference numerals may be used to denote like features. Detailed Description
[0027] The present disclosure describes cross-component intra prediction of video data in the MH-CCP mode, wherein each sample of a plurality of samples of a first color component is determined based on one or more associated samples of a second color component. The MH-CCP mode corresponds to a multi-tap model including multiple (N) taps. Each tap is selected from co-located samples of the second color component, one or more associated neighboring luma samples of the second color component, a non-linear term, and an offset term. The selected taps are combined in a weighted manner to determine a sample of the second color component. In some embodiments, the sample of the first color component is one of a chroma sample and a luma sample, and the sample of the second color component is a luma sample. For example, the chroma sample is a weighted combination of terms selected from corresponding co-located luma samples, one or more neighboring luma samples, a non-linear term, and an offset term. In some implementations, the number of taps can be selected from 2, 3, 4, and 5. In one aspect of the present application, a computing device receives a video bitstream that includes a current encoded block of a current image frame and a first syntax element for the MH-CCP mode, and the computing device identifies a five-tap model configured to determine a chroma sample of the current encoded block in the MH-CCP mode. A pair of neighboring luma samples of a first luma sample is determined based on the five-tap model, and the first luma sample, the non-linear term, and the offset term are used to generate a first chroma sample co-located with the first luma sample.
[0028] Figure 1 FIG. 4 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), and the source device 102 and the plurality of electronic devices 120 are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for use with video-enabled applications such as video conferencing applications, digital television (TV) applications, media storage and / or distribution applications.
[0029] The source device 102 includes a video source 104 (e.g., a camera assembly or media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may have a higher data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less transmission bandwidth and less storage space for storage. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to send uncompressed video to one or more networks 110).
[0030] The one or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. The one or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks may include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0031] The one or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode the encoded video bitstreams 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encode based on the encoded video bitstreams 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to crop the encoded video bitstreams 108 for customizing potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0032] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to produce an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices in the electronic device 120 do not include a display component (e.g., the electronic device is communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0033] The source device and / or the plurality of electronic devices 120 may also be referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0034] In an example operation of the communication system 100, the source device 102 sends the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using the codec component 114. For example, the server system 112 may encode video data that is more suitable for network transmission and / or storage. The server system 112 may send the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video picture.
[0035] Figure 2AFIG. 0 is a block diagram showing example elements of an encoder assembly 106 in accordance with some embodiments. The encoder assembly 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder assembly includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder assembly 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a different device than the encoder assembly 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In some embodiments, the video source 104 may be a storage device storing previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0036] The encoding component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder assembly 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or lambda value of rate-distortion optimization techniques, etc.), picture size, group of picture (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily discover other functions of the controller 204 as these functions may be related to the encoder assembly 106 optimized for a particular system design.
[0037] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and a reference picture) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to that of the (remote) decoder (when the compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory 208 also correspond bit-accurately between the local encoder and the remote encoder. As a result, the prediction part of the encoder interprets the reference picture samples as exactly the same sample values as the decoder would interpret when using prediction during decoding. This reference picture synchronization principle (and the drift that occurs when synchronization cannot be maintained, for example due to channel errors) is known to those of ordinary skill in the art.
[0038] The operation of decoder 210 may be combined with, for example, Figure 2B The remote decoders such as decoder assembly 122 described in detail are the same. However, brief reference is made to Figure 2B , when symbols are available and the entropy encoder 214 and the parser 254 are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.
[0039] In addition to parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverse to decoder techniques.
[0040] As part of this operation, the source encoder 202 may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes the input frame with reference to one or more previously encoded frames from the video sequence that are designated as reference frames. In this manner, the encoding engine 212 encodes the differences between the pixel blocks of the input frame and the pixel blocks of one or more reference frames that may be selected as prediction references for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.
[0041] The decoder 210 decodes the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 202. Advantageously, the operation of the encoding engine 212 can be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 duplicates the decoding process that can be performed by a remote video decoder on the reference frame and can cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder assembly 106 locally stores a copy of the reconstructed reference frame that has the same content (in the absence of transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0042] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as a candidate reference pixel block) or some metadata, such as a reference picture motion vector, block shape, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel-block basis of sample blocks to find a suitable prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.
[0043] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.
[0044] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214, thus preparing for transmission over a communication channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The transmitter may be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such other data being, for example, encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / signal noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Video Usability Information (VUI) parameter set segments, and the like.
[0045] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, pictures may generally be classified as: intra pictures (I pictures), predictive pictures (P pictures), or bi-predictive pictures (B pictures). Intra pictures may be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and features, and thus the variants and their corresponding applications and features are not repeated herein. Predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0046] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the block can be predictively encoded with reference to the already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded with reference to a previously encoded reference picture, either through spatial prediction or through temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures, either through spatial prediction or through temporal prediction.
[0047] The video captured can be a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0048] The encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard described herein, for example. In operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.
[0049] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 shown is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to a loop filter 256 and is configured to send data to the display 124 (e.g., via a wired or wireless connection).
[0050] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, each encoded video sequence is decoded independently of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective consuming entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0051] According to some embodiments, decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.
[0052] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to prevent network jitter). While it may not be necessary to configure buffer memory 252 or the buffer memory may be made smaller when receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network. For use on a best-effort network such as the Internet, buffer memory 252 may be required, which may be relatively large and / or have an adaptive size and may be implemented at least partially in the operating system or a similar element external to decoder component 122.
[0053] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols may include, for example, information for managing the operation of the decoder components 122 and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI) messages or parameter set fragments of the Video Usability Information (VUI) (not labeled). The parser 254 may parse (entropy decode) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and the like.
[0054] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol 270 may involve multiple different units. Which units are involved and the way these units are involved may be controlled by subgroup control information parsed by the parser 254 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 254 and the multiple units below are not described.
[0055] The decoder component 122 may be conceptually divided into multiple functional units, and in some embodiments, these functional units interact closely with each other and may be at least partially integrated with each other. However, for the sake of clarity, the functionally divided units are retained herein conceptually.
[0056] The Scaler / Inverse Transform Unit 258 receives the quantized transform coefficients and control information (e.g., which transform mode to use, block size, quantization factor, and / or quantization scaling matrix, etc.) as the symbol 270 from the parser 254. The Scaler / Inverse Transform Unit 258 may output a block including sample values, which may be input into the aggregator 268.
[0057] In some cases, the output samples of the scaler / inverse transform unit 258 may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra picture prediction unit 262. The intra picture prediction unit 262 may generate a block having the same size and shape as the block being reconstructed using surrounding reconstructed information extracted from the current (partially reconstructed) picture in the current picture buffer 264. The aggregator 268 may add, on a per-sample basis, the prediction information generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.
[0058] In other cases, the output samples of the scaler / inverse transform unit 258 may belong to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After motion-compensating the extracted samples according to the sign 270 belonging to the block, these samples may be added, via the aggregator 268, to the output of the scaler / inverse transform unit 258 (referred to as residual samples or a residual signal in this case) to generate output sample information. The address from which the motion compensation prediction unit 260 obtains the prediction samples from within the reference picture memory 266 may be controlled by a motion vector. The motion vector may be available to the motion compensation prediction unit 260 in the form of the sign 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of the sample values extracted from the reference picture memory 266, a motion vector prediction mechanism, etc., when using sub-sample accurate motion vectors.
[0059] The output samples of the aggregator 268 may be employed by various loop filtering techniques in the loop filter unit 256. The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and that may be available to the loop filter unit 256 as the sign 270 from the parser 254. However, the video compression technique may also respond to meta-information obtained during the decoding of a previously (in decoding order) part of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device such as the display 124 and stored in the reference picture memory 266 for subsequent inter-picture prediction.
[0060] Once reconstructed, certain encoded pictures can be used as reference pictures for subsequent prediction. Once an encoded picture has been fully reconstructed and the encoded picture (by, for example, parser 254) is identified as a reference picture, the current reference picture can be made part of the reference picture memory 266 and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0061] Decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be recorded in a standard (such as any of the standards described herein). An encoded video sequence can conform to the syntax specified by the video compression technique or standard being used in the sense that the encoded video sequence follows the syntax specified in the video compression technique document or standard (in particular, the syntax of the video compression technique or standard specified in the profile document thereof). Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0062] Figure 3 is a block diagram showing a server system 112 according to some embodiments. Server system 112 includes control circuitry 302, one or more network interfaces 304, memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, control circuitry 302 includes one or more processors (e.g., a central processing unit (CPU), a Graphics Processing Unit (GPU), and / or a Data Processing Unit (DPU)). In some embodiments, the control circuitry includes one or more Field-Programmable Gate Arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).
[0063] The network interface 304 can be configured to connect to one or more communication networks (e.g., wireless network, wired network, and / or optical network). The communication network can be a local area network, wide area network, metropolitan area network, vehicle and industrial network, real-time network, low-latency network, etc. Examples of communication networks include local area networks such as Ethernet, wireless local area networks (LANs), cellular networks including Global System for Mobile Communications (GSM), 3rd generation (3G), 4th generation (4G), 5th generation (5G), Long Term Evolution (LTE), etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial television broadcasting, vehicle and industrial television including Controller Area Network Bus (CANbus), etc. Such communication can be one-way reception only (e.g., television broadcasting), one-way transmission only (e.g., CANbus connected to certain CANbus devices), or two-way (e.g., connecting to other computer systems using a local area network or wide area digital network). Such communication can include communication to one or more cloud computing networks.
[0064] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.
[0065] Memory 314 may include high-speed random access memory (such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDRRAM), and / or other random access solid-state memory devices), and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices remote from the control circuit 302. Memory 314 or, optionally, the non-volatile solid-state memory device within memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, memory 314 or the non-transitory computer-readable storage medium of memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · Operating system 316, including procedures for handling various basic system services and for performing hardware-related tasks; · Network communication module 318, for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); ● Codec module 320, for performing various functions related to encoding and / or decoding data (such as video data). In some embodiments, codec module 320 is an instance of codec component 114. Codec module 320 includes, but is not limited to, one or more of the following modules: o Decoding module 322, for performing various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and o Encoding module 340, for performing various functions related to encoding data, such as those previously described with respect to encoder component 106; and · Picture memory 352, for storing pictures and picture data, e.g., for use with codec module 320. In some embodiments, picture memory 352 includes one or more of the following memories: reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.
[0066] In some embodiments, the decoding module 322 includes: a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform the various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).
[0067] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform the various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, the shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0068] Each of the above-mentioned modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The above-mentioned modules (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, optionally, the encoding module 320 does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above-mentioned modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.
[0069] Although Figure 3 a server system 112 is shown in accordance with some embodiments, however, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems rather than as a structural diagram of the embodiments described herein. In practice, and as will be appreciated by those of ordinary skill in the art, the separately shown items can be combined, and some items can be separated. For example, Figure 3 some of the items separately shown in can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among these servers will vary depending on the implementation, and optionally, will partly depend on the data traffic processed by the server system during peak usage periods as well as during average usage periods.
[0070] Figure 4 illustrates an example scheme 400 for generating a first chroma sample 402C based on a plurality of luminance samples 404 (e.g., 404C and 404X) in the MH-CCP mode according to some embodiments. In some embodiments, the current coding block 406A of the current image frame 408 is encoded in the Cross-Component Intra Prediction (CCIP) mode. In the CCIP mode, the decoder 122 ( Figure 2B ) determines each chroma sample among the plurality of chroma samples 402 of the current coding block 406A based on one or more reconstructed luminance samples 404. In some cases, the CCIP mode includes the Cross-Component Linear model Mode (CCLM), under which the first chroma sample 402C is transformed from the reconstructed luminance sample 404C co-located with the chroma sample based on a linear model. Alternatively, in some cases, the CCIP mode includes the Convolutional Cross-Component Mode (CCCM), under which the first chroma sample 402C is directly predicted from a plurality of reconstructed luminance samples 404X adjacent to the first luminance sample 404C based on the filter shape of a filter. Alternatively and additionally, in some cases, the CCIP mode includes the MH-CCP mode, under which the first chroma sample 402C is generated by combining at least the first luminance sample 404C co-located with the first chroma sample 402C and a plurality of hypothesis values 410 using a plurality of weighting factors 416. A plurality of adjacent luminance samples 404X of the first luminance sample 404C are combined using a plurality of coefficients to generate a plurality of hypothesis values 410.
[0071] In some embodiments, the video bitstream 116 includes a first syntax element 420 for the MH-CCP mode. The first chroma sample 402C of the current coding block 406A is configured to be generated by combining at least the first luminance sample 404C co-located with the first chroma sample 402C and one or more adjacent luminance samples 404X of the first luminance sample 404C using a plurality of weighting factors 416 (e.g., wi, wP, wB, where i is equal to 0 or a positive integer). When it is determined that the MH-CCP mode is enabled, the first chroma sample 402C is predicted based on the following equation: predChromaVal = ∑w i L i + w P P + w B B (1) where predChromaVal is the predicted chroma value of the first chroma sample 402C; ∑w i L i represents the weighted sum of one or more luminance values L of one or more luminance samples 404 i ; i ranges from 0 to M; M is the number of linear terms of adjacent luminance samples 404X; P is a non - linear term; B is an offset term; w P and w B are weighted factors associated with the non - linear term and the offset term respectively. In an example, the non - linear term P equals (L0×L0 + B)>>bit depth, where L0 is the luminance value of the first luminance sample 404C, and "bit depth" is the number of bits required to represent internal data values during encoding and decoding. In some embodiments, the non - linear term P is determined based on a subset of a set composed of multiple adjacent luminance samples 404X and the first luminance sample 404C. In some embodiments, B is the middle luminance value of the luminance value range (e.g., 255 in the range of [0, 511]). Alternatively, in some embodiments, B is the median luminance value or the average luminance value of the luminance samples 404 of the current encoding block 406A.
[0072] In some embodiments, each of the one or more adjacent luminance samples 404X of the first luminance sample 404C is adjacent to the first luminance sample 404C and shares at least one corresponding side or vertex with the first luminance sample 404C. In some embodiments, the one or more adjacent luminance samples 404X include a subset or all of the following adjacent luminance samples: a north - adjacent luminance sample (also referred to as an upper luminance sample) 404N, a south - adjacent luminance sample (also referred to as a lower luminance sample) 404S, a west - adjacent luminance sample (also referred to as a left luminance sample) 404W, an east - adjacent luminance sample (also referred to as a right luminance sample) 404E, a northwest - adjacent luminance sample (also referred to as an upper - left luminance sample) 404NW, a southeast - adjacent luminance sample (also referred to as a lower - right luminance sample) 404SE, a southwest - adjacent luminance sample (also referred to as a lower - left luminance sample) 404SW, and a northeast - adjacent luminance sample (also referred to as an upper - right luminance sample) 404NE. In some embodiments, the offset term B is generated based on the average value of a subset of the one or more adjacent luminance samples 404X of the first luminance sample 404C, and the first chroma sample 402C is generated based on the offset term B.
[0073] In some embodiments, the luminance samples 404 and chrominance samples 402 of the current coding block have different resolutions corresponding to a chrominance subsampling scheme (e.g., 4:2:2 or 4:2:0). Each luminance sample 404 includes a downsampled luminance sample generated from the reconstructed luminance samples using a downsampling filter. Alternatively, in some embodiments, each luminance sample 404 includes an original sample or a reconstructed luminance sample without any downsampling. That is, the first luminance sample 404C is reconstructed or downsampled to the resolution of the chrominance samples according to the resolution of the luminance samples. Thus, the adjacent luminance samples 404X (e.g., 404N, 404W, 404E, 404S, 404NW, 404NE, 404SW, 404SE) are also reconstructed or downsampled to the resolution of the chrominance samples according to the resolution of the luminance samples.
[0074] In some embodiments, a subset of the terms in Equation (1) is used to predict a first chrominance sample 402C that is co-located with the first luminance sample 404C. The number (N) of model parameters used in the MH-CCP mode is determined at least for the current coding block. One or more adjacent luminance samples 404X of the first luminance sample 404C are identified based on the number (N) of model parameters, and the one or more adjacent luminance samples 404X are applied to generate the first chrominance sample 402C that is co-located with the first luminance sample 404C. Additionally, in some embodiments, the number (N) of model parameters is selected from 2, 3, 4, and 5, and the number (N) of model parameters corresponds to a two-parameter model 414A, a three-parameter model 414B, a four-parameter model 414C, and a five-parameter model 414D, respectively. The two-parameter model 414A, the three-parameter model 414B, the four-parameter model 414C, and the five-parameter model 414D will be applied to reconstruct the first chrominance sample 402C of the current coding block 406A in the MH-CCP mode. More specifically, in some embodiments, if the number (N) of model parameters is equal to 2, two terms are selected in Equation (1) to provide the two-parameter model 414A for determining the first chrominance sample 402C. In some embodiments, if the number (N) of model parameters is equal to 3, three terms are selected in Equation (1) to provide the three-parameter model 414B for determining the first chrominance sample 402C. In some embodiments, if the number (N) of model parameters is equal to 4, four terms are selected in Equation (1) to provide the four-parameter model 414C for determining the first chrominance sample 402C. In some embodiments, if the number (N) of model parameters is equal to 5, five terms are selected in Equation (1) to provide the five-parameter model 414D for determining the first chrominance sample 402C. Alternatively, in some embodiments, if the number (N) of model parameters is equal to or greater than 6, six or more terms are selected in Equation (1) to provide a related model for determining the first chrominance sample 402C.
[0075] In one example, the quantity (N) equals 3, and equation (1) includes at least a single linear term of the first luminance sample 404C co-located with the first chrominance sample 402C. L0 is the luminance value of the first luminance sample 404C, and w0 is the associated weighting factor. In another example, the quantity (N) is equal to or greater than 4, and equation (1) includes a single linear term of the first luminance sample 404C and a quantity (e.g., M) of linear terms of adjacent luminance samples 404X of the first luminance sample 404C. L0 to LM represent the luminance values of the luminance samples 404, and w0 to wM are the associated weighting factors, where M is a positive integer and equal to N - 3.
[0076] In some embodiments, the video bitstream 116 further includes a second syntax element 422 that indicates the quantity (N) of the model parameters used in the MH-CCP mode at least for the current coding block 406A. Additionally, in some embodiments, the second syntax element 422 is signaled using one of a sequence header, a picture header, a sub-picture header, a slice header, and a tile header. In some embodiments, the second syntax element 422 is signaled as a block-level syntax element using one of a maximum coding block level, a coding block level, a prediction block level, a transform block level, and a predefined fixed block size level.
[0077] In contrast, in some embodiments, the video bitstream 116 does not include a second syntax element 422 that indicates the quantity (N) of the model parameters used in the MH-CCP mode for the current coding block 406A. The quantity (N) of the model parameters is determined based on the coding information shared between the encoder 106 and the decoder 122. The coding information includes one or more of the following items: the frame resolution of the current image frame 408, the quantization parameter of the current coding block 406A, the block size, the block shape, and the luminance prediction mode.
[0078] In some embodiments, the quantity (N) of the model parameters used in the MH-CCP mode is selected from a plurality of predefined values (e.g., 2, 3, 4, 5). One or more of the adjacent luminance samples 404X and the first luminance sample 404C are selected based on the quantity (N) of the model parameters. At least two of the plurality of predefined values correspond to two different selected luminance samples. The non-linear term P is determined based on the selected one of the one or more adjacent luminance samples 404X and the first luminance sample 404C. Further, the first chrominance sample 402C is generated based on the non-linear term P. In other words, the non-linear term P can be determined based on one or more luminance values L0 to LM of the luminance samples 404, where M is the quantity of the adjacent luminance samples 404X used to determine the first chrominance sample 402C.
[0079] In some embodiments, when it is determined that the number (N) of model parameters is equal to a first value (e.g., 3), a first non - linear term P is generated based on the first luminance sample 404C and the first luminance sample (e.g., the first luminance sample 404C) among one or more adjacent luminance samples 404X. When it is determined that the number (N) of model parameters is equal to a second value (e.g., 4) different from the first value (e.g., 3), a second non - linear term P is generated based on the first luminance sample 404C and a different second luminance sample (e.g., the average of adjacent luminance samples 404W and 404E) among one or more adjacent luminance samples 404X.
[0080] In some embodiments, the number (N) of model parameters corresponds to a plurality of hypothesized tap combinations. One or more adjacent luminance samples 404X of the first luminance sample 404C are identified by selecting one hypothesized tap combination among the plurality of hypothesized tap combinations corresponding to the number (N) of model parameters and identifying one or more adjacent luminance samples 404X of the first luminance sample 404C based on the one hypothesized tap combination among the plurality of hypothesized tap combinations. For example, the number (N) of model parameters is equal to 5, and the number of model parameters corresponds to one hypothesized tap combination among horizontal tap hypothesized tap combinations and vertical hypothesized tap combinations. When corresponding to the horizontal tap hypothesized tap combination, the first chroma sample 402C is predicted based on the following five - parameter model: predChromaVal = w0C + w1W + w2E + w P P + w B B (2) where C, W, and E are the luminance value of the first luminance value 404C, the luminance value of the left - adjacent luminance sample 404W, and the luminance value of the right - adjacent luminance sample 404E, respectively. Alternatively, when corresponding to the vertical tap hypothesized tap combination, the first chroma sample 402C is predicted based on the following five - parameter model: predChromaVal = w0C + w1N + w2S + w P P + w B B (3) where N and S are the luminance values of the upper - adjacent luminance sample 404N and the lower - adjacent luminance sample 404S, respectively.
[0081] In some embodiments, a plurality of weighting factors 416 (e.g., w i 、w P 、w B)。The reference region 412 is located in the current image frame 408. Additionally, in some embodiments, based on Equation (1), the reference luminance samples 404R of the reference region 412 are used to generate one or more chrominance samples 402C. In some embodiments, a set of one or more co-located reference chrominance samples 402R is compared with one or more regenerated chrominance samples to generate a Least Mean Square (LMS) value. The plurality of weighting factors 416 (e.g., w i , w P , w B ) are iteratively adjusted to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., the LMS value is less than a threshold LMS value, or the LMS value is minimized).
[0082] In some embodiments, at least one weighting factor 416 is derived based on chrominance samples and luminance samples within the reference region 412 of the current coding block 406A, and the reference region 412 includes one or more coding blocks that were decoded prior to the current coding block 406A (e.g., Figure 4 8 coding blocks are shown in Figure 4 ). In some embodiments, a subset of one or more coding blocks is adjacent to the current coding block 406A. In some embodiments, a subset of one or more coding blocks is separated from the current coding block 406A by one or more coding blocks. In some embodiments, the reference region 412 includes at least a portion of one or more rows above the current coding block 406A and / or a portion of one or more columns to the left of the current coding block 406A. For example, referring to Figure 4 , the reference region 412 includes 7 rows of luminance samples 404 above the current coding block 406A and 9 columns of luminance samples 404 to the left of the current coding block 406A.
[0083] In some embodiments, at least one weighting factor 416 is determined by minimizing the Mean Square Error (MSE) between the predicted chrominance samples 402 and the reconstructed chrominance samples 402 in the reference region 412. The MSE is minimized by calculating the autocorrelation matrix of the luminance samples 404 and the cross-correlation vector between the luminance samples 404R and the chrominance samples 402R of the reference region 412. The autocorrelation matrix is processed with LDL decomposition, and back substitution is used to calculate the plurality of weighting factors 416. This process generally follows the calculation of the filter coefficients of the adaptive loop filter (ALF) in the enhanced compression model (ECM) video coding. The LDL decomposition does not use square root operations and only uses integer arithmetic operations.
[0084] Figure 5FIG. 500 illustrates an example scenario for generating a first chroma sample 402C in an MH - CCP mode corresponding to a two - parameter model 414A with two model parameters. The number of model parameters (N) is equal to 2 and corresponds to the two - parameter model 414A, which is applied to reconstruct the first chroma sample 402C of the current coding block 406A in the MH - CCP mode, and the first chroma sample 402C is generated based on a weighted combination of a non - linear term P and an offset term B. In other words, the first chroma sample 402C is predicted based on the following two - parameter model: predChromaVal = w P P + w B B (4)
[0085] Additionally, in some embodiments, the non - linear term P is determined based on at least one of the eight neighboring luma samples 404X of the first luma sample 402C.
[0086] Figure 6 FIG. 600 illustrates an example scenario for generating a first chroma sample 402C in an MH - CCP mode corresponding to a three - parameter model 414B with three model parameters. The number of model parameters (N) is equal to 3 and corresponds to the three - parameter model 414B, which is applied to reconstruct the first chroma sample 402C of the current coding block 406A in the MH - CCP mode. The first chroma sample is generated based on a weighted combination of a first number K1 of linear terms (e.g., w i L i in Equation (1)), a second number K2 of non - linear terms (e.g., w P P in Equation (1)), and a third number K3 of offset terms (e.g., w B in Equation (1)). The sum of the first number K1, the second number K2, and the third number K3 is equal to 3, and the first number, the second number, and the third number are integers selected from the ranges [0, 3], [0, 3], and [0, 1], respectively. For example, the first chroma sample 402C is predicted based on any one of the following four - parameter models: predChromaVal=w0C+w1W+w B B (5.1) predChromaVal=w0C+w P1 P1+w P2 P2 (5.2) predChromaVal=w0C+w P P+w B B (5.3) predChromaVal = w P1 P1 + w P2 P2 + w B B(5.4) predChromaVal = w0W + w P P + w B B(5.5)
[0087] In some embodiments, the number (N) of the model parameters is equal to 3 and corresponds to the three-parameter model 414B, and the three-parameter model 414B is applied to reconstruct the first chroma sample 402C of the current coded block 406A based on a weighted sum of the first luma sample 404C, the non-linear term w P P, and the offset term w B B in the MH-CCP mode. The non-linear term w P P is determined based on at least one of the eight neighboring luma samples 404X (e.g., 404W, 404N, 404E, 404S, 404NW, 404NE, 404SW, 404SE) of the first luma sample 404C and the first luma sample 404C
[0088] Alternatively, in some embodiments, the number (N) of the model parameters is equal to 3 and corresponds to the three-parameter model 414B, and the three-parameter model 414B is applied to reconstruct the first chroma sample 402C of the current coded block 406A based on the first neighboring luma sample, the non-linear term w P P, and the offset term w B B in the MH-CCP mode. For example, the first neighboring luma sample is one of the eight neighboring luma samples 404X (e.g., 404W, 404N, 404E, 404S, 404NW, 404NE, 404SW, 404SE) of the first luma sample 404C. The non-linear term w P P is determined based on at least one of the first luma sample 404C, the first neighboring luma sample, and the set consisting of the remaining seven neighboring luma samples of the first luma sample. In other words, the non-linear term w P P is determined based on at least one of the eight neighboring luma samples 404X (e.g., 404W, 404N, 404E, 404S, 404NW, 404NE, 404SW, 404SE) of the first luma sample 404C and the first luma sample 404C
[0089] Figure 7FIG. 700 shows an example scenario for generating a first chroma sample 402C in the MH-CCP mode corresponding to a four-parameter model 414C with four model parameters. The number (N) of model parameters is equal to 4 and corresponds to the four-parameter model 414C, which is applied to reconstruct the first chroma sample 402C of the current coding block 406A in the MH-CCP mode. The first chroma sample 402C is generated based on a weighted combination of a first number K1 of linear terms, a second number K2 of non-linear terms, and a third number K3 of offset terms, where the sum of the first number K1, the second number K2, and the third number K3 is equal to 3, and the first number K1, the second number K2, and the third number K3 are integers selected from the ranges [0, 4], [0, 4], and [0, 1], respectively. For example, the first chroma sample 402C is predicted based on any one of the following four-parameter models: predChromaVal = w0C + w1W + w P P + w B B (6.1) predChromaVal = w0C + w1E + w P1 P1 + w P2 P2 (6.2) predChromaVal = w0C + w P1 P1 + w P2 P2 + w B B (6.3) predChromaVal = w P1 P1 + w P2 P2 + w P3 P3 + w P4 P4 (6.4)
[0090] In some embodiments (e.g., the embodiments related to Equation (6.1)), the number (N) of model parameters is equal to 4 and corresponds to the four-parameter model 414C, which is applied to reconstruct the first chroma sample 402C of the current coding block 406A in the MH-CCP mode based on a weighted sum of the first luminance sample 404C, one of the one or more adjacent luminance samples (e.g., the left adjacent luminance sample 404W), non-linear terms, and offset terms.
[0091] Figure 8Shows an example position 800 of a current coding block 406A relative to a current superblock 802 according to some embodiments. The current coding block 406A is included in the current superblock 802, which is the largest coding block of the current image frame 408. When it is determined that the current coding block 406A is adjacent to the upper boundary of the current superblock 802, a first upper reference region 412A is applied to determine a plurality of weighting factors for generating a first chrominance sample 402C. The first upper reference region 412A has a first number of rows (e.g., 6 rows) and is located immediately above the current coding block. When it is determined that the current coding block 406A is not adjacent to the upper boundary of the current superblock 802, a second upper reference region 412B having a second number of rows (e.g., 4 rows) and located immediately above the current coding block 406A is applied to determine a plurality of weighting factors for generating the first chrominance sample 402C. The first number is greater than the second number. In other words, when the current coding block 406A is adjacent to the upper boundary of the current superblock 802, more rows of luminance samples 404R are used to determine the plurality of weighting coefficients.
[0092] In some embodiments, the resolution of the first luminance sample 404C and one or more adjacent luminance samples 404X is different from the resolution of the first chrominance sample 402C. In some embodiments, the current superblock 802 further includes an alternative coding block 406B different from the current coding block 406A. When it is determined that the alternative coding block is adjacent to the upper boundary of the current superblock 802, the MH-CCP mode for the alternative coding block 406B is aborted. In some embodiments, the reference region 412 of the current coding block includes an upper reference region 412A having a first number of rows (e.g., 4 rows) and a left reference region 412C having a second number of columns (e.g., 8 columns). The first number is less than the second number. A plurality of weighting factors are determined based on a set of one or more reference samples 404R and / or 402R in the reference region 412 for generating the first chrominance sample 402C.
[0093] In some embodiments, the reference region 412 includes a subset of multiple rows of samples 804 located immediately above the current coding block. The multiple rows of samples 804 include a number of rows, and the number of rows is selected from a predefined set of integers (e.g., 1, 2, 3, 4, 5, 6).
[0094] Figure 9is a flowchart showing a video decoding method 900 according to some embodiments. The method 900 may be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, the method 900 is executed by executing instructions stored in the memory of the computing system (e.g., memory 314). In some embodiments, the method 900 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, in the MH-CCP mode, a linear or non-linear weighted sum of multiple types of co-located luma samples 404C ( Figure 4 ) is used to predict the chroma value 402C. Multiple types of co-located luma samples 404C are derived from the co-located luma samples 404C or filtered co-located luma samples using adjacent luma samples 404X (e.g., 404W, 404N, 404E, 404S, 404NW, 404NE, 404SW, 404SE) as filtering inputs. Each input to the weighted sum (e.g., corresponding to a respective type of co-located luma sample) is referred to as a hypothesis. In some cases, the value of each hypothesis is fed into a least mean square calculation kernel to derive the weight of the corresponding hypothesis used in MH-CCP. In some embodiments, the first luma sample 404C has eight adjacent luma samples 404X ( Figure 5 ) adjacent to the first luma sample 404C. Additionally, in some embodiments, when luma and chroma have different dimensions (e.g., 4:2:2 or 4:2:0), each luma sample 404 (e.g., 404C, 404X) includes a downsampled luma sample generated using a downsampling filter. Alternatively, each luma sample 404 includes the original co-located luma sample without any downsampling.
[0095] In one aspect of the present application, an adaptive MH-CCP mode is applied to support switching between different numbers (N) of model parameters. In one example, a two-parameter model 414A, a three-parameter model 414B, a four-parameter model 414C, or a five-parameter model 414D ( Figure 4)。In some embodiments, the selection for the number of model parameters is signaled explicitly in the video bitstream 116. In one example, the number of model parameters (N) is signaled in a high-level syntax, which includes but is not limited to sequence header, picture header, sub-picture header, slice header, tile header. In another example, the number of model parameters (N) is explicitly signaled at a block level, which includes but is not limited to maximum coding block level, coding block level, prediction block level, transform block level, predefined fixed block size level. In yet another example, for the current coding block 406A, the number of model parameters (N) is signaled both at a high-level syntax and at a block level, and the number of model parameters (N) is signaled at a block level rather than at a high-level syntax level.
[0096] In contrast, in some embodiments, the number of model parameters (N) is implicitly derived. For example, the number of model parameters (N) is implicitly derived based on coding information (including but not limited to frame resolution, quantization parameter, block size, block shape, luminance prediction mode) that is known to both the encoder and the decoder.
[0097] In some embodiments, different non-linear terms P are applied for different options of the number of model parameters (N). In some embodiments, different options of the same number (N) of model parameters correspond to different input positions (e.g., different adjacent luminance samples 404X).
[0098] In some embodiments, the two-parameter model 414A is applied in the MH-CCP mode, and for example, according to Equation (4), the first chrominance sample 402C is determined based on the weighted sum of the non-linear term P and the offset term B. In other words, the two-tap model corresponds to the two-parameter model 414A, and the two-tap model includes the non-linear term P and the offset term B. In some embodiments, the non-linear term B in the two-parameter model 414A is determined based on one of eight adjacent luminance samples 404X (e.g., 404N, 404W, 404E, 404S, 404NW, 404NE, 404SW, 404SE).
[0099] In some embodiments, a three-parameter model 414B is applied in the MH-CCP mode. The three-parameter model 414B has a variable number (e.g., 0, 1, 2, 3) of linear terms, a variable number (e.g., 0, 1, 2, 3) of non-linear terms P, and a variable number (e.g., 0, 1) of offset terms B. In some embodiments, the three-parameter model 414B includes a first luminance sample 404C, a non-linear term P, and an offset term B. The first chrominance sample 402C is determined by Equation (5.3). In some embodiments, the three-parameter model 414B includes one of eight adjacent luminance values 404X (e.g., having a luminance value X), a non-linear term P, and an offset term B. In some embodiments, the non-linear term P is determined based on the luminance value X. For example, the non-linear term P is equal to (XX + B) >> bit depth. In some embodiments, the non-linear term P is determined based on the first luminance sample 404C. In some embodiments, the linear terms correspond to one or more adjacent luminance samples 404X of the first luminance sample 404C, as shown, for example, in Equation (5.5).
[0100] In some embodiments, a four-parameter model 414C is applied in the MH-CCP mode. The four-parameter model 414C has a variable number (e.g., 0, 1, 2, 3, 4) of linear terms, a variable number (e.g., 0, 1, 2, 3, 4) of non-linear terms P, and a variable number (e.g., 0, 1) of offset terms B. In some embodiments, the four-parameter model 414 includes a first linear term of at least one of eight adjacent luminance samples 404X (having a luminance value X), a second linear term of the first luminance sample 404C co-located with the first chrominance sample 402C, a non-linear term P, and an offset term B. In one example, for instance, the left adjacent luminance sample 404W is applied according to Equation (6.1) to provide the first linear term. In some embodiments, a four-parameter model 414C is applied in the MH-CCP mode, wherein the first linear term or the offset term B is equal to the average Avg of a subset of adjacent luminance samples 404X of the first luminance sample 404C. The first chrominance sample 402C (predChromaVal) is predicted as: predChromaVal = w0C + w1Avg + w P P + w B B (7.1) predChromaVal = w0C + w1X + w P P + w B Avg (7.2)
[0101] In another aspect, the MH-CCP mode is applied to the largest coding block located in the current image frame 408 (e.g., Figure 8The current coded block 406A at the upper boundary of the current superblock 802). The number of rows of the upper reference region 412A corresponding to the current coded block 406A is less than the number of rows of the upper reference region 412B corresponding to a different coded block 406C that is not at the upper boundary of the current superblock 802. In other words, the current coded block 406A located at the upper boundary of the current superblock 802 uses less row buffer space. In some embodiments, the luminance samples 404 of the current coded block 406A are not downsampled. In another embodiment, the number of columns in the left reference region 412C is different from the number of rows of the upper reference region 412A or the upper reference region 412B. This column number is greater than this row number. In some embodiments, the number of rows in the upper reference region 412A is less than a threshold number (e.g., 4, 5), and the samples of the number of rows less than the threshold number in the upper reference region 412A are applied to determine a plurality of weighting factors.
[0102] In yet another aspect, for an alternative coded block 406B located at the upper boundary of the current superblock 802, the MH-CCP mode is automatically disabled.
[0103] Although Figure 7 A number of logical stages are shown in a particular order, but stages that are not order-dependent can be reordered, and other stages can be combined or broken apart. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0104] Some example embodiments are now described.
[0105] (A1)In some embodiments, method 900 is implemented for decoding video data. Method 900 includes: receiving (operation 902) a video bitstream that includes a current coded block of a current picture frame, wherein the video bitstream includes (operation 904) a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode; based on the first syntax element in the video bitstream, determining to enable (operation 906) the MH-CCP mode to reconstruct each chrominance sample using at least a corresponding luma sample co-located with each of a plurality of chrominance samples of the current coded block and one or more neighboring luma samples corresponding to the corresponding luma sample; determining (operation 908) at least for the current coded block a number of model parameters to be used in the MH-CCP mode; identifying (operation 910) one or more neighboring luma samples of a first luma sample based on the number of the model parameters; generating (operation 912) a first chrominance sample co-located with the first luma sample based on the first luma sample and the one or more neighboring luma samples; and reconstructing (operation 914) the current coded block including the first chrominance sample.
[0106] (A2)In some embodiments of A1, the number of the model parameters is selected from 2, 3, 4, and 5, and the number of the model parameters corresponds to a two-parameter model, a three-parameter model, a four-parameter model, and a five-parameter model, and the two-parameter model, the three-parameter model, the four-parameter model, and the five-parameter model are to be applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode.
[0107] (A3)In some embodiments of A1 or A2, the video bitstream further includes a second syntax element that indicates the number of the model parameters to be used in the MH-CCP mode at least for the current coded block.
[0108] (A4)In some embodiments of A3, the second syntax element is signaled as an advanced syntax element by one of a sequence header, a picture header, a sub-picture header, a slice header, and a tile header.
[0109] (A5)In some embodiments of A3, the second syntax element is signaled as a block-level syntax element at one of a maximum coded block level, a coded block level, a prediction block level, a transform block level, and a predefined fixed block size level.
[0110] (A6)In some embodiments of A1 or A2, instead of signaling the number of the model parameters for the current coding block, determining the number of the model parameters used in the MH-CCP mode for the current coding block at least further includes: determining the number of the model parameters based on coding information shared between an encoder and a decoder, where the coding information includes one or more of the following items: frame resolution of the current image frame, quantization parameter of the current coding block, block size, block shape, and luminance prediction mode.
[0111] (A7)In some embodiments of any one of A1 to A6, the number of the model parameters is selected from a plurality of predefined values, and generating the first chrominance sample includes: selecting one luminance sample from the one or more neighboring luminance samples and the first luminance sample based on the number of the model parameters, where at least two of the plurality of predefined values correspond to two different selected luminance samples; and generating a non-linear term based on the one luminance sample selected from the one or more neighboring luminance samples and the first luminance sample, where the first chrominance sample is generated based on the non-linear term.
[0112] (A8)In some embodiments of any one of A1 to A7, generating the first chrominance sample includes: when determining that the number of the model parameters is equal to a first value, generating a first non-linear term based on the first luminance sample and the first luminance sample among the one or more neighboring luminance samples; and when determining that the number of the model parameters is equal to a second value different from the first value, generating a second non-linear term based on the first luminance sample and a different second luminance sample among the one or more neighboring luminance samples.
[0113] (A9)In some embodiments of any one of A1 to A8, where the number of the model parameters corresponds to a plurality of hypothesis tap combinations, and identifying the one or more neighboring luminance samples of the first luminance sample further includes: selecting one hypothesis tap combination from the plurality of hypothesis tap combinations corresponding to the number of the model parameters; and identifying the one or more neighboring luminance samples of the first luminance sample based on the one hypothesis tap combination among the plurality of hypothesis tap combinations.
[0114] (A10)In some embodiments of any one of A1 to A9, where the number of the model parameters is equal to 2 and corresponds to a two-parameter model, the two-parameter model is applied to reconstruct the first chrominance sample of the current coding block in the MH-CCP mode, and the first chrominance sample is generated based on a weighted combination of a non-linear term P and an offset term B.
[0115] (A11)In some embodiments of A10, method 900 further includes: determining the non-linear term P based on at least one neighboring luminance sample from a set of eight neighboring luminance samples of the first luminance sample.
[0116] (A12) In some embodiments of any one of A1 to A9, the number of the model parameters is equal to 3 and corresponds to a three-parameter model, the three-parameter model is applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode, and the first chrominance sample is generated based on a weighted combination of a first number of linear terms, a second number of non-linear terms, and a third number of offset terms. The sum of the first number, the second number, and the third number is 3, and the first number, the second number, and the third number are respectively integers selected from the ranges [0, 3], [0, 3], and [0, 1].
[0117] (A13) In some embodiments of any one of A1 to A9, the number of the model parameters is equal to 3 and corresponds to a three-parameter model, the three-parameter model is applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode based on a weighted sum of the first luminance sample, the non-linear term, and the offset term, and generating the first chrominance sample further includes: determining the non-linear term based on eight adjacent luminance samples of the first luminance sample and at least one luminance sample among the first luminance sample.
[0118] (A14) In some embodiments of any one of A1 to A9, the number of the model parameters is equal to 3 and corresponds to a three-parameter model, the three-parameter model is applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode based on a weighted sum of the first adjacent luminance sample, the non-linear term, and the offset term, and generating the first chrominance sample further includes: determining the non-linear term based on at least one luminance sample among a set composed of the first luminance sample, the first adjacent luminance sample, and the remaining seven adjacent luminance samples of the first luminance sample.
[0119] (A15) In some embodiments of any one of A1 to A9, the number of the model parameters is equal to 4 and corresponds to a four-parameter model, the four-parameter model is applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode, and the first chrominance sample is generated based on a weighted combination of a first number of linear terms, a second number of non-linear terms, and a third number of offset terms. The sum of the first number, the second number, and the third number is 3, and the first number, the second number, and the third number are respectively integers selected from the ranges [0, 4], [0, 4], and [0, 1].
[0120] (A16)In some embodiments of any one of A1 to A9, the number of the model parameters is equal to 4 and corresponds to a four-parameter model, and the four-parameter model is applied to reconstruct the first chrominance sample of the current coding block based on a weighted sum of the first luminance sample, an adjacent luminance sample among the one or more adjacent luminance samples, a non-linear term, and an offset term in the MH-CCP mode.
[0121] (A17)In some embodiments of any one of A1 to A16, generating the first chrominance sample further includes: generating an offset term based on an average value of a subset of the one or more adjacent luminance samples of the first luminance sample, and the first chrominance sample is generated based on the offset term.
[0122] (A18)In some embodiments of any one of A1 to A17, the current coding block is included in a current superblock having a maximum coding block size in the current image frame. The method 900 further includes: when it is determined that the current coding block is adjacent to the upper boundary of the current superblock, applying a first upper reference region to determine a plurality of weighting factors for generating the first chrominance sample, the first upper reference region having a first number of rows and being located immediately above the current coding block; and when it is determined that the current coding block is not adjacent to the upper boundary of the current superblock, applying a second upper reference region to determine a plurality of weighting factors for generating the first chrominance sample, the second upper reference region having a second number of rows and being located immediately above the current coding block, the first number being greater than the second number.
[0123] (A19)In some embodiments of any one of A1 to A17, the current coding block is included in a current superblock having a maximum coding block size in the current image frame, and the first luminance sample and the one or more adjacent luminance samples have a resolution different from that of the first chrominance sample.
[0124] (A20)In some embodiments of any one of A1 to A17, the current coding block is included in a current superblock having a maximum coding block size in the current image frame, and the current superblock further includes an alternative coding block different from the current coding block. The method 900 further includes: when it is determined that the alternative coding block is adjacent to the upper boundary of the current superblock, aborting the MH-CCP mode for the alternative coding block.
[0125] (A21)In some embodiments of any one of A1 to A17, the method 900 further includes: identifying a reference region of the current coded block, the reference region including an upper reference region having a first number of rows and a left reference region having a second number of columns, wherein the first number is less than the second number; and determining the plurality of weighting factors based on a set of one or more reference samples in the reference region for generating the first chrominance sample.
[0126] (A22)In some embodiments of any one of A1 to A17, the method 900 further includes: identifying a reference region, the reference region including a subset of a plurality of rows of samples immediately above the current coded block, wherein the plurality of rows of samples includes a number of rows, and the number of rows is selected from a predefined set of integers.
[0127] (A23)A computing system includes: a control circuit; and a memory storing one or more programs configured to be executed by the control circuit, the one or more programs further including instructions for: receiving video data including a current coded block of a current image frame; encoding the current image frame according to intra prediction; determining to enable a multi-hypothesis cross-component prediction (MH-CCP) mode to determine each chrominance sample based on a corresponding luma sample co-located with each chrominance sample of the current coded block and one or more adjacent luma samples corresponding to the corresponding luma sample, wherein the MH-CCP mode is associated with a number of model parameters for identifying one or more adjacent luma samples of a first luma sample; transmitting the encoded current image frame via a video bitstream; and signaling a first syntax element via the video bitstream to indicate application of the MH-CCP mode to reconstruct a first chrominance sample co-located with the first luma sample based at least on the first luma sample and the one or more adjacent luma samples.
[0128] (A24)A non-transitory computer-readable storage medium stores one or more programs executed by a control circuit of a computing system. The one or more programs include instructions for: obtaining a source video sequence including a current encoded block of a current image frame; and performing a transformation between the source video sequence and a video bitstream. The video bitstream includes: the current encoded block of the current image frame; and a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct each chrominance sample based at least on a corresponding luma sample co-located with each chrominance sample of the current encoded block and one or more neighboring luma samples corresponding to the corresponding luma sample. The MH-CCP mode is associated with a number of model parameters for identifying one or more neighboring luma samples of a first luma sample, and the number of model parameters is applied to reconstruct a first chrominance sample co-located with the first luma sample based at least on the first luma sample and the one or more neighboring luma samples.
[0129] (A25)A method for decoding video data includes: receiving a video bitstream including a current encoded block of a current image frame. The video bitstream includes a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode. Based on the first syntax element in the video bitstream, determine to enable the MH-CCP mode to reconstruct each sample of a second color component of the current encoded block based at least on a corresponding sample of a first color component co-located with each sample of the second color component of the current encoded block and one or more neighboring samples of the second color component corresponding to the corresponding sample of the first color component. At least for the current encoded block, determine a number of model parameters used in the MH-CCP mode. Identify one or more neighboring samples of a first sample of the first color component based on the number of model parameters. Generate a first sample of the second color component co-located with the first sample of the first color component based on the first sample of the first color component and the one or more neighboring samples of the first color component. And reconstruct the current encoded block including the first sample of the second color component.
[0130] In another aspect, some embodiments include a computing system (e.g., server system 112) including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit. The memory stores one or more instruction sets configured to be executed by the control circuit, the one or more instruction sets including instructions for performing any method described herein (e.g., A1 to A25 above).
[0131] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets for execution by a control circuit of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A25 above).
[0132] Unless otherwise specified, any syntax elements described herein may be high-level syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements may be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, picture header, tile header, and / or coding tree unit (CTU) header.
[0133] It will be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the listed features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0134] As used herein, depending on context, the term "if" can be interpreted to mean "when" or "in the case where" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the described precondition is true. Similarly, depending on context, the phrase "if it is determined [that the described precondition is true]" or "if [the described precondition is true]" or "when [the described precondition is true]" can be interpreted to mean "upon determining" or "in response to determining" or "in accordance with determining" or "upon detecting" or "in response to detecting" that the described precondition is true.
[0135] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments selected and described were chosen in order to best explain the operating principles and practical applications, thereby enabling others skilled in the art to practice.
Claims
1. A method for decoding video data, comprising: receiving a video bitstream, the video bitstream including a current coded block of a current picture frame, wherein the video bitstream includes a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode; determining, based on the first syntax element in the video bitstream, to enable the MH-CCP mode to reconstruct each of the chrominance samples by using at least a corresponding luma sample co-located with each of a plurality of chrominance samples of the current coded block and one or more neighboring luma samples corresponding to the corresponding luma sample; determining, at least for the current coded block, a number of model parameters to be used in the MH-CCP mode; identifying, based on the number of model parameters, one or more neighboring luma samples of a first luma sample; generating, based on the first luma sample and the one or more neighboring luma samples, a first chrominance sample co-located with the first luma sample; and reconstructing the current coded block including the first chrominance sample.
2. The method according to claim 1, wherein The number of model parameters is selected from 2, 3, 4, and 5, and the number of model parameters corresponds to a two-parameter model, a three-parameter model, a four-parameter model, and a five-parameter model, and the two-parameter model, the three-parameter model, the four-parameter model, and the five-parameter model will be applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode.
3. The method according to claim 1, wherein, The video bitstream further includes a second syntax element that indicates the number of model parameters to be used in the MH-CCP mode at least for the current coded block.
4. The method according to claim 3, wherein The second syntax element is signaled as an advanced syntax element by one of a sequence header, a picture header, a sub-picture header, a slice header, and a tile header.
5. The method according to claim 3, wherein The second syntax element is signaled as a block-level syntax element at one of a maximum coded block level, a coded block level, a prediction block level, a transform block level, and a predefined fixed block size level.
6. The method according to claim 1, wherein Without signaling the number of model parameters for the current coded block, determining, at least for the current coded block, the number of model parameters to be used in the MH-CCP mode further includes: determining the number of model parameters based on coding information shared between an encoder and a decoder, the coding information including one or more of the following items: a frame resolution of the current picture frame, a quantization parameter of the current coded block, a block size, a block shape, and a luma prediction mode.
7. The method according to claim 1, wherein, The number of model parameters is selected from a plurality of predefined values, and generating the first chrominance sample includes: selecting, based on the number of model parameters, one luma sample from the one or more neighboring luma samples and the first luma sample, at least two of the plurality of predefined values corresponding to two different selected luma samples; and generating a non-linear term based on the one luma sample selected from the one or more neighboring luma samples and the first luma sample, wherein the first chrominance sample is generated based on the non-linear term.
8. The method according to claim 1, wherein Generating the first chrominance sample includes: When it is determined that the number of the model parameters is equal to a first value, generate a first non - linear term based on the first luminance sample and the first luminance sample among the one or more adjacent luminance samples; and When it is determined that the number of the model parameters is equal to a second value different from the first value, generate a second non - linear term based on the first luminance sample and a different second luminance sample among the one or more adjacent luminance samples.
9. The method according to claim 1, wherein The number of the model parameters corresponds to a plurality of hypothesized tap combinations, and identifying the one or more adjacent luminance samples of the first luminance sample further includes: Selecting one hypothesized tap combination from the plurality of hypothesized tap combinations corresponding to the number of the model parameters; and Identifying the one or more adjacent luminance samples of the first luminance sample based on the one hypothesized tap combination among the plurality of hypothesized tap combinations.
10. The method according to claim 1, wherein, The number of the model parameters is equal to 2 and corresponds to a two - parameter model, the two - parameter model is applied to reconstruct the first chrominance sample of the current coding block in the MH - CCP mode, and the first chrominance sample is generated based on a weighted combination of a non - linear term P and an offset term B.
11. The method according to claim 10, the method further includes: Determining the non - linear term P based on at least one adjacent luminance sample from a set of eight adjacent luminance samples of the first luminance sample.
12. The method according to claim 1, wherein, The number of the model parameters is equal to 3 and corresponds to a three - parameter model, the three - parameter model is applied to reconstruct the first chrominance sample of the current coding block in the MH - CCP mode, and the first chrominance sample is generated based on a weighted combination of a first number of linear terms, a second number of non - linear terms, and a third number of offset terms, wherein the sum of the first number, the second number, and the third number is 3, and the first number, the second number, and the third number are integers selected from the ranges [0, 3], [0, 3], and [0, 1] respectively.
13. The method according to claim 1, wherein, The number of the model parameters is equal to 3 and corresponds to a three - parameter model, the three - parameter model is applied to reconstruct the first chrominance sample of the current coding block in the MH - CCP mode based on a weighted sum of the first luminance sample, the non - linear term, and the offset term, and generating the first chrominance sample further includes: Determining the non - linear term based on eight adjacent luminance samples of the first luminance sample and at least one luminance sample of the first luminance sample.
14. The method according to claim 1, wherein The number of the model parameters is equal to 3 and corresponds to a three - parameter model, the three - parameter model is applied to reconstruct the first chrominance sample of the current coding block in the MH - CCP mode based on a weighted sum of a first adjacent luminance sample, the non - linear term, and the offset term, and generating the first chrominance sample further includes: Determining the non - linear term based on at least one luminance sample among the first luminance sample, the first adjacent luminance sample, and a set of the remaining seven adjacent luminance samples of the first luminance sample.
15. The method according to claim 1, wherein The number of the model parameters is equal to 4 and corresponds to a four-parameter model, the four-parameter model is applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode, and the first chrominance sample is generated based on a weighted combination of a first number of linear terms, a second number of non-linear terms, and a third number of offset terms, wherein the sum of the first number, the second number, and the third number is 3, and the first number, the second number, and the third number are integers selected from the ranges [0, 4], [0, 4], and [0, 1], respectively.
16. The method according to claim 1, wherein, The number of the model parameters is equal to 4 and corresponds to a four-parameter model, the four-parameter model is applied to reconstruct the first chrominance sample of the current coded block in the MH-CCP mode based on a weighted sum of the first luminance sample, one of the one or more adjacent luminance samples, non-linear terms, and offset terms.
17. The method according to claim 1, wherein Generating the first chrominance sample further includes: generating an offset term based on an average value of a subset of the one or more adjacent luminance samples of the first luminance sample, and the first chrominance sample is generated based on the offset term.
18. The method according to claim 1, wherein The current coded block is included in a current superblock, the current superblock has a maximum coded block size in the current image frame, and the method further includes: when it is determined that the current coded block is adjacent to the upper boundary of the current superblock, applying a first upper reference region to determine a plurality of weighting factors for generating the first chrominance sample, the first upper reference region having a first number of rows and being located immediately above the current coded block; and when it is determined that the current coded block is not adjacent to the upper boundary of the current superblock, applying a second upper reference region to determine the plurality of weighting factors for generating the first chrominance sample, the second upper reference region having a second number of rows and being located immediately above the current coded block, wherein the first number is greater than the second number.
19. The method according to claim 1, wherein, The current coded block is included in a current superblock, the current superblock has a maximum coded block size in the current image frame, and the first luminance sample and the one or more adjacent luminance samples have a resolution different from the resolution of the first chrominance sample.
20. The method according to claim 1, wherein, The current coded block is included in a current superblock, the current superblock has a maximum coded block size in the current image frame, and the current superblock further includes an alternative coded block different from the current coded block, and the method further includes: when it is determined that the alternative coded block is adjacent to the upper boundary of the current superblock, aborting the MH-CCP mode for the alternative coded block.
21. The method according to claim 1, the method further includes: identifying a reference region of the current coded block, the reference region including an upper reference region having a first number of rows and a left reference region having a second number of columns, wherein the first number is less than the second number; and Determine the plurality of weighting factors based on a set of one or more reference samples in the reference region for generating the first chrominance sample.
22. The method according to claim 1, the method further comprising: Identifying a reference region including a subset of a plurality of rows of samples immediately above the current coding block, wherein the plurality of rows of samples includes a number of rows, and the number of rows is selected from a predefined set of integers.
23. A computing system, comprising: A control circuit; And A memory storing one or more programs configured to be executed by the control circuit, the one or more programs further including instructions for: Receiving video data including a current coding block of a current image frame; Encoding the current image frame according to intra prediction; Determining to enable a multi-hypothesis cross-component prediction (MH-CCP) mode to determine each chrominance sample based on a corresponding luma sample co-located with each chrominance sample of the current coding block and one or more adjacent luma samples corresponding to the corresponding luma sample, wherein the MH-CCP mode is associated with a number of model parameters for identifying one or more adjacent luma samples of a first luma sample; Transmitting the encoded current image frame via a video bitstream; and Signaling, via the video bitstream, a first syntax element indicating application of the MH-CCP mode to reconstruct a first chrominance sample co-located with the first luma sample based at least on the first luma sample and the one or more adjacent luma samples.
24. A non-transitory computer-readable storage medium storing one or more programs executed by a control circuit of a computing system, the one or more programs including instructions for: Obtain a source video sequence, where the source video sequence includes a current coding block of a current image frame; And Performing a conversion between the source video sequence and a video bitstream, wherein The video bitstream includes: The current coding block of the current image frame; and A first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct each chrominance sample based at least on a corresponding luma sample co-located with each chrominance sample of the current coding block and one or more adjacent luma samples corresponding to the corresponding luma sample; Wherein the MH-CCP mode is associated with a number of model parameters for identifying one or more adjacent luma samples of a first luma sample, and the number of model parameters is applied to reconstruct a first chrominance sample co-located with the first luma sample based at least on the first luma sample and the one or more adjacent luma samples.