Multi-hypothesis cross-component prediction differentiation model
By adopting the multi-assumption cross-component prediction (MH-CCP) model and five-tap model in video encoding technology, the problems of video quality degradation and low resource utilization efficiency in the prior art are solved, and efficient video data compression and quality maintenance are achieved.
Patent Information
- Application Number
- CN202480004775.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2024-05-21
- Publication Date
- 2025-06-20
AI Technical Summary
When processing video data, it is difficult for existing video encoding technologies to effectively utilize the advantages of multi-assumption cross-component prediction, resulting in video quality degradation during compression and low bandwidth and storage space utilization efficiency.
Multi-assumption Cross-Component Prediction (MH-CCP) mode is adopted to predict chromaticity samples using the weighted sum of multiple brightness samples in intra prediction through a five-tap model, thereby achieving efficient compression of video data.
It improves the compression efficiency of video data, reduces the bandwidth and storage space requirements, and maintains video quality, and can provide better video encoding performance under limited resource conditions.
Smart Images

Figure CN120188482A_ABST
Abstract
Description
Cross-reference
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 588,920, filed on October 9, 2023, entitled "Multi-Hypothesis Cross Component Prediction Different Model", and this application is a continuation-in-part of and claims priority to U.S. Patent Application No. 18 / 622,836, filed on March 29, 2024, entitled "Multi-Hypothesis Cross Component Prediction Different Model". Technical Field
[0002] Embodiments of the present disclosure generally relate to video coding and decoding, including but not limited to systems and methods for processing video data using multi-hypothesis cross-component prediction (MH-CCP). Background Art
[0003] A variety of electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. These electronic devices send and receive digital video data over a communication network or otherwise transmit digital video data, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video coding can be used to compress video data according to one or more video coding standards before the video data is transmitted or stored. Video coding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.
[0004] Video coding is typically performed using prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficency Video Coding (HEVC / H.265) is a video compression standard designed as part of the Moving Picture Experts Group-High Efficiency (MPEG-H) project. The International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4), respectively. Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2), respectively. The Alliance for Open Media (AOMedia) Video (AV1) is an open video coding format designed as an alternative to HEVC. A verified version 1.0.0 containing errata 1 was released on January 8, 2019. Summary of the Invention
[0005] As described above, encoding (compression) reduces the need for bandwidth and / or storage space. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal through a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not completely retained during the encoding process and cannot be fully restored during the decoding process. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal can be used for the intended application. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications can tolerate higher distortion compared to users of movie or television broadcast applications. The compression ratio achieved by a particular encoding algorithm can be selected or adjusted to reflect the various distortion tolerances: higher tolerable distortion generally allows for encoding and decoding algorithms with higher losses and higher compression ratios.
[0006] The present disclosure describes a video compression method using intra prediction. A linear or non-linear weighted sum of multiple types of luma samples is used to predict chroma samples, for example, in multi-hypothesis cross-component prediction (MH-CCP). The multiple types of luma samples include luma sample C co-located with the chroma sample and filtered luma samples determined based on adjacent luma samples and used as filter inputs. Each filter input of the weighted sum is referred to as a hypothesis. Each hypothesis is associated with a weighting factor in MH-CCP. In one aspect of the present application, the weighting factors are applied to generate a linear or non-linear weighted sum of different types of luma samples. These weighting factors are determined for each coding block based on the reference region of the coding block. In some embodiments, these weighting factors are determined by applying a least mean square calculation kernel to process the reconstructed samples of the reference block of each coding block.
[0007] In other words, in some embodiments, samples of a second color component are predicted as a linear or non-linear weighted sum of samples of the second color component co-located with the samples of the second color component and one or more associated neighboring luminance samples according to a multi-tap model associated with MH-CCP. The multi-tap model includes multi (N) taps selected from co-located samples of the second color component, one or more associated neighboring luminance samples of the second color component, non-linear terms, and offset terms. The multi-tap model corresponds to the same number (N) of selected items that are combined to determine samples of the second color component. In some embodiments, samples of a first color component are one of a chrominance sample and a luminance sample, and samples of the second color component are luminance samples. For example, the chrominance sample is a weighted combination of terms selected from corresponding co-located luminance samples, one or more neighboring luminance samples, non-linear terms, and offset terms. Alternatively, in some embodiments, the first color component is one of red, green, and blue, and the second color component is another of red, green, and blue. Alternatively, in some embodiments, the first color component and the second component correspond to a color format different from the YCbCr color format and the RGB color format.
[0008] According to some embodiments, a method of video decoding is provided. The method includes: receiving a video bitstream that includes a current encoded block of a current image frame. The video bitstream includes a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode. The method further includes: determining, based on the first syntax element in the video bitstream, to enable the MH-CCP mode to reconstruct a chrominance sample based at least on a luminance sample co-located with each of a plurality of chrominance samples of the current encoded block and one or more neighboring luminance samples corresponding to the luminance sample. The method further includes: identifying a five-tap model configured to determine a chrominance sample of the current encoded block in the MH-CCP mode; identifying, based on the five-tap model, a pair of neighboring luminance samples of a first luminance sample; generating a first chrominance sample co-located with the first luminance sample based at least on the first luminance sample and the pair of neighboring luminance samples; and reconstructing the current encoded block including the first chrominance sample.
[0009] According to some embodiments, a video encoding method is provided. The method includes: receiving video data including a current coding block of a current picture frame; and encoding the current picture frame according to intra prediction. The method further includes: determining to enable a multi-hypothesis cross-component prediction (MH-CCP) mode to determine a chrominance sample based on at least a luminance sample co-located with each of a plurality of chrominance samples of the current coding block and a sum of one or more neighboring luminance samples corresponding to the luminance sample. The MH-CCP mode is associated with a five-tap model for identifying a pair of neighboring luminance samples of a first luminance sample. The method further includes: transmitting the encoded current picture frame via a video bitstream; and signaling, via the video bitstream, a first syntax element to indicate application of the MH-CCP mode to reconstruct a first chrominance sample co-located with the first luminance sample based on at least the first luminance sample and the pair of neighboring luminance samples.
[0010] According to some embodiments, a bitstream conversion method is provided. The method includes: obtaining a source video sequence including a current coding block of a current picture frame; and converting between the source video sequence and a video bitstream. The video bitstream includes a current coding block of the current picture frame and a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode. The first syntax element for the MH-CCP mode indicates whether to reconstruct the chrominance sample based on at least a luminance sample co-located with each of a plurality of chrominance samples of the current coding block and a sum of one or more neighboring luminance samples corresponding to the luminance sample. The MH-CCP mode is associated with a five-tap model for identifying a pair of neighboring luminance samples of a first luminance sample, and the MH-CCP mode is applied to reconstruct a first chrominance sample co-located with the first luminance sample based on at least the first luminance sample and the pair of neighboring luminance samples.
[0011] According to some embodiments, a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).
[0012] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets executable by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0013] Accordingly, devices and systems are disclosed that use methods for encoding and decoding video. Such methods, devices, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding. The features and advantages described in this specification need not be all-inclusive, and in particular, given the figures, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Further, it should be noted that the language used in the specification has been selected primarily for readability and instructional purposes and not necessarily to delimit or circumscribe the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To enable a more detailed understanding of the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings only illustrate the relevant features of the present disclosure, and these features need not be considered restrictive, as those skilled in the art will understand upon reading this disclosure that the description may admit other effective features.
[0015] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0016] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0017] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0018] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0019] Figure 4A shows an example scheme for generating a first chrominance sample from a plurality of luminance samples according to some embodiments.
[0020] Figure 4B and Figure 4C is a schematic diagram of two hypothesized tap combinations according to some embodiments, where each hypothesized tap combination includes two adjacent luminance samples of a first luminance sample.
[0021] Figure 5 shows another example scheme for generating a first chrominance sample from a plurality of luminance samples according to some embodiments.
[0022] Figure 6A 、 Figure 6B and Figure 6C are three examples of block size group mapping tables according to some embodiments.
[0023] Figure 7is a flowchart showing another example method of video coding according to some embodiments.
[0024] By convention, the various features shown in the drawings need not be drawn to scale, and like reference numerals may be used throughout the specification and drawings to denote like features. Detailed Description
[0025] The present disclosure describes cross-component intra prediction of video data in the MH-CCP mode, where each sample among a plurality of samples of a first color component is determined based on one or more associated samples of a second color component. The MH-CCP mode corresponds to a multi-tap model including multiple (N) taps. Each tap is selected from a co-located sample of the second color component, one or more associated adjacent luma samples of the second color component, a non-linear term, and an offset term. The selected taps are combined in a weighted manner to determine a sample of the second color component. In some embodiments, the sample of the first color component is one of a chroma sample and a luma sample, and the sample of the second color component is a luma sample. For example, the chroma sample is a weighted combination of terms selected from a corresponding co-located luma sample, one or more adjacent luma samples, a non-linear term, and an offset term. In some embodiments, the number of taps may be selected from 2, 3, 4, and 5. In one aspect of the present application, a computing device receives a video bitstream including a current encoded block of a current image frame and a first syntax element for the MH-CCP mode, and the computing device identifies a five-tap model configured to determine a chroma sample of the current encoded block in the MH-CCP mode. A pair of adjacent luma samples of a first luma sample is determined based on the five-tap model, and the first luma sample, the non-linear term, and the offset term are used to generate a first chroma sample co-located with the first luma sample.
[0026] Figure 1 is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), and the source device 102 and the plurality of electronic devices 120 are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for use with video-enabled applications such as video conferencing applications, digital television (TV) applications, media storage, and / or distribution applications.
[0027] The source device 102 includes a video source 104 (e.g., a camera assembly or a media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may have a higher data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less transmission bandwidth and less storage space for storage. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to send uncompressed video to one or more networks 110).
[0028] The one or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. The one or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks may include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0029] The one or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode the encoded video bitstreams 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or perform encoding based on the encoded video bitstreams 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to crop the encoded video bitstreams 108 for customizing potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0030] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to produce an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices in the electronic device 120 do not include a display component (e.g., the electronic device is communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0031] The source device and / or the plurality of electronic devices 120 may also be referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0032] In an example operation of the communication system 100, the source device 102 sends the encoded video stream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using the codec component 114. For example, the server system 112 may encode video data that is more suitable for network transmission and / or storage. The server system 112 may send the encoded video data 116 (e.g., one or more encoded video streams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0033] Figure 2Ais a block diagram showing example elements of an encoder assembly 106 according to some embodiments. The encoder assembly 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder assembly includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder assembly 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a different device than the encoder assembly 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In some embodiments, the video source 104 may be a storage device storing previously acquired / prepared video. In some embodiments, the video source 104 is a camera that acquires local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can easily understand the relationship between pixels and samples.
[0034] The encoding component 106 is configured to encode and / or compress pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder assembly 106 is configured to convert between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques, etc.), picture size, group of picture (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can easily find other functions of the controller 204, as these functions may be related to the encoder assembly 106 optimized for a specific system design.
[0035] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder creates sample data (when the compression between the symbol stream and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exactly corresponding between the local encoder and the remote encoder. Thus, the prediction part of the encoder interprets the reference picture samples as exactly the same sample values as the decoder will interpret when using the prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those of ordinary skill in the art.
[0036] The operation of the decoder 210 can be the same as that of a remote decoder such as the decoder component 122 described in detail below, for example. However, briefly referring to Figure 2B When the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210. Figure 2B
[0037] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of the encoder techniques can be simplified because the encoder techniques are inverse to the decoder techniques.
[0038] As part of this operation, the source encoder 202 can perform motion-compensated predictive coding. Referring to one or more previously encoded frames in the video sequence designated as reference frames, this motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 212 encodes the difference between a pixel block of the input frame and pixel blocks of one or more reference frames, which can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.
[0039] The decoder 210 decodes the encoded video data of a frame that can be specified as a reference frame based on the symbols created by the source encoder 202. Advantageously, the operation of the encoding engine 212 can be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on the reference frame and can cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame that has the same content (without transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0040] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as a candidate reference pixel block) or some metadata, such as a reference picture motion vector, block shape, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 206 can operate block by block on the sample blocks to find a suitable prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.
[0041] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.
[0042] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission over a communication channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / signal noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Video Usability Information (VUI) parameter set fragments, and the like.
[0043] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, pictures may generally be classified as: intra pictures (I pictures), predictive pictures (P pictures), or bi-predictive pictures (B pictures). Intra pictures may be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics, and thus these variants and their corresponding applications and characteristics are not repeated herein. Predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0044] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the block can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded with reference to a previously encoded reference picture, either through spatial prediction or through temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures, either through spatial prediction or through temporal prediction.
[0045] The video captured can be a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In an example, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded with a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0046] The encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard described herein, for example. In operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.
[0047] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The illustrated decoder component 122 is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to a loop filter 256 and is configured to send data to the display 124 (e.g., via a wired or wireless connection).
[0048] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, each encoded video sequence is decoded independently of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective consuming entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0049] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.
[0050] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 inside decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is also provided outside decoder component 122 (e.g., to prevent network jitter). While it may not be necessary to configure buffer memory 252, or the buffer memory may be made smaller, when receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network. For use on a best-effort network such as the Internet, buffer memory 252 may be required, which may be relatively large and / or have an adaptive size and may be implemented at least partially in the operating system or a similar element outside decoder component 122.
[0051] The parser 254 is configured to reconstruct symbols 270 from an encoded video sequence. The symbols may include, for example, information for managing the operation of the decoder components 122 and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of supplementary enhancement information (SEI) messages or parameter set fragments (not labeled) of the VUI. The parser 254 may parse (entropy decode) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), and the like. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and the like.
[0052] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., an inter-picture and an intra-picture, an inter-block and an intra-block) and other factors, the reconstruction of the symbol 270 may involve multiple different units. Which units are involved and the way these units are involved may be controlled by subgroup control information parsed by the parser 254 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 254 and the multiple units below are not described.
[0053] The decoder components 122 may be conceptually subdivided into multiple functional units, and in some embodiments, these functional units interact closely with each other and may be at least partially integrated with each other. However, for clarity, the conceptually subdivided functional units are retained herein.
[0054] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as symbols 270 and control information (e.g., which transform mode to use, block size, quantization factor, and / or quantization scaling matrix, etc.). The scaler / inverse transform unit 258 may output a block including sample values, which may be input into the aggregator 268.
[0055] In some cases, the output samples of the scaler / inverse transform unit 258 may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra picture prediction unit 262. The intra picture prediction unit 262 may generate a block having the same size and shape as the block being reconstructed using surrounding reconstructed information extracted from the current (partially reconstructed) picture in the current picture buffer 264. The aggregator 268 may add the prediction information generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a per-sample basis.
[0056] In other cases, the output samples of the scaler / inverse transform unit 258 may belong to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After motion-compensating the extracted samples according to the sign 270 belonging to the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case), thereby generating output sample information. The address from which the motion compensation prediction unit 260 obtains prediction samples from within the reference picture memory 266 may be controlled by a motion vector. The motion vector is available to the motion compensation prediction unit 260 in the form of the sign 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory 266, a motion vector prediction mechanism, etc., when using sub-sample accurate motion vectors.
[0057] The output samples of the aggregator 268 may be employed by various loop filtering techniques in the loop filter unit 256. The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters may be available to the loop filter unit 256 as the sign 270 from the parser 254. However, the video compression technique may also respond to meta-information obtained during decoding of a previously (in decoding order) part of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device such as the display 124 and stored in the reference picture memory 266 for subsequent inter picture prediction.
[0058] Once reconstructed, certain coded pictures can be used as reference pictures for subsequent prediction. Once a coded picture has been fully reconstructed and the coded picture (by, for example, parser 254) is identified as a reference picture, the current reference picture can be made part of the reference picture memory 266 and a new current picture memory can be reallocated before starting to reconstruct subsequent coded pictures.
[0059] The decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard (such as any of the standards described herein). The coded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the coded video sequence follows the syntax specified in the video compression technique document or standard (specifically, the syntax of the video compression technique or standard specified in the profile document thereof). Additionally, in order to conform to some video compression techniques or standards, the complexity of the coded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the coded video sequence.
[0060] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a central processing unit (CPU), a Graphics Processing Unit (GPU), and / or a Data Processing Unit (DPU)). In some embodiments, the control circuit includes one or more Field-Programmable Gate Arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).
[0061] The network interface 304 can be configured to connect to one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network can be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a low-latency network, etc. Examples of communication networks include local area networks such as Ethernet, wireless local area networks (LANs), cellular networks including the global system for mobile (GSM), 3rd generation (3G), 4th generation (4G), 5th generation (5G), long term evolution (LTE), etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including the controller area network bus (CANbus), and so on. Such communication can be only one-way reception (e.g., broadcast television), only one-way transmission (e.g., CANbus connected to certain CANbus devices), or two-way (e.g., connecting to other computer systems using a local area network or a wide area digital network). Such communication can include communication to one or more cloud computing networks.
[0062] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.
[0063] Memory 314 may include high-speed random access memory (such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDRRAM), and / or other random access solid-state memory devices), and / or non-volatile memory (such as one or more disk storage devices, optical disc storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices remote from control circuit 302. Memory 314 or, optionally, the non-volatile solid-state memory device within memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, memory 314 or the non-transitory computer-readable storage medium of memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● Operating system 316, including procedures for handling various basic system services and for performing hardware-related tasks; ● Network communication module 318, for connecting server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · Codec module 320, for performing various functions related to encoding and / or decoding data (such as video data). In some embodiments, codec module 320 is an instance of codec component 114. Codec module 320 includes, but is not limited to, one or more of the following modules: o Decoding module 322, for performing various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and o Encoding module 340, for performing various functions related to encoding data, such as those previously described with respect to encoder component 106; and · Picture memory 352, for storing pictures and picture data, e.g., for use with codec module 320. In some embodiments, picture memory 352 includes one or more of the following memories: reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.
[0064] In some embodiments, the decoding module 322 includes: a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform the various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).
[0065] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform the various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, the shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0066] Each of the aforementioned modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The aforementioned modules (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, optionally, the encoding module 320 does not include separate decoding and encoding modules, but rather uses the same set of modules to perform the two sets of functions. In some embodiments, the memory 314 stores a subset of the aforementioned modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0067] Although Figure 3 a server system 112 is shown in accordance with some embodiments, however, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems rather than as a structural schematic of the embodiments described herein. In practice, and as would be recognized by one of ordinary skill in the art, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the items shown separately may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among these servers will vary depending on the implementation, and optionally, will depend in part on the data traffic processed by the server system during peak usage periods as well as during average usage periods.
[0068] Figure 4A FIG. 400 shows an example scenario 400 for generating a first chroma sample 402C from a plurality of luminance samples 404 (e.g., 404C and 404X) in the MH-CCP mode according to some embodiments. Figure 4B and Figure 4C FIGS. 416A and 416B are schematic diagrams of two example hypothesis tap combinations 416A and 416B according to some embodiments, where each hypothesis tap combination includes a pair of adjacent luminance samples 404X of a first luminance sample 402C. In some embodiments, the current encoded block 406A of the current image frame 408 is encoded in the Cross-Component Intra Prediction (CCIP) mode. In the CCIP mode, the decoder 122 ( Figure 2B ) determines each chroma sample of the plurality of chroma samples 402 of the current encoded block 406A based on one or more reconstructed luminance samples 404. In some cases, the CCIP mode includes the Cross-Component Linear model Mode (CCLM), in which the first chroma sample 402C is transformed from the reconstructed luminance sample 404C co-located with the chroma sample based on a linear model. Alternatively, in some cases, the CCIP mode includes the Convolutional Cross-Component Mode (CCCM), in which the first chroma sample 402C is directly predicted from a plurality of reconstructed luminance samples 404X near the first luminance sample 404C based on the filter shape of the filter. Alternatively and additionally, in some cases, the CCIP mode includes the MH-CCP mode, in which the first chroma sample 402C is generated by combining at least the first luminance sample 404C co-located with the first chroma sample 402C and a plurality of hypothesis values 410 using a plurality of weighting factors. A plurality of adjacent luminance samples 404X of the first luminance sample 404C are combined using a plurality of coefficients to generate a plurality of hypothesis values 410.
[0069] In some embodiments, the video bitstream 116 includes a first syntax element 420 for the MH-CCP mode. The first chroma sample 402C of the current encoded block 406A is configured to be generated by combining at least the first luminance sample 404C co-located with the first chroma sample 402C and one or more adjacent luminance samples 404X of the first luminance sample 404C using a plurality of weighting factors (e.g., c0 to c6). When it is determined that the MH-CCP mode is enabled, the first chroma sample 402C is predicted according to the following equation: predChromaVal = c0C + c1N1 + c2N2 + c3P + c4B (1) Wherein, predChromaVal is the predicted chroma value of the first chroma sample 402C; C is the luma value of the first luma sample 404C co-located with the first chroma sample 402C; N1 and N2 are two adjacent luma samples 404X of the first luma sample 404C; P is a non-linear term, for example, equal to (C×C + B) >> bit depth; B is an offset; c0-c4 are weighting factors. In some embodiments, B is the median luma value or the average luma value of the luma samples 404 of the current coding block 406A. Equation (1) includes five terms and represents a five-tap model, which is used to determine the first chroma sample 402C of the current coding block 406A in the MH-CCP mode.
[0070] In some embodiments, each of one or more adjacent luma samples 404X of the first luma sample 404C is adjacent to the first luma sample 404C and shares at least one corresponding side or vertex with the first luma sample 404C. In some embodiments, the one or more adjacent luma samples 404X include a subset or all of the following adjacent luma samples: a north adjacent luma sample (also referred to as an upper luma sample) 404N, a south adjacent luma sample (also referred to as a lower luma sample) 404S, a west adjacent luma sample (also referred to as a left luma sample) 404W, an east adjacent luma sample (also referred to as a right luma sample) 404E, a northwest adjacent luma sample (also referred to as an upper left luma sample) 404NW, a southeast adjacent luma sample (also referred to as a lower right luma sample) 404SE, a southwest adjacent luma sample (also referred to as a lower left luma sample) 404SW, and a northeast adjacent luma sample (also referred to as an upper right luma sample) 404NE.
[0071] In some embodiments, the luma samples 404 and chroma samples 402 of the current coding block have different resolutions corresponding to a chroma subsampling scheme (e.g., 4:2:2 or 4:2:0). Each luma sample 404 includes a downsampled luma sample generated from the reconstructed luma samples using a downsampling filter. Alternatively, in some embodiments, each luma sample 404 includes an original sample or a reconstructed luma sample without any downsampling. That is, the first luma sample 404C is reconstructed or downsampled to the resolution of the chroma samples according to the resolution of the luma samples. Therefore, the adjacent luma samples 404X (e.g., 404N, 404W, 404E, 404S, 404NW, 404NE, 404SW, 404SE) are also reconstructed or downsampled to the resolution of the chroma samples according to the resolution of the luma samples.
[0072] In some embodiments, based on a five-tap model (e.g., the five-tap model in Equation (1)), the first luminance sample 404C, two adjacent luminance samples 404X, a non-linear term P, and an offset B are applied to predict the first chroma sample 402C, enabling the reconstruction of the current coded block. The non-linear term P is determined based on a plurality of adjacent luminance samples and a subset of the first luminance samples. In some embodiments, the video bitstream 116 further includes a second syntax element 422 for a hypothesized tap index that selects at least one of two hypothesized tap combinations 416 corresponding to the five-tap model for the current coded block 406A, and the two hypothesized tap combinations 416 include a horizontal hypothesized tap combination 416A and a vertical hypothesized tap combination 416B. Additionally, in some embodiments, the second syntax element 422 is signaled at one of block level, superblock level, frame level, keyframe level, and sequence level. In contrast, in some embodiments, the video bitstream 116 does not include a second syntax element 422 for a hypothesized tap index that defines a hypothesized tap combination corresponding to the five-tap model for the current coded block 406A. One of the horizontal hypothesized tap combination 416A and the vertical hypothesized tap combination 416B is applied by default.
[0073] In some embodiments not shown, if the second syntax element 422 is signaled in the video bitstream 116, the second syntax element includes a first bit and a second bit. The first bit indicates whether the horizontal hypothesized tap combination 416A is enabled, and the second bit indicates whether the vertical hypothesized tap combination 416B is enabled. The two hypothesized tap combinations 416A and 416B are controlled independently of each other. Alternatively, in some embodiments, the second syntax element 422 includes a single bit that has: (1) a first value (e.g., “0”) that indicates the horizontal hypothesized tap combination 416A is enabled, and (2) a second value (e.g., “1”) that indicates the vertical hypothesized tap combination 416B is enabled. Only one of the two hypothesized tap combinations 416A and 416B is enabled at a time.
[0074] In an example, the horizontal hypothesized tap combination 416A is identified based on the hypothesized tap index. When it is determined that the five-tap model corresponds to the horizontal hypothesized tap combination 416A, the left adjacent luminance sample 404W and the right adjacent luminance sample 404E are identified as a pair of adjacent luminance samples 404X. According to the five-tap model (e.g., as described in Equation (1)), the first chroma sample 402C is equal to the weighted sum of the first luminance sample 404C, the left adjacent luminance sample 404W, the right adjacent luminance sample 404E, the non-linear term P, and the offset term B. Equation (1) is modified to: predChromaVal = c0C + c1W + c2E + c3P + c4B (2) Wherein, W and E are the luminance values of the left adjacent luminance sample 404W and the right adjacent luminance sample 404E, respectively.
[0075] Alternatively, in another example, a vertical hypothesis tap combination 416B is identified based on a hypothesized tap index. When it is determined that the five-tap model corresponds to this vertical hypothesis tap combination 416B, the upper adjacent luminance sample 404N and the lower adjacent luminance sample 404S are identified as a pair of adjacent luminance samples 404X. According to this five-tap model (e.g., as described in Equation (1)), the first chroma sample 402C is equal to the weighted sum of the first luminance sample 404C, the upper adjacent luminance sample 404N, the lower adjacent luminance sample 404S, the non-linear term P, and the offset term B. Equation (1) is modified to: predChromaVal = c0C + c1N + c2S + c3P + c4B (3) Wherein, N and S are the luminance values of the upper adjacent luminance sample 404N and the lower adjacent luminance sample 404S, respectively.
[0076] In some embodiments, the first chroma sample 402C is determined based on the luminance Direct Current (DC) value lumaDC associated with the current coding block 406A. A plurality of hypothesized values 424 are generated by subtracting the luminance DC value from each of the luminance samples in the first luminance sample 404C and a pair of adjacent luminance samples 404X. The plurality of hypothesized values 424, the non-linear term P of a subset of the plurality of hypothesized values 424, and the offset term B are combined based on a plurality of weighting factors c0 - c4 to generate the first chroma sample 402C co-located with the first luminance sample 404C. In some embodiments, referring to Figure 4B , based on the five-tap model, the horizontal hypothesis tap combination 416A is applied to generate the first chroma sample 402C, as shown in the following equation: predChromaVal = c0e + c1a + c2b + c3P + c4B (4) Wherein, the first hypothesized value 424C(e) is equal to C - lumaDC, the left hypothesized value 424W(a) is equal to W - lumaDC, and the right hypothesized value 424E(b) is equal to E - lumaDC. Alternatively, in some embodiments, referring to Figure 4C , based on the five-tap model, the vertical hypothesis tap combination 416B is applied to generate the first chroma sample 402C, as shown in the following equation: predChromaVal = c0e + c1a’ + c2b’ + c3P + c4B (5) Wherein, the first hypothesized value 424C(e) is equal to C-lumaDC, the upper hypothesized value 424N(a') is equal to N-lumaDC, and the lower hypothesized value 424S(b') is equal to S-lumaDC. In some embodiments, the non-linear term P is expressed as: P = (C × C + B-lumaDC × lumaDC) >> bitdepth (bit depth) (6) Wherein, C is the luminance value of the first luminance sample 404C co-located with the first chrominance sample 402C.
[0077] In some embodiments, the luminance DC value lumaDC and a plurality of weighting factors c0 to c4 are determined based on a set of one or more reference luminance samples 404R and a set of one or more co-located reference chrominance samples 402R within the reference region 412 of the current coding block 406A. The reference region 412 is located in the current image frame 408. Further, in some embodiments, the reference luminance samples 404R of the reference region 412 are used to generate corresponding reference hypothesized values 424, and these reference hypothesized values 424 are further combined to regenerate one or more chrominance samples 402C based on Equation (4) or (5). In some embodiments, a set of one or more co-located reference chrominance samples 402R is compared with one or more regenerated chrominance samples to generate a Least Mean Square (LMS) value. The plurality of weighting factors c0 to c4 are iteratively adjusted to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., the LMS value is less than a threshold LMS value, or the LMS value is minimized).
[0078] In some embodiments, at least one of the weighting factors c0 to c4 is derived based on chrominance samples and luminance samples within the reference region 412 of the current coding block 406A, and the reference region 412 includes one or more coding blocks decoded before the current coding block 406A (e.g., Figure 4A 8 coding blocks are shown). In some embodiments, a subset of one or more coding blocks is adjacent to the current coding block 406A. In some embodiments, a subset of one or more coding blocks is separated from the current coding block 406A by one or more coding blocks. In some embodiments, the reference region 412 includes at least a portion of one or more rows above and / or a portion of one or more columns to the left of the current coding block 406A. For example, reference Figure 4A , the reference region 412 includes 7 rows of luminance samples 404 above the current coding block 406A and 9 columns of luminance samples 404 to the left of the current coding block 406A.
[0079] In some embodiments, at least one of the weighting factors c0 to c4 is determined by minimizing the mean square error (MSE) between the predicted chroma samples 402 and the reconstructed chroma samples 402 in the reference region 412. The MSE is minimized by calculating the autocorrelation matrix of the luminance samples 404 and the cross-correlation vector between the luminance samples 404R of the reference region 412 and the chroma samples 402R. The autocorrelation matrix is processed with LDL decomposition, and back substitution is used to calculate the plurality of weighting factors. This process generally follows the calculation of the filter coefficients of the adaptive loop filter (ALF) in the enhanced compression model (ECM) video coding. The LDL decomposition does not use square root operations and only uses integer arithmetic operations.
[0080] Figure 5 Another exemplary scheme 500 for generating a first chroma sample 402C from a plurality of luminance samples 404 (e.g., 404C and 404X) is shown in accordance with some embodiments. In some embodiments, the video bitstream 116 includes a first syntax element 420 for the MH-CCP mode. The first chroma sample 402C of the current coding block 406A is configured to be generated by combining at least the first luminance sample 404C co-located with the first chroma sample 402C and one or more adjacent luminance samples 404X of the first luminance sample 404C using a plurality of weighting factors (e.g., c0 to c6). When it is determined that the MH-CCP mode is enabled, the first chroma sample 402C is predicted according to the following equation: predChromaVal = c0e + c1a + c2b + c3c + c4d + c5P + c6B (7) where predChromaVal is the predicted chroma value of the first chroma sample 402C; e is the hypothesized value of the first luminance sample 404C co-located with the first chroma sample 402C; a, b, c, and d are the hypothesized values of the adjacent luminance samples 404X of the first luminance sample 404C; P is a non-linear term; B is an offset; and c0 to c6 are weighting factors. In some embodiments, B is the middle luminance value of the luminance value range (e.g., 255 in the range of [0, 511]). In some embodiments, B is the median luminance value or the average luminance value of the luminance samples 404 of the current coding block 406A. Equation (7) includes seven terms and represents a seven-tap model for determining the first chroma sample 402C of the current coding block 406A in the MH-CCP mode.
[0081] In some embodiments, it is assumed that values a, b, c, and d correspond to a north adjacent luminance sample 404N, a south adjacent luminance sample 404S, a west adjacent luminance sample 404W, and an east adjacent luminance sample 404E, respectively. Each of the assumed values e, a, b, c, and d is equal to the difference between the luminance value of the corresponding adjacent luminance sample 404X (e.g., C, N, S, W, S) and the luminance DC value lumaDC. Alternatively, in some embodiments, it is assumed that values a, b, c, and d correspond to a northwest adjacent luminance sample 404NW, a southeast adjacent luminance sample 404SE, a southwest adjacent luminance sample 404SW, and a northeast adjacent luminance sample 404NE, respectively. Each of the assumed values e, a, b, c, and d is equal to the difference between the luminance value of the corresponding adjacent luminance sample 404X (e.g., C, NW, SE, SW, NE) and the luminance DC value lumaDC.
[0082] In some embodiments, the luminance DC value lumaDC is signaled in the video bitstream 116. Alternatively, in some embodiments, lumaDC is not signaled. In some embodiments, the luminance DC value lumaDC and a plurality of weighting factors c0 to c4 are determined based on a set of one or more co-located reference chrominance samples 402R and a set of one or more reference luminance samples 404R within a reference region 412 of the current coding block 406A. The plurality of weighting factors c0 to c6 are iteratively adjusted to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., the LMS value is less than a threshold LMS value or is minimized by the LMS value).
[0083] Figure 6A , Figure 6B and Figure 6C are three examples 600, 620, and 640 of a block size group mapping table according to some embodiments. The video bitstream includes a first syntax element 420 for the MH-CCP mode that indicates whether to reconstruct the chrominance samples 402 based at least on the luminance samples 404 co-located with each of the plurality of chrominance samples 402 of the current coding block 406A ( Figure 4A ) and one or more adjacent luminance samples corresponding to the luminance sample. The video bitstream further includes a second syntax element 422 for an assumed tap index that selects at least an assumed tap combination (e.g., corresponding to a five-tap model) for the current coding block 406A. In some embodiments ( Figure 4A ), the assumed tap combination is selected from two assumed tap combinations 416 that include a horizontal assumed tap combination 416A and a vertical assumed tap combination 416B. In some embodiments, the second syntax element 422 is signaled at one of a block level, a superblock level, a frame level, a keyframe level, and a sequence level.
[0084] In some embodiments, the second syntax element 422 is encoded in the video bitstream 116 based on the context 602 of the current coding block 406C. The context 602 of the current coding block 406 is determined based on a block size 604 corresponding to one of the following items of the current coding block 406C: block width, block height, minimum block width, minimum block height, maximum block width, maximum block height, and the product of the block width and the block height. For example, referring to Figure 6A , the block size 604 is selected from 22 block size options, such as 4×4, 4×8,..., and 64×16. The current coding block 406C including the first chroma sample 402C is reconstructed based on the context 602 of the current coding block 406C. Further, in some embodiments, the context 602 corresponds to a plurality of predefined contexts identified by a plurality of predefined context group identifiers 606. Based on the block size 604 of the current coding block 406C, one of the plurality of predefined context group identifiers 606 is selected to identify one of the plurality of predefined contexts. Based on one of the plurality of predefined context group identifiers 606, the corresponding context 602 stored in the decoder 122 is extracted.
[0085] In some embodiments, the block size group mapping tables 600, 620, or 640 map a plurality of block sizes 604 to a plurality of predefined context group identifiers 606 representing a plurality of predefined contexts. According to the block size group mapping table, the context 602 of the current coding block 406A is determined based on the block size 604 of the current coding block 406A. More specifically, in some embodiments, a group of pictures (GOP) includes a set of available block sizes 604. The plurality of block sizes 604 of the block size group mapping tables 600, 620, or 640 correspond to less than all of the available block sizes. The MH-CCP mode is applied to a subset of the coding blocks having a plurality of block sizes in the block size group mapping tables 600, 620, or 640. Alternatively, in some embodiments, the plurality of block sizes 604 of the block size group mapping tables 600, 620, or 640 correspond to all of the available block sizes in the GOP. In some embodiments, the plurality of block sizes are uniquely associated with the plurality of predefined context group identifiers 606 according to the block size group mapping table 640.
[0086] In some embodiments, the video bitstream 116 further includes a third syntax element 426 for a context flag that indicates whether entropy coding is based on the context 602 of the current coding block 406C, and the context 602 is selected from a plurality of predefined contexts based on the block size 604 of the current coding block 406C. In some embodiments, one of the corresponding hypothesis tap combinations 416A and 416B of the two hypothesis tap combinations corresponding to the five-tap model is selected for the current coding block 406A based on the context 602.
[0087] In some embodiments, when it is determined to enable the MH-CCP mode, a predefined context 602 is employed to perform entropy coding on the current coding block 406. The predefined context 602 is not signaled in the video bitstream 116. Further, in some embodiments, a second syntax element 422 is encoded in the video bitstream 116 based on the predefined context 602 of the current coding block 406C.
[0088] Figure 7 is a flowchart illustrating a video decoding method 700 according to some embodiments. The method 700 may be performed at a computing system (e.g., the server system 112, the source device 102, or the electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 700 is performed by executing instructions stored in the memory of the computing system (e.g., the memory 314). In some embodiments, the method 800 is applied in conjunction with one or more video codecs including, but not limited to, H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, in the MH-CCP mode, a linear or non-linear weighted sum of multiple types of co-located luma samples 404C ( Figure 4A ) is used to predict the chroma value 402C. The multiple types of co-located luma samples 404C are derived from the co-located luma samples 404C or filtered co-located luma samples using adjacent luma samples 404X (e.g., 404W, 404N, 404E, 404S, 404NW, 404NE, 404SW, 404SE) as filtering inputs. Each input to the weighted sum (e.g., corresponding to a respective type of co-located luma sample) is referred to as a hypothesis. In some cases, the value of each hypothesis is fed into a least mean square calculation kernel to derive the weight of the corresponding hypothesis used in the MH-CCP. In some embodiments, a first luma sample 404C has eight adjacent samples 404X ( Figure 5 ) adjacent to the first luma sample 404C. Further, in some embodiments, when luma and chroma have different dimensions (e.g., 4:2:2 or 4:2:0), each luma sample 404 (e.g., 404C, 404X) includes a downsampled luma sample generated using a downsampling filter. Alternatively, each luma sample 404 includes the original co-located luma sample without any downsampling.
[0089] In one aspect of the present invention, the MH-CCP mode is applied based on the luma DC value lumaDC associated with the current coding block 406A. In some embodiments, the luma DC value DClumaDC is determined based on the average luminance value of a set of one or more reference samples 404R in the reference region 412, and this luma DC value is also applied to determine the scaling factor parameters (also referred to as the weighting factors c0 to c6 in Equation (7)) during model derivation. In an example, the input (e.g., the hypothesis value) in model derivation is: where C is the luminance value of the reference luminance sample 404R co-located with the reference chroma sample 402R; N and S are the luminance values of the upper adjacent luminance sample and the lower adjacent luminance sample of the reference luminance sample 404R respectively; W and E are the luminance values of the left adjacent luminance sample and the right adjacent luminance sample of the reference luminance sample 404R respectively. P is a non-linear value, and the intermediate value corresponds to a predefined luminance value range. B is an offset term.
[0090] In some embodiments, after the weighting factors c0 to c6 are known and during sample prediction, the luma DC value DC is further applied to predict the first chroma sample 402C as a combination of the first luminance sample 404C and the associated adjacent luminance sample 404X, as follows: predChromaVal = c0e + c1a + c2b + c3c + c4d + c5P + c6B + chromDC (9) where, except that C is the luminance value of the first luminance sample 404C co-located with the first chroma sample 402C, a, b, c, d, e, and P are all determined based on the luma DC value lumaDC using Equation (8); N and S are the luminance values of the upper adjacent luminance sample 404N and the lower adjacent luminance sample 404S of the first luminance sample 404C respectively; W and E are the luminance values of the left adjacent luminance sample 404W and the right adjacent luminance sample 404E of the first luminance sample 404C respectively.
[0091] The video bitstream 116 is received (operation 702), and a first syntax element for the MH-CCP mode is signaled (operation 704). In some embodiments, for the current coding block 406A, the luma DC value is signaled in the video bitstream 116. Alternatively, in some embodiments, the luma DC value is not signaled in the video bitstream 116. For example, the luma DC value lumaDC is determined based on the average luminance value of the luminance samples 404 in the current coding block 406A, and this luma DC value is applied during the convolution process to obtain a predicted value.
[0092] On the other hand, the five-tap model is applied (operations 706 and 708) in the MH-CCP mode to reduce the complexity of deriving and predicting model parameters. In other words, both the encoder 106 and the decoder 122 adopt Equation (1) (e.g., with five terms) in the MH-CCP mode. In some embodiments, two five-tap models are adopted, and the two five-tap models correspond to the horizontal hypothesis tap combination 416A (where the luminance samples 404 are arranged in the horizontal direction) and the vertical hypothesis tap combination 416B (where the luminance samples 404 are arranged in the vertical direction). In some embodiments, according to Equation (2), the five-tap model includes (operation 710) three terms associated with the luminance samples 404W, 404C, and 404E, a non-linear term P, and an offset term B. Alternatively, in some embodiments, according to Equation (3), the five-tap model includes three terms associated with the luminance samples 404N, 404C, and 404S, a non-linear term P, and an offset term B. Thus, according to Equation (2) or Equation (3), a chrominance value 402C is generated (operation 712) based on at least two adjacent luminance samples 404X, enabling the reconstruction of the current encoded block 406A including the first chrominance sample 402C.
[0093] In some embodiments, only the five-tap model corresponding to the horizontal hypothesis tap combination 416A is applied, e.g., without any signaling in the video bitstream 116. Alternatively, in some embodiments, only the five-tap model corresponding to the horizontal hypothesis tap combination 416A is applied, e.g., without any signaling in the video bitstream 116. In some embodiments, the video bitstream 116 signals a first syntax element 420 for the MH-CCP mode and a second syntax element 422 for the hypothesis tap index, and the hypothesis tap index selects at least one of the two hypothesis tap combinations 416 corresponding to the five-tap model for the current encoded block 406A.
[0094] In some embodiments, the five-tap model is applied in the MH-CCP mode, and each term is determined based on the luminance DC value lumaDC, for example, during the model derivation process and the sample prediction process. According to Equations (4) and (5), the five-tap model is modified based on the luminance DC values lumaDC of the horizontal hypothesis tap combination 416A and the vertical hypothesis tap combination 416B, respectively.
[0095] In yet another aspect, a hypothesized tap combination is selected from multiple combinations of the cross-component samples 404 (e.g., corresponding to multiple hypothesized tap combinations). The selection of the combination of the unfiltered or filtered luma samples 404 is signaled in the video bitstream 116 and parsed at the decoder 1222. In some embodiments, two hypothesized tap combinations 416 are applied, which include a first tap combination 416A (corresponding to five terms associated with N, S, C, P, B) and a second tap combination 416B (corresponding to five terms associated with W, E, C, P, B). The video bitstream 116 includes a second syntax element 422 of a hypothesized tap index for selecting one of the two hypothesized tap combinations 416. Optionally, the second syntax element 422 is signaled at least at one of block level, superblock level, frame level, keyframe level, and sequence level.
[0096] In some embodiments, the hypothesized tap combination is encoded based on the context 602 of the current coding block 406A ( Figures 6A to 6C ). The context 602 is determined based on the block size 604. In some embodiments, the block size 604 is one of the following: block width, block height, minimum block width and minimum block height, maximum block width and maximum block height, product of block width and block height. In some embodiments, different block sizes 604 are grouped based on the context 602. Referring to Figure 6A , in an example, 22 block sizes 604 are grouped into 7 context groups associated with context group identifiers 606 (e.g., 0 to 6). Referring to Figure 6B , in another example, 9 block sizes 604 are grouped into 4 context groups associated with context group identifiers 606 (e.g., 0 to 3). Referring to Figure 6C , in another example, 9 block sizes 604 are grouped into 9 context groups associated with context group identifiers 606 (e.g., 0 to 8), i.e., the block sizes 604 are uniquely associated with the context groups such that multiple block sizes 604 respectively match multiple context groups. The encoder 106 and the decoder 122 store the same block size group mapping tables (e.g., Figures 6A to 6C the tables 600, 620, and 640 in), enabling the identification of the same context 602 for encoding and decoding the current coding block 406A.
[0097] In some embodiments, the MH-CCP mode is applied to a subset of the block sizes of an image frame or GOP (e.g., less than all block sizes). For example, the number of block sizes of the coded blocks of an image frame associated with the block size group mapping table 620 is more than nine block sizes, and the MH-CCP mode is only applied to the nine block sizes 604 listed in the mapping table 620. In contrast, in some embodiments, the MH-CCP mode is applied to all block sizes of an image frame or GOP. In the above example, the coded blocks of the image frame associated with the block size group mapping table 620 only have nine block sizes. Additionally, in some embodiments, four context group identifiers 606 (e.g., Figure 6B 0-3 in
[0098] correspond to four different hypothesis tap combinations 416. Alternatively, in some embodiments, the four context group identifiers 606 correspond to two different hypothesis tap combinations 416 (e.g., 416A and 416B). Figure 7 Although multiple logical stages are shown in a particular order, stages that are not order-dependent can be reordered, and other stages can be combined or separated. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0099] Some example embodiments will now be described.
[0100] (A1) In some implementations, method 800 is implemented for decoding video data. Method 800 includes: receiving (operation 702) a video bitstream that includes a current coded block of a current image frame, where the video bitstream includes (operation 704) a first syntax element for the multi-hypothesis cross-component prediction (MH-CCP) mode. Method 800 further includes: determining (operation 706) to enable the MH-CCP mode based on the first syntax element in the video bitstream to reconstruct a chrominance sample based at least on a luminance sample co-located with each of a plurality of chrominance samples of the current coded block and one or more adjacent luminance samples corresponding to the luminance sample. Method 800 further includes: identifying (operation 708) a five-tap model configured to determine a chrominance sample of the current coded block in the MH-CCP mode; identifying (operation 710) a pair of adjacent luminance samples of a first luminance sample based on the five-tap model; generating (operation 712) a first chrominance sample co-located with the first luminance sample based at least on the first luminance sample and the pair of adjacent luminance samples; and reconstructing (operation 714) the current coded block including the first chrominance sample.
[0101] (A2) In some embodiments of A1, generating the first chroma sample based at least on the first luminance sample and the pair of adjacent luminance samples further includes: combining the first luminance sample, the pair of adjacent luminance samples, non-linear terms of a plurality of adjacent luminance samples and a subset of the first luminance sample, and an offset term based on a plurality of weighting factors.
[0102] (A3) In some embodiments of A1 or A2, the video bitstream further includes a second syntax element for a hypothesized tap index, the hypothesized tap index selecting at least one hypothesized tap combination corresponding to a five-tap model for the current coding block, and the two hypothesized tap combinations including a horizontal hypothesized tap combination and a vertical hypothesized tap combination.
[0103] (A4) In some embodiments of A3, identifying the pair of adjacent luminance samples of the first luminance sample further includes: identifying a horizontal hypothesized tap combination corresponding to the five-tap model based on the hypothesized tap index; and when it is determined that the five-tap model corresponds to the horizontal hypothesized tap combination, identifying the left adjacent luminance sample and the right adjacent luminance sample as the pair of adjacent luminance samples.
[0104] (A5) In some embodiments of A3, identifying the pair of adjacent luminance samples of the first luminance sample further includes: identifying a vertical hypothesized tap combination corresponding to the five-tap model based on the hypothesized tap index; and when it is determined that the five-tap model corresponds to the vertical hypothesized tap combination, identifying the upper adjacent luminance sample and the lower adjacent luminance sample as the pair of adjacent luminance samples.
[0105] (A6) In some embodiments of any one of A3 to A5, the second syntax element includes a first bit and a second bit. The first bit indicates whether the horizontal hypothesized tap combination is enabled, and the second bit indicates whether the vertical hypothesized tap combination is enabled.
[0106] (A7) In some embodiments of any one of A3 to A5, the second syntax element includes a single bit having: (1) a first value that indicates enabling the horizontal hypothesized tap combination; and (2) a second value that indicates enabling the vertical hypothesized tap combination.
[0107] (A8) In some embodiments of any one of A1 to A7, the five-tap model corresponds to the horizontal hypothesized tap combination, in which the pair of adjacent luminance samples includes a left adjacent luminance sample and a right adjacent luminance sample. According to the five-tap model, the first chroma sample is equal to the weighted sum of the first luminance sample, the left adjacent luminance sample, the right adjacent luminance sample, the non-linear term, and the offset term.
[0108] (A9)In some embodiments of any one of A1 to A7, the five-tap model corresponds to a vertical hypothesis tap combination in which the pair of adjacent luminance samples includes an upper adjacent luminance sample and a lower adjacent luminance sample; and according to the five-tap model, the first chrominance sample is equal to the weighted sum of the first luminance sample, the upper adjacent luminance sample, the lower adjacent luminance sample, a non-linear term, and an offset term.
[0109] (A10)In some embodiments of any one of A1 to A9, the video bitstream does not include a second syntax element for a hypothesis tap index that is used to define a hypothesis tap combination corresponding to the five-tap model for the current coding block, and one of the horizontal hypothesis tap combination and the vertical hypothesis tap combination is applied by default.
[0110] (A11)In some embodiments of any one of A1 to A9, generating the first chrominance sample based on the first luminance sample and the pair of adjacent luminance samples further includes: determining a luminance DC value of the current coding block; generating a plurality of hypothesis values by subtracting the luminance DC value from each of the first luminance sample and the pair of adjacent luminance samples; and combining the plurality of hypothesis values, a non-linear term of a subset of the plurality of hypothesis values, and an offset term based on a plurality of weighting factors to generate a first chrominance sample co-located with the first luminance sample.
[0111] (A12)In some embodiments of A11, method 700 further includes: identifying a reference region of the current coding block; determining the luminance DC value based on an average luminance value of a set of one or more reference samples in the reference region; and determining a plurality of weighting factors based on the set of one or more reference samples in the reference region and the luminance DC value.
[0112] (A13)In some embodiments of A12, determining the plurality of weighting factors further includes: determining a least mean square (LMS) value based on the set of one or more reference samples and the luminance DC value; and iteratively adjusting the plurality of weighting factors to reduce the LMS value until the LMS value meets a predefined criterion.
[0113] (A14)In some embodiments of any one of A1 to A13, the video bitstream further includes a second syntax element for a hypothesis tap index that selects at least one of two hypothesis tap combinations corresponding to the five-tap model for the current coding block, and the two hypothesis tap combinations include a horizontal hypothesis tap combination and a vertical hypothesis tap combination, and wherein the second syntax element is signaled at one of block level, super-block level, frame level, key-frame level, and sequence level.
[0114] (A15)In some embodiments of A14, the second syntax element is encoded in the video bitstream based on the context of the current coding block. The method 800 further includes: determining the context of the current coding block based on a block size corresponding to one of the following items of the current coding block, and these items are: block width, block height, minimum block width, minimum block height, maximum block width, maximum block height, and the product of the block width and the block height. The current coding block including the first chrominance sample is reconstructed based on the context of the current coding block.
[0115] (A16)In some embodiments of A15, determining the context of the current coding block based on the block size further includes: based on the block size of the current coding block, selecting one predefined context group identifier from a plurality of predefined context group identifiers associated with a plurality of predefined contexts; and based on one predefined context group identifier among the plurality of predefined context group identifiers, extracting the context stored in the decoder.
[0116] (A17)In some embodiments of A15 or A16, selecting one predefined context group identifier from a plurality of predefined context group identifiers further includes: obtaining a block size group mapping table that maps a plurality of block sizes to a plurality of predefined context group identifiers. The context of the current coding block is determined based on the block size group mapping table based on the block size of the current coding block.
[0117] (A18)In some embodiments of A17, a group of pictures (GOP) includes a set of available block sizes, and the plurality of block sizes of the block size group mapping table correspond to less than all of the available block sizes, and wherein the MH-CCP mode is applied to a subset of coding blocks having a plurality of block sizes.
[0118] (A19)In some embodiments of A17 or A18, the plurality of block sizes are uniquely associated with a plurality of predefined context group identifiers according to the block size group mapping table.
[0119] (A20)In some embodiments of any one of A1 to A19, the video bitstream further includes a third syntax element for a context flag that indicates whether entropy coding is based on the context of the current coding block, and the context is selected from a plurality of predefined contexts based on the block size of the current coding block. A corresponding one of the two hypothesized tap combinations corresponding to the five-tap model is selected for the current coding block based on the context.
[0120] (A21)In some embodiments of any one of A1 to A20, when it is determined to enable the MH-CCP mode, a predefined context is used to perform entropy coding on the current coding block.
[0121] (A22)A computing system includes a control circuit and a memory. The memory stores one or more programs configured to be executed by the control circuit. The one or more programs further include instructions for: receiving video data including a current coded block of a current image frame; encoding the current image frame according to intra prediction; determining to enable a multi-hypothesis cross-component prediction (MH-CCP) mode to determine a chrominance sample based on at least a luminance sample co-located with each of a plurality of chrominance samples and a sum of one or more neighboring luminance samples corresponding to the luminance sample, wherein the MH-CCP mode is associated with a five-tap model for identifying a pair of neighboring luminance samples of a first luminance sample; transmitting the encoded current image frame via a video bitstream; and signaling a first syntax element via the video bitstream to indicate application of the MH-CCP mode to reconstruct a first chrominance sample co-located with the first luminance sample based at least on the first luminance sample and the pair of neighboring luminance samples.
[0122] (A23)A non-transitory computer-readable storage medium stores one or more programs executed by a control circuit of a computing system. The one or more programs include instructions for: obtaining a source video sequence including a current coded block of a current image frame; and converting between the source video sequence and a video bitstream, wherein the video bitstream includes: the current coded block of the current image frame; and a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode indicating whether to reconstruct the chrominance sample based at least on a luminance sample co-located with each of a plurality of chrominance samples of the current coded block and one or more neighboring luminance samples corresponding to the luminance sample. The MH-CCP mode is associated with a five-tap model for identifying a pair of neighboring luminance samples of a first luminance sample, and the MH-CCP mode is applied to reconstruct a first chrominance sample co-located with the first luminance sample based at least on the first luminance sample and the pair of neighboring luminance samples.
[0123] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit. The memory stores one or more instruction sets configured to be executed by the control circuit. The one or more instruction sets include instructions for performing any of the methods described herein (e.g., A1 to A23 above).
[0124] In another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets for execution by a control circuit of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A23 above).
[0125] Unless otherwise specified, any syntactic elements described herein may be high-level syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements may be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, picture header, tile header, and / or coding tree unit (CTU) header.
[0126] It will be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the listed related items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the listed features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0127] As used herein, depending on context, the term "if" can be interpreted to mean "when", or "in the case where", or "in response to determining", or "in accordance with determining", or "in response to detecting" that the described precondition is true. Similarly, depending on context, the phrase "if it is determined [that the described precondition is true]" or "if [the described precondition is true]" or "when [the described precondition is true]" can be interpreted to mean "upon determining", or "in response to determining", or "in accordance with determining", or "upon detecting", or "in response to detecting" that the described precondition is true.
[0128] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments chosen and described are to best explain the principles of operation and practical application, and to enable others skilled in the art to implement.
Claims
1. A method for decoding video data, comprising: Receiving a video bitstream, the video bitstream comprising a current coding block of a current image frame, wherein the video bitstream comprises a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode; Based on the first syntax element in the video bitstream, determine to enable the MH-CCP mode to reconstruct the chroma samples based on at least a luma sample co-located with each chroma sample of a plurality of chroma samples of the current coding block and one or more adjacent luma samples corresponding to the luma sample; identifying a five-tap model configured to determine chroma samples of the current coding block in the MH-CCP mode; Based on the five-tap model, identifying a pair of adjacent brightness samples of a first brightness sample; generating a first chroma sample co-located with the first luma sample based at least on the first luma sample and the pair of adjacent luma samples; and The current coding block including the first chroma samples is reconstructed.
2. The method according to claim 1, wherein: Generating the first chrominance sample based at least on the first luma sample and the pair of adjacent luma samples comprises: The first luma sample, the pair of adjacent luma samples, a non-linear term of the plurality of adjacent luma samples and a subset of the first luma samples, and an offset term are combined based on a plurality of weighting factors.
3. The method according to claim 1, wherein: The video bitstream also includes a second syntax element for assuming a tap index, wherein the assumed tap index selects at least one of two assumed tap combinations corresponding to the five-tap model for the current coding block, and the two assumed tap combinations include a horizontal assumed tap combination and a vertical assumed tap combination.
4. The method according to claim 3, wherein: Identifying a pair of adjacent luma samples of the first luma sample further comprises: identifying, based on the hypothesis tap index, a horizontal hypothesis tap combination corresponding to the five-tap model; and When it is determined that the five-tap model corresponds to the horizontal hypothesis tap combination, a left-neighboring luma sample and a right-neighboring luma sample are identified as the pair of adjacent luma samples.
5. The method according to claim 3, wherein: Identifying a pair of adjacent luma samples of the first luma sample further comprises: identifying, based on the hypothesis tap index, a vertical hypothesis tap combination corresponding to the five-tap model; and When it is determined that the five-tap model corresponds to the vertical hypothetical tap combination, an upper adjacent luma sample and a lower adjacent luma sample are identified as the pair of adjacent luma samples.
6. The method according to claim 3, wherein: The second syntax element includes a first bit and a second bit, and wherein the first bit indicates whether the horizontal hypothesis tap combination is enabled, and the second bit indicates whether the vertical hypothesis tap combination is enabled.
7. The method according to claim 3, wherein: The second syntax element includes a single bit having: (1) a first value indicating that the horizontal hypothesis tap combination is enabled; and (2) a second value indicating that the vertical hypothesis tap combination is enabled.
8. The method according to claim 1, wherein: The five-tap model corresponds to a horizontal hypothesis tap combination in which the pair of adjacent luminance samples includes a left adjacent luminance sample and a right adjacent luminance sample; and According to the five-tap model, the first chrominance sample is equal to a weighted sum of the first luma sample, the left-neighboring luma sample, the right-neighboring luma sample, a nonlinear term, and an offset term.
9. The method according to claim 1, wherein: The five-tap model corresponds to a vertical hypothesis tap combination in which the pair of adjacent luminance samples includes an upper adjacent luminance sample and a lower adjacent luminance sample; and According to the five-tap model, the first chrominance sample is equal to a weighted sum of the first luma sample, the upper adjacent luma sample, the lower adjacent luma sample, a nonlinear term, and an offset term.
10. The method according to claim 1, wherein: The video bitstream does not include a second syntax element for a hypothetical tap index, wherein the hypothetical tap index is used to define a hypothetical tap combination corresponding to the five-tap model for the current coding block, and one of the horizontal hypothetical tap combination and the vertical hypothetical tap combination is applied by default.
11. The method according to claim 1, wherein: Generating the first chrominance sample based on the first luma sample and the pair of adjacent luma samples further comprises: Determining a luminance direct current (DC) value of the current coding block; generating a plurality of hypothetical values by subtracting the luma DC value from the first luma sample and each of the pair of adjacent luma samples; and The plurality of hypothetical values, a non-linear term for a subset of the plurality of hypothetical values, and an offset term are combined based on a plurality of weighting factors to generate the first chrominance sample co-located with the first luma sample.
12. The method according to claim 11, further comprising: Identifying a reference region of the current coding block; determining the luminance DC value based on an average luminance value of a set of one or more reference samples in the reference area; as well as The plurality of weighting factors are determined based on the set of the one or more reference samples in the reference region and the luminance DC value.
13. The method according to claim 12, wherein: Determining the plurality of weighting factors further comprises: determining a least mean square (LMS) value based on the set of the one or more reference samples and the luminance DC value; and The plurality of weighting factors are iteratively adjusted to reduce the LMS value until the LMS value meets a predefined criterion.
14. The method according to claim 1, wherein: The video bitstream also includes a second syntax element for assuming a tap index, wherein the assumed tap index selects at least one of two assumed tap combinations corresponding to the five-tap model for the current coding block, and the two assumed tap combinations include a horizontal assumed tap combination and a vertical assumed tap combination, and wherein the second syntax element is signaled at one of a block level, a super block level, a frame level, a key frame level, and a sequence level.
15. The method according to claim 14, wherein: The second syntax element is encoded in the video bitstream based on the context of the current coding block, and the method further includes: Determine a context of the current coding block based on a block size corresponding to one of the following items of the current coding block, the items being: a block width, a block height, a minimum block width, a minimum block height, a maximum block width, a maximum block height, and a product of the block width and the block height, wherein the current coding block including the first chroma sample is reconstructed based on the context of the current coding block.
16. The method according to claim 15, wherein: Determining the context of the current coding block based on the block size also includes: selecting, based on the block size of the current encoding block, one of a plurality of predefined context group identifiers associated with a plurality of predefined contexts; and The context stored in the decoder is extracted based on a predefined context group identifier among the plurality of predefined context group identifiers.
17. The method according to claim 15, wherein: Selecting one of the plurality of predefined context group identifiers further comprises: A block size group mapping table is obtained, the block size group mapping table mapping a plurality of block sizes to the plurality of predefined context group identifiers, wherein the context of the current coding block is determined based on the block size of the current coding block according to the block size group mapping table.
18. The method according to claim 17, wherein: A group of pictures (GOP) includes a set of available block sizes, and the plurality of block sizes of the block size group map corresponds to less than all of the available block sizes, and wherein the MH-CCP mode is applied to a subset of coding blocks having the plurality of block sizes.
19. The method according to claim 17, wherein: The plurality of block sizes are uniquely associated with the plurality of predefined context group identifiers according to the block size group mapping table.
20. The method according to claim 1, wherein: The video bitstream also includes a third syntax element for a context flag, wherein the context flag indicates whether the entropy coding is based on the context of the current coding block, and the context is selected from multiple predefined contexts based on the block size of the current coding block, and wherein a corresponding one of the two hypothetical tap combinations corresponding to the five-tap model is selected for the current coding block based on the context.
21. The method according to claim 1, further comprising: When it is determined to enable the MH-CCP mode, a predefined context is used to perform entropy encoding on the current coding block.
22. A computing system comprising: Control circuit; as well as A memory storing one or more programs configured to be executed by the control circuit, the one or more programs further comprising instructions for: Receiving video data, the video data comprising a current coding block of a current image frame; Encoding the current image frame according to intra-frame prediction; determining to enable a multi-hypothesis cross-component prediction (MH-CCP) mode to determine a chroma sample based on a sum of at least a luma sample co-located with each chroma sample of a plurality of chroma samples of the current coding block and one or more neighboring luma samples corresponding to the luma sample, wherein the MH-CCP mode is associated with a five-tap model for identifying a pair of neighboring luma samples of a first luma sample; Sending the encoded current image frame via a video code stream; as well as A first syntax element is signaled via the video code stream to indicate application of the MH-CCP mode to reconstruct a first chrominance sample co-located with the first luma sample based on at least the first luma sample and the pair of adjacent luma samples.
23. A non-transitory computer-readable storage medium storing one or more programs executed by control circuitry of a computing system, the one or more programs comprising instructions for: Obtaining a source video sequence, wherein the source video sequence includes a current coding block of a current image frame; as well as Convert between the source video sequence and the video code stream, wherein: The video code stream includes: The current coding block of the current image frame; and a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct each chroma sample of a plurality of chroma samples of the current coding block based on at least a luma sample co-located with the chroma sample and one or more neighboring luma samples corresponding to the luma sample; wherein the MH-CCP mode is associated with a five-tap model, the five-tap model being used to identify a pair of adjacent luma samples of a first luma sample, and the MH-CCP mode being applied to reconstruct a first chroma sample co-located with the first luma sample based at least on the first luma sample and the pair of adjacent luma samples.