Video coding and decoding method, device and medium
By adopting the multi-assumption cross-component prediction (MH-CCP) mode in video encoding and decoding, non-linear prediction of chrominance samples is used to use multiple inputs of the brightness sample to solve the problems of encoding efficiency and quality loss in the prior art, and more efficient video encoding and better video quality are achieved.
Patent Information
- Application Number
- CN202411984973.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-13
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-18
AI Technical Summary
When compressing video data, existing video encoding and decoding technologies are difficult to effectively utilize the redundant information in the video data, resulting in low encoding efficiency and large quality loss. Especially in the case of lossy compression, it is difficult to maintain video quality while ensuring a higher compression ratio.
Multi-assumption Cross-Component Prediction (MH-CCP) mode is adopted to predict chroma samples by multiple inputs of luminance samples, and the chroma samples are reconstructed using multiple nonlinear terms and offset terms, reducing direct encoding of chroma samples, thereby improving encoding efficiency and video quality.
Through the MH-CCP mode, the direct encoding requirement for chroma samples during video encoding is reduced, the encoding efficiency is improved, and the prediction accuracy of chroma samples is improved through multiple nonlinear terms, maintaining the quality of the video.
Smart Images

Figure BDA0005222367160000191 
Figure BDA0005222367160000211 
Figure BDA0005222367160000221
Abstract
Description
Cross-Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 622,535, entitled "Multi-Hypothesis Cross Component Prediction Models," filed on January 18, 2024, and U.S. Provisional Patent Application No. 18 / 803,213, entitled "Multi-Hypothesis Cross Component Prediction Model," filed on August 13, 2024, which are hereby incorporated by reference in their entirety. Technical Field
[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for processing video data using multi-hypothesis cross-component prediction (MH-CCP). Background Art
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video conferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transfer digital video data via a communication network and / or store the digital video data on a storage device. Due to the limited bandwidth capacity of the communication network and the limited memory resources of the storage device, video data can be compressed according to one or more video coding standards using video coding before the video data is transferred or stored. Video coding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.
[0004] Video coding typically utilizes prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a reduction in video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to be a successor to HEVC. The ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, a specification verification version 1.0.0 containing errata 1 was released. Summary of the Invention
[0005] As described above, encoding (compression) reduces the bandwidth and / or storage space requirements. As will be described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully preserved during encoding and cannot be fully recovered during decoding. When lossy compression is used, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. The degree of tolerable distortion depends on the application. For example, users of some consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows for coding algorithms that produce higher losses and higher compression ratios.
[0006] The present disclosure describes a video compression method using intra prediction. A linear or non-linear weighted sum of multiple inputs of luminance samples is used to predict chrominance samples, for example, in multi-hypothesis cross-component prediction (MH-CCP). The multiple inputs of luminance samples include a luminance sample C collocated with the chrominance sample and a filtered luminance sample determined based on adjacent luminance samples and applied as a filtering input. Each filtering input of the weighted sum is called a hypothesis and is fed into a least mean square calculation kernel. In other words, the multiple inputs may include multiple reconstructed luminance samples located around the collocated position and non-linear terms of the reconstructed luminance samples. Some implementations of the present application involve applying multiple non-linear terms in MH-CCP, where each non-linear term is formed based on a luminance sample collocated with the chrominance sample and a subset of one or more associated adjacent luminance samples.
[0007] In some embodiments, according to a multi-tap model associated with MH-CCP, samples of a second color component are predicted as a linear or non-linear weighted sum of multiple inputs determined based on samples of a first color component collocated with the samples of the second color component and one or more associated adjacent samples of the first color component. The multi-tap model includes multiple (N) taps selected from collocated samples of the first color component, one or more associated adjacent samples of the first color component, multiple non-linear terms, and an offset term. The multi-tap model corresponds to the same number (N) of options, which are combined to determine samples of the second color component. Some implementations of the present application involve applying multiple non-linear terms in MH-CCP, where each non-linear term is formed based on samples of the first color component and a subset of one or more associated adjacent samples.
[0008] In some embodiments, samples of the first color component are luminance samples, and samples of the second color component are blue-difference chrominance (Cb) samples or red-difference chrominance (Cr) components. For example, chrominance samples are a weighted combination of terms selected from corresponding collocated luminance samples, one or more adjacent luminance samples, non-linear terms, and an offset term. Alternatively, in some embodiments, the first color component is one of red, green, and blue, and the second color component is another of red, green, and blue. Alternatively, in some embodiments, the first color component and the second color component correspond to a color format different from the YCbCr color format and the RGB color format.
[0009] According to some embodiments, a video decoding method is provided. The method includes receiving a video bitstream that includes a current coded block of a current picture frame, and the video bitstream includes a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode. The method further includes determining, based on the first syntax element, to enable the MH-CCP mode to reconstruct a first chrominance sample of the current coded block based at least on a first luminance sample and associated neighboring luminance samples. The first luminance sample is co-located with the first chrominance sample. The method further includes identifying the first luminance sample and one or more neighboring luminance samples in the current coded block; generating a plurality of non-linear terms based at least on the first luminance sample and a subset of the one or more neighboring luminance samples; predicting the first chrominance sample co-located with the first luminance sample in the current coded block based on the plurality of non-linear terms; and reconstructing the current picture frame including the current coded block.
[0010] According to some embodiments, a video encoding method is provided. The method includes receiving video data including a current coded block of a current picture frame, encoding the current picture frame, transmitting the encoded current picture frame via a video bitstream, and writing, via the video bitstream, a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct a first chrominance sample of the current coded block based on a first luminance sample and associated neighboring luminance samples. The first luminance sample is co-located with the first chrominance sample. When the MH-CCP mode is enabled, a plurality of non-linear terms are determined based at least on the first luminance sample and a subset of one or more neighboring luminance samples of the first luminance sample, and the first chrominance sample co-located with the first luminance sample is predicted based on the plurality of non-linear terms.
[0011] According to some embodiments, a bitstream conversion method is provided. The method includes obtaining a source video sequence including a current picture frame having a current coded block, and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes a current picture frame having a current coded block and a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct a first chrominance sample of the current coded block based on a first luminance sample and associated neighboring luminance samples. The first luminance sample is co-located with the first chrominance sample. When the MH-CCP mode is enabled, a plurality of non-linear terms are determined based at least on the first luminance sample and a subset of one or more neighboring luminance samples of the first luminance sample, and the first chrominance sample co-located with the first luminance sample is predicted based on the plurality of non-linear terms.
[0012] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a code converter).
[0013] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0014] Thus, devices and systems having methods for encoding and decoding video are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all-inclusive, and in particular, considering the drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in the specification is primarily selected for readability and guidance purposes and is not necessarily selected to depict or delimit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To be able to understand the present disclosure in more detail, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the drawings. However, the drawings only show the relevant features of the present disclosure and should not be considered restrictive, as those skilled in the art will understand, after reading the present disclosure, that the description may also have other valid features.
[0016] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0017] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0018] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0019] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0020] Figure 4 shows an example scheme for generating a first chrominance sample from a plurality of luminance samples in the MH-CCP mode according to some embodiments.
[0021] Figure 5A illustrates an example process for predicting a first chrominance sample based on multiple non - linear terms of one or more luminance samples in the MH - CCP mode according to some embodiments.
[0022] Figure 5B and Figure 5C are two example look - up tables according to some embodiments, from which multiple model parameters are identified.
[0023] Figure 6 illustrates another example process for predicting a first chrominance sample 402A based on multiple non - linear terms of one or more luminance samples in the MH - CCP mode according to some embodiments.
[0024] Figure 7 is a flowchart illustrating an example method for decoding video according to some embodiments.
[0025] According to conventional practice, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and drawings to indicate the same features. Detailed Description
[0026] The present disclosure describes cross - component intra - prediction of video data in the MH - CCP mode, in which each of multiple samples of a second color component is determined based on one or more associated samples of a first color component. The MH - CCP mode corresponds to a multi - tap model including multiple (N) taps. Each tap is selected from co - located samples of the first color component, one or more associated adjacent samples of the first color component, multiple non - linear terms, and an offset term. The selected taps are combined in a weighted manner to determine the samples of the second color component. In some embodiments, the samples of the first color component are luminance samples and the samples of the second color component are chrominance samples. For example, a chrominance sample is a weighted combination of terms selected from corresponding co - located luminance samples, one or more adjacent luminance samples, multiple non - linear terms, and an offset term. When a syntax element indicates that the MH - CCP mode is enabled, multiple non - linear terms are determined based at least on a co - located luminance sample and a subset of one or more adjacent luminance samples of the co - located luminance sample, and a chrominance sample is predicted based on the multiple non - linear terms. In the MH - CCP mode, it is not necessary to transmit the samples of the second color component of the current encoded block in the video bitstream, thus saving the communication bandwidth of the video codec. In addition, MH - CCP involves multiple non - linear terms, which can provide a high level of accuracy for sample values compared to a single non - linear term.
[0027] Figure 1is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0028] The source device 102 includes a video source 104 (e.g., a camera component or a media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The data volume of the video stream from the video source 104 may be higher compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to send uncompressed video to the (one or more) networks 110).
[0029] The one or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. The one or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0030] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as an encoded video stream from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the codec component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or multiple video encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0031] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on the display or other type of presentation device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0032] The source device and / or the multiple electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, one or more of the source device 102 and / or the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0033] In an example operation of communication system 100, source device 102 sends encoded video stream 108 to server system 112. For example, source device 102 may encode a picture stream captured by the source device. Server system 112 receives encoded video stream 108 and may use codec component 114 to decode and / or encode encoded video stream 108. For example, server system 112 may apply encoding to the video data to be more suitable for network transmission and / or storage. Server system 112 may send encoded video data 116 (e.g., one or more encoded video streams) to one or more of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video pictures.
[0034] Figure 2A is a block diagram showing example elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from encoder component 106). Video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device storing previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. Those of ordinary skill in the art can easily understand the relationship between pixels and samples.
[0035] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between a source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or lambda value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may be related to the encoder component 106 optimized for a particular system design.
[0036] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols (e.g., a symbol stream) based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data when the compression between the symbols and the encoded video bitstream is lossless. The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. In this way, the reference picture samples interpreted by the prediction part of the encoder are exactly the same as the sample values interpreted by the decoder when using prediction during decoding. This principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those of ordinary skill in the art.
[0037] The operation of the decoder 210 may be the same as that of a remote decoder such as the decoder component 122, which will be described in detail below in connection with Figure 2B However, briefly referring to Figure 2B , when symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into an encoded video sequence losslessly, the entropy decoding part of the decoder component 122 including the buffer memory 252 and the parser 254 may not be fully implemented in the local decoder 210.
[0038] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a corresponding encoder in substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. Since encoder techniques are inverse to decoder techniques, the description of encoder techniques can be simplified.
[0039] As part of its operation, the source encoder 202 can perform motion-compensated predictive coding, which predictively encodes an input frame by referring to one or more previously encoded frames designated as reference frames in a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the (one or more) reference frames, which (one or more) reference frames can be selected as the (one or more) prediction references for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0040] The decoder 210 decodes the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is decoded in a video decoder ( Figure 2A not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on a reference frame and can cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame that has the same content (in the absence of transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0041] The predictor 206 can perform a prediction search on the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate block by block based on sample blocks to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have a prediction reference extracted from multiple reference pictures stored in the reference picture memory 208.
[0042] The outputs of all the above functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by various functional units according to techniques well known to those of ordinary skill in the art (such as Huffman coding, variable-length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.
[0043] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter 440 may be configured to buffer the (one or more) encoded video sequences created by the entropy encoder 214, in preparation for transmission over the communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such other data being, for example, encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.
[0044] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture may be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are familiar with those variations of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. A predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata to reconstruct a single block.
[0045] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks determined by the encoding assignments of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-predictively encoded with reference to a previously encoded reference picture, either through spatial prediction or through temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures, either through spatial prediction or through temporal prediction.
[0046] Video can be acquired as multiple source pictures (video pictures) in a time sequence. Intra picture prediction (usually abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded (referred to as the current picture) is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. In the case of using multiple reference pictures, the motion vector points to the reference block in the reference picture and can have a third dimension identifying the reference picture.
[0047] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (e.g., any video coding technique or standard described herein). The encoder component 106 can perform various compression operations in its operation, including predictive coding operations that utilize temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.
[0048] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).
[0049] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data while receiving other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. The additional data may be part of the (one or more) encoded video sequences. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0050] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.
[0051] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 inside decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (e.g., to prevent network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be needed, or may be made smaller. For use on a service packet network such as the Internet, buffer memory 252 may also be needed, which may be relatively large and / or have an adaptive size and may be implemented at least partially in an operating system or similar element outside decoder component 122.
[0052] The parser 254 is configured to reconstruct a symbol 270 from the encoded video sequence. The symbol may include, for example, information for managing the operation of the decoder component 122 and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set segment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 254 may extract a subgroup parameter set for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), and so on. The parser 254 may also extract information such as transform coefficients, quantizer parameter values, motion vectors from the encoded video sequence.
[0053] The reconstruction of the symbol 270 may involve multiple different units, depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and the way these units are involved may be controlled by subgroup control information parsed by the parser 254 from the encoded video sequence. For the sake of brevity, this subgroup control information flow between the parser 254 and the multiple units below is not described.
[0054] The decoder component 122 may be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and may be at least partially integrated with each other. However, for the sake of clarity, the conceptually subdivided functional units are retained here.
[0055] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as the symbol 270 and control information (e.g., which transform mode to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 may output a block including sample values, which may be input into the aggregator 268.
[0056] In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra-coded block; that is: a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed part of the current picture. Such predictive information may be provided by the intra prediction unit 262. The intra prediction unit 262 uses the surrounding reconstructed information extracted from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 may add, on a per-sample basis, the prediction information generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.
[0057] In other cases, the output samples of the scaler / inverse transform unit 258 may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After motion-compensating the extracted samples according to the syntax element 270 associated with the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to in this case as the residual samples or residual signal), thereby generating the output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 extracts the prediction samples may be controlled by a motion vector. The motion vector may be in the form of a syntax element 270 for use by the motion compensation prediction unit 260, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of the sample values extracted from the reference picture memory 266 when using sub-sampled accurate motion vectors, a motion vector prediction mechanism, and the like.
[0058] The output samples of the aggregator 268 may be employed by various loop filter techniques in the loop filter unit 256. Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and that are available to the loop filter unit 256 as a syntax element 270 from the parser 254. However, video compression techniques may also respond to meta-information obtained during the decoding of a previously (in decoding order) part of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device such as the display 124 and may be stored in the reference picture memory 266 for subsequent inter-picture prediction.
[0059] Once reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture (by, for example, parser 254) is identified as a reference picture, the current reference picture can become part of the reference picture memory 266 and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0060] The decoder component 122 can perform decoding operations according to a predetermined video compression technique, which can be recorded in a standard, such as any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence complies with the syntax of the video compression technique or standard, as specified in the video compression technique document or standard, particularly in the profiles thereof. In addition, to conform to some video compression techniques or standards, the complexity of the encoded video sequence is within the range defined by the levels of the video compression technique or standard. In some cases, the level can limit the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limitations set by the level can be further restricted by assuming a Hypothetical Reference Decoder (HRD) specification and metadata written into the encoded video sequence for HRD buffer management.
[0061] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).
[0062] One or more network interfaces 304 may be configured to interact with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network may be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, and so on. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, and so on. Such communication may be unidirectional receive-only (e.g., broadcast television), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., using a local digital network or a wide area digital network to connect to other computer systems). Such communication may include communication to one or more cloud computing networks.
[0063] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. One or more input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. One or more output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., monitors or displays), etc.
[0064] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices located remotely from the control circuit 302. The memory 314, or alternatively the one or more non-volatile solid-state memory devices within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures or subsets or supersets thereof: · An operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks; · A network communication module 318, which is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via a wired connection and / or a wireless connection); · A codec module 320, which is used to perform various functions regarding encoding and / or decoding data (such as video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes one or more of the following: ○ A decoding module 322, which is used to perform various functions regarding decoding the encoded data, such as those functions previously described regarding the decoder component 122; and ○ An encoding module 340, which is used to perform various functions regarding encoding data, such as those functions previously described regarding the encoder component 106; and · A picture memory 352, which is used to store pictures and picture data, such as for use by the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0065] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described regarding the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described regarding the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described regarding the motion compensation prediction unit 260 and / or the intra prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described regarding the loop filter 256).
[0066] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described regarding the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described regarding the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, the shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0067] Each of the modules identified above stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., the instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.
[0068] Although Figure 3 server system 112 is shown in accordance with some embodiments, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems rather than a structural schematic of the embodiments described herein. In practice, and as will be appreciated by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately may be implemented on a single server, and a single item may also be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation, and optionally, in part, on the data traffic processed by the server system during peak usage periods as well as during average usage periods.
[0069] Figure 4 Example scenario 400 for generating a first chroma sample 402A based on a plurality of luminance samples 404 (e.g., 404A and 404X) in the MH-CCP mode is shown. In some embodiments, the current coded block 406C of the current image frame 408 is decoded in a cross-component intra prediction (CCIP) mode. In the CCIP mode, the decoder 122 ( Figure 2BDetermine each of a plurality of chrominance samples 402 of a current coding block 406C based on one or more reconstructed luminance samples 404. In some cases, the CCIP mode includes a cross-component linear model (CCLM) mode, in which, based on a linear model, a first chrominance sample 402A is transformed according to a reconstructed luminance sample 404A collocated with the chrominance sample 402A. Alternatively, in some cases, the CCIP mode includes a convolutional cross-component mode (CCCM), in which, based on a filter shape of a filter, a first chrominance sample 402A is directly predicted according to a plurality of reconstructed luminance samples 404X located near a first luminance sample 404A. Alternatively and additionally, in some cases, the CCIP mode includes an MH-CCP mode, in which a first chrominance sample 402A is generated by combining at least a first luminance sample 404A collocated with the first chrominance sample 402A and a plurality of hypothesis values by using a plurality of weighting factors. A plurality of adjacent luminance samples 404X of the first luminance sample 404A are combined by using a plurality of coefficients to generate a plurality of hypothesis values.
[0070] In other words, in some embodiments associated with the MH-CCP mode, a plurality of model parameters 410 (which are associated with weighting factors and coefficients) are used to combine a first luminance sample 404A and a plurality of adjacent luminance samples 404X to generate a first chrominance sample 402A. The first chrominance sample 402A is a blue-difference chrominance (Cb) sample or a red-difference chrominance (Cr) component. In some embodiments, the video bitstream 116 includes a current coding block 406C of a current image frame 408 and a first syntax element 420 for the MH-CCP mode. The first syntax element 420 indicates whether a first chrominance sample 402A of the current coding block 406C is reconstructed by combining a set of luminance samples 404 including the first luminance sample 404A based on a plurality of model parameters 410. In some embodiments, for the current coding block 406C, the first syntax element 420 is written in the video bitstream 116 at one of the following levels: block level, superblock level, image frame level, slice level, tile level, and image sequence level.
[0071] In some embodiments, the video bitstream 116 includes a first syntax element 420 for the MH-CCP mode. The first chrominance sample 402A of the current coding block 406C is configured to be generated by using a plurality of model parameters (e.g., w i 、w p 、w B) is generated by combining at least the first luminance sample 404A that is co-located with the first chrominance sample 402A and one or more adjacent luminance samples 404X of the first luminance sample 404. Based on determining that the MH-CCP mode is applied, the first chrominance sample 402A is predicted according to the following model: where predChromaVal is the predicted chrominance value of the first chrominance sample 402A; Num is the total number of adjacent luminance samples 404X; S i is the luminance value of the first luminance sample 404A (where i equals 0) or the adjacent luminance sample 404X (where i is greater than 0), which is indexed by i; P is the non-linear element 416; B is the offset term; and w i 、w p 、w B are model parameters. In one example, the non-linear element 416 (P) is equal to (C × C + B) >> bit_depth (bit depth), where bit_depth is the number of bits required to represent the luminance samples of the current image frame 408 during encoding and decoding. In some embodiments, B is the median luminance value, the intermediate luminance value, or the average luminance value of the luminance samples 404 of the current coding block 406C. In another example, B is equal to 1 << (bit_depth - 1).
[0072] In some embodiments, each of one or more adjacent luminance samples 404X of the first luminance sample 404A is adjacent to the first luminance sample 404A and shares at least one corresponding side or vertex with the first luminance sample 404A. In some embodiments, the one or more adjacent luminance samples 404X include a subset or all of the following: a north adjacent luminance sample (also referred to as a top luminance sample) 404N, a south adjacent luminance sample (also referred to as a bottom luminance sample) 404S, a west adjacent luminance sample (also referred to as a left side luminance sample) 404W, an east adjacent luminance sample (also referred to as a right side luminance sample) 404E, a northwestern adjacent luminance sample (also referred to as an upper left luminance sample) 404NW, a southeastern adjacent luminance sample (also referred to as a lower right luminance sample) 404SE, a southwestern adjacent luminance sample (also referred to as a lower left luminance sample) 404SW, and a northeastern adjacent luminance sample (also referred to as an upper right luminance sample) 404NE.
[0073] In some embodiments, in the MH-CCP mode, Equation (1) includes five terms and represents a five-tap model; the five-tap model is used to determine the first chrominance sample 402A of the current coding block 406C based on three linear terms (e.g., associated with the first luminance sample 404A and adjacent luminance samples 404W and 404E), a non-linear element 416(P), and an offset term B. Alternatively, in some embodiments, in the MH-CCP mode, Equation (1) includes seven terms and represents a seven-tap model; the seven-tap model is used to determine the first chrominance sample 402A of the current coding block 406C based on three linear terms (e.g., associated with luminance samples 404A, 404W, 404E, 404N, and 404S), a non-linear element 416(P), and an offset term B.
[0074] In some embodiments, the luminance samples 404 and chrominance samples 402 of the current coding block have different resolutions corresponding to a chrominance subsampling scheme (e.g., 4:2:2 or 4:2:0). Each luminance sample 404 includes a downsampled luminance sample generated from the reconstructed luminance samples using a downsampling filter. Alternatively, in some embodiments, each luminance sample 404 includes an original or reconstructed luminance sample without any downsampling. That is, the first luminance sample 404A is reconstructed at the resolution of the luminance samples or downsampled to the resolution of the chrominance samples. The adjacent luminance samples 404X (e.g., 404N, 404W, 404E, 404S, 404NW, 404NE, 404SW, 404SE) are also reconstructed at the resolution of the luminance samples or downsampled to the resolution of the chrominance samples.
[0075] In some embodiments, a plurality of model parameters w i 、w p and w B are determined based on a set of one or more reference luminance samples 404R and a set of one or more co-located reference chrominance samples 402R within the reference region 412 of the current coding block 406C. The reference region 412 is located within the current image frame 408. Additionally, in some embodiments, the reference luminance samples 404R of the reference region 412 are combined based on Equation (1) to regenerate one or more chrominance samples 402A. In some embodiments, a set of one or more co-located reference chrominance samples 402R is compared with one or more regenerated chrominance samples to generate a least mean square (LMS) value. The plurality of model parameters w i 、w p and w B are iteratively adjusted to reduce the LMS value until the LMS value meets a predetermined criterion (e.g., where the LMS value is below a threshold LMS value or is minimized).
[0076] In some embodiments, a plurality of model parameters w are derived based at least in part on chrominance samples and luminance samples within a reference region 412 of a current coding block 406C. i 、w p or w B . The reference region 412 includes one or more coding blocks that were decoded prior to the current coding block 406C (e.g., coding blocks 412LT, 412T, 412RT, 412L, and 412LB). In some embodiments, a subset of the one or more coding blocks is adjacent to the current coding block 406C. In some embodiments, a subset of the one or more coding blocks is separated from the current coding block 406C by one or more coding blocks. In some embodiments, the reference region 412 includes at least a portion of one or more rows above the current coding block 406C and / or at least a portion of one or more columns to the left of the current coding block 406C. For example, Figure 4 , the reference region 412 includes seven rows of luminance samples 404R above the current coding block 406C and nine columns of luminance samples 404R to the left of the current coding block 406C. The reference region 412 may include padded rows and padded columns (e.g., as shaded in Figure 4 ).
[0077] In some embodiments, the model represented by equation (1) includes a plurality of non-linear terms 414, which have: N + 1 non-linear terms. The non-linear prediction w p ·P is represented as follows: where P i represents the non-linear term 414 indexed by i and corresponds to the model parameter w Pi . After a first luminance sample 404A and one or more adjacent luminance samples 404X are identified, a plurality of non-linear terms 414 (P i ) are determined based at least on the first luminance sample 404A and a subset of the one or more adjacent luminance samples 404X. A first chrominance sample 402A that is collocated with the first luminance sample 404A is predicted in the current coding block based on the plurality of non-linear terms 414 (P i ). In some embodiments, one of the non-linear terms 414 (P i ) corresponds to the Mth power of a luminance sample, where M is an integer greater than 1. Additionally, in some embodiments, one of the non-linear terms 414 (P i ) corresponds to the Mth power of a single luminance sample. Alternatively, in some embodiments, the non-linear term 414 (P i) corresponds to the product of the mth power of the first luma sample and the nth power of the second luma sample, where m and n are positive integers, and the sum of m and n is equal to M. In one example, M is equal to 2. In another example, M is equal to 3.
[0078] In some embodiments, in addition to the first syntax element 420 associated with the MH-CCP mode, the video code stream 116 also includes a non-linear usage syntax element 418. The non-linear usage syntax element 418 indicates whether at least one non-linear term 414 is used in the MH-CCP mode or whether more than one non-linear term 414 is used in the MH-CCP mode.
[0079] Figure 5A An example process 500 is shown for predicting a first chroma sample 402A based on a plurality of non-linear terms 414 of one or more luma samples 404 in MH-CCP mode, in accordance with some embodiments. Figure 5B and Figure 5C are two example lookup tables 502 from which a plurality of model parameters 410 are identified according to some embodiments. Process 500 is performed on a computing system (specifically, Figure 1 The computing system receives a video bitstream 116 including a current coding block 406C of a current image frame 408. The video bitstream 116 includes a first syntax element 420 for an MH-CCP mode. Based on the first syntax element 420, it is determined to enable the MH-CCP mode to reconstruct a first chroma sample 402A of the current coding block 406C based on at least a first luma sample 404A co-located with the first chroma sample 402A and an associated adjacent luma sample 404X. The first luma sample 404A and one or more adjacent luma samples 404X are identified in the current coding block 406C and used to generate a plurality of non-linear terms 414. The first chroma sample 402A is predicted based on the plurality of non-linear terms 414 (e.g., based on equations (1) and (2)). The current image frame 408 including the current coding block 406C is reconstructed based on the first chroma sample 402A.
[0080] In some embodiments, each of the plurality of nonlinear terms 414 is one of the following: the square of the first brightness sample 404A (eg, S0 2 ), the square of each adjacent brightness sample 404X (eg, S i 2 , where i is a positive integer), the Mth power of the first brightness sample 404A (eg, S0 M ), where M is an integer greater than 2, the Mth power of each adjacent brightness sample 404X (e.g., S i M) The product of the first luminance sample 404A and each subset of one or more adjacent luminance samples 404X (e.g., S0S1, S0S1S2), and the product of each subset of two or more adjacent luminance samples 404X (e.g., S1S2, S1S2S3).
[0081] In some embodiments, the computing system combines multiple non - linear terms 414 and an offset term B to generate the first chrominance sample 402A, and excludes any linear terms S associated with the first luminance sample 404A and one or more adjacent luminance samples 404X when generating the first chrominance sample 402A i . For example, the multiple non - linear terms 414 and the offset term B are combined in a weighted manner to generate the first chrominance sample 402A in the following way: where the offset term B does not depend on any luminance sample, and the linear terms of the first luminance sample 404A or adjacent luminance samples 404X are not applied to determine the first chrominance sample 402A.
[0082] In some embodiments, the computing system combines multiple linear terms S i and multiple non - linear terms 414(P i ) to generate the first chrominance sample 402A, where the offset term B is not applied to determine the first chrominance sample 402A. Each linear term S i corresponds to a respective luminance sample selected from a set of luminance samples including the first luminance sample 404A and one or more adjacent luminance samples 404X. Additionally, in some embodiments, the multiple linear terms S i and multiple non - linear terms 414(P i ) are combined based on multiple model parameters 410. The multiple model parameters are based on a look - up table 502, and the look - up table 502 includes one or more valid model parameter values 510 for each of the linear terms S i and non - linear terms 414(P i ).
[0083] Additionally, in some embodiments, the video bitstream 116 includes a weight syntax element 504 that identifies a target weight index. Based on the weight syntax element 504, multiple model parameters 410 are selected from the look - up table 502. The look - up table 502 maps multiple weight indices to multiple sets of model parameters. For example, referring to Figure 5B , the multiple model parameters 410 can be addressed as a whole; the target weight index identifies one of the multiple model parameter combinations 506 in the look - up table 502 as the multiple model parameters 410. In another example, referring to Figure 5C, multiple model parameters 410 can be addressed in subsets; the lookup table 502 maps a first set of weight indices to values of a first subset 410A of the model parameters and maps a second set of weight indices to values of a second subset 410B of the model parameters. The target weight index includes two indicators for selecting values of the first subset of the model parameters and values of the second subset of the model parameters, respectively. Figure 5C The subsets 410A and 410B of the model parameters shown in Figure 5C are merely examples and are not intended to limit the combinations of the model parameters in each subset.
[0084] In some embodiments, the computing system identifies a reference region 412 corresponding to the current coding block 406C. Based on the reference samples 402R and 404R of the reference region 412, each of the multiple model parameters 410 is selected from one or more valid model parameter values in the lookup table 502 (e.g., denoted by "xx" in Figure 5B and Figure 5C ). Additionally, in some embodiments, a least mean square (LMS) value is determined based on the samples 402A and 404A of the reference region 412, and the multiple model parameters 410 are determined iteratively to reduce the LMS value until the LMS value meets a predetermined criterion. More details regarding the determination of the multiple model parameters 410 are explained above with reference to Figure 4 . Figure 5B and Figure 5C In some embodiments, after determining the multiple model parameters 410 based on the lookup table 502, the computing system applies at least one of a rounding operation 508, a shift operation 512, and a clipping operation 514 to one of the multiple model parameters 410 based on one or more predetermined weighting thresholds. Figure 4 are explained.
[0085] Reference Figure 5A Figure 5A , in some embodiments, after determining the multiple model parameters 410 based on the lookup table 502, the computing system applies at least one of a rounding operation 508, a shift operation 512, and a clipping operation to one of the multiple model parameters 410 based on one or more predetermined weighting thresholds.
[0086] In some embodiments, the computing system identifies a set 516 of predetermined non - linear terms based on the first luminance sample 404A and one or more adjacent luminance samples 404X, and selects multiple non - linear terms 414(P i ) from the set 516 of predetermined non - linear terms.
[0087] In some embodiments, the current image frame 408 also includes an alternative coding block 406A ( Figure 4 ) different from the current coding block 406C. When the MH - CCP mode is enabled for the alternative coding block 406A, the alternative chrominance samples are located in the alternative coding block 406A and are predicted based on a plurality of alternative linear terms S i associated with the alternative luminance sample 404A', where the alternative luminance sample 404A' is co - located with the alternative chrominance samples. When determining the MH - CCP mode, any non - linear terms 414(Pi )。In other words, multiple linear terms S i and the offset term B are combined in a weighted manner to generate an alternative chroma sample as follows: where predChromaVal’ is the predicted chroma value of the alternative chroma sample, and the offset term B does not depend on any luma samples. The non - linear terms 414(P i ) of the alternative luma sample 404A’ or adjacent luma samples are not applied to determine the alternative chroma sample.
[0088] Reference Figure 5A , in some embodiments, the computing system identifies a set 516 of predetermined non - linear terms and a set 518 of predetermined linear terms based on a first luma sample 404A and one or more adjacent luma samples 404X. The video bitstream includes a second syntax element 520. When the second syntax element 520 has a first value indicating a first combination, multiple linear terms S i are selected from the set 518 of predetermined linear terms, and multiple non - linear terms 414(P i ) are selected from the set 516 of predetermined non - linear terms. When the second syntax element 520 has a second value indicating a second combination, multiple linear terms S i are selected from the set 518 of predetermined linear terms, and the non - linear terms 414(P i ) are not selected from the set 516 of predetermined non - linear terms. When the second syntax element 520 has a third value indicating a third combination, the linear terms S i are not selected from the set 518 of predetermined linear terms, and multiple non - linear terms 414(P i ) are selected from the set 516 of predetermined non - linear terms. Multiple non - linear terms 414(P i ) are combined, multiple linear terms S i are combined, or multiple non - linear terms 414(P i ) and multiple linear terms S i are combined to generate the first chroma sample 402A. Additionally, in some embodiments, the second syntax element 520 is written at the following levels: coding block level, super - block level, tile level, slice level, frame level, or picture sequence level.
[0089] Figure 6 shows another example process 600 for predicting the first chroma sample 402A based on multiple non - linear terms 414(P i ) of one or more luma samples 404 in the MH - CCP mode. Process 600 is in a computing system (specifically, Figure 1It is implemented in the decoder 122 of the electronic device 120. The computing system receives a video bitstream 116, which includes a current encoded block 406C of the current image frame 408. And the computing system identifies a first luma sample 404A and one or more neighboring luma samples 404X in the current encoded block 406C. A plurality of non-linear terms 414 are generated based on the first luma sample 404A and the one or more neighboring luma samples 404X. A first chroma sample 402A is predicted based on the plurality of non-linear terms 414 (e.g., based on equations (1) and (2)).
[0090] In some embodiments, the luma samples 404 of the current image frame 408 correspond to a full range 602. If each of the luma samples 404 has 8 bits, the full range 602 can be from 0 to 255. In one example, the full range 602 is defined by the minimum and maximum values of the luma samples 404 of the current image frame 408. The full range 602 of the luma samples 404 of the current image frame 408 is divided into a plurality of regions 604. For the first non-linear term 410-1 (e.g., P1), the computing system determines a corresponding target luma sample 404T based on the first luma sample 404A and the one or more neighboring luma samples 404X. The corresponding target luma sample 404T corresponds to one of the plurality of regions 604. Based on the region of the corresponding target luma sample 404T, a subset of the predetermined functions 608 is selected to form a non-linear function 606. For example, when the corresponding target luma sample 404T is in the range from 0 to 127, the non-linear function 606 is sin(S0), where S0 is the first luma sample 404A. When the corresponding target luma sample 404T is in the range from 128 to 255, the non-linear function 606 is cos(S0). In some embodiments, the non-linear function 606 also includes a combination of a subset of the predetermined functions 608 and a linear function. For example, the non-linear function 606 is expressed as sin(S0)+S0, which is applied to the corresponding target luma sample 404T to generate the first non-linear term 410-1.
[0091] In some embodiments, for each non-linear term 414 (P i) The computing system determines a corresponding target luminance sample 404T based on the first luminance sample 404A and one or more adjacent luminance samples 404X. In some embodiments, the target luminance sample 404T is selected from the first luminance sample 404A and one or more adjacent luminance samples 404X. Alternatively, in some embodiments, the corresponding target luminance sample 404T is determined based on the difference between the first luminance sample 404A and one of the one or more adjacent luminance samples 404X. In one example, the corresponding target luminance sample 404T represents the difference between the first luminance sample 404A and one of the one or more adjacent luminance samples 404X.
[0092] In some embodiments, the computing system applies at least one of the following functions to the corresponding target luminance sample 404T: Sigmoid function 608A, hyperbolic function 608B, cosine function 608C, sine function 608D, exponential function 608E, and cubic function 608F. In one example, one of the multiple non - linear terms 414(P i ) is the Sigmoid function 608A of the luminance sample 404 (e.g., S i ), expressed as follows: In another example, one of the multiple non - linear terms 414(P i ) is the hyperbolic function 608B of the target luminance sample 404T (e.g., S i ), expressed as follows: In yet another example, one of the multiple non - linear terms 414(P i ) is expressed as follows: where S0 is the first luminance sample 404A and S1 is the associated adjacent luminance sample 404X (e.g., the left - hand luminance sample 404W). In some embodiments, one of the functions 608A to 608F is applied to the corresponding target luminance sample 404T (e.g., the first luminance sample 404A, the corresponding target luminance sample 404T).
[0093] Figure 7is a flowchart showing an example method 700 for decoding video according to some embodiments. Method 700 may be performed in a computing system having a control circuit and a memory (e.g., server system 112, source device 102, or electronic device 120), the memory storing instructions for execution by the control circuit. In some embodiments, method 700 is performed by executing instructions stored in the memory of the computing system (e.g., memory 314). In some embodiments, method 700 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, a weighted sum of multiple versions of the co-located luma samples 404A ( Figure 4 ) is used to predict the chroma value 402A in the MH-CCP mode. The multiple versions of the co-located luma samples 404A are derived based on the co-located luma samples 404A or based on the filtered co-located luma samples obtained using the adjacent luma samples 404X (e.g., 404W, 404N, 404E, 404S, 404NW, 404NE, 404SW, 404SE) as the filtering input. Each input to the weighted sum (e.g., the respective versions of the co-located luma samples) is referred to as a hypothesis. In some cases, the values of the reference luma sample 404R and the reference chroma sample 402R are fed into a least mean square calculation kernel to derive the model parameters 410 used in the MH-CCP mode. In some embodiments, the first luma sample 404A has eight adjacent samples 404X ( Figure 4 ) adjacent to the first luma sample 404A. Additionally, in some embodiments, when the luma and chroma have different dimensions (e.g., 4:2:2 vs. 4:2:0), each luma sample 404 (e.g., 404A, 404X) includes a downsampled luma sample generated using a downsampling filter. Alternatively, each luma sample 404 includes an original co-located luma sample without any downsampling.
[0094] According to method 700, in the MH-CCP mode, multiple non-linear terms 414 of the luma samples 404 are determined and applied to generate the first chroma sample 402A. The non-linear terms 414 can be derived based on the co-located sample of the first color component (e.g., the first luma sample 404A) and one or more associated adjacent samples of the co-located sample (e.g., the adjacent luma samples 404X). The sample of the second color component (e.g., the first chroma sample 402A) corresponds to the co-located sample of the first color component (e.g., the first luma sample 404A) and is determined using Equation (1). In some embodiments, w pi is the model parameter of the non-linear term, and P iDenotes the non - linear term 414 indexed by i. In some embodiments, each non - linear term 414(P i ) may include: the m - th power of adjacent samples of the first color component (e.g., S1 m , where m is greater than 1), the n - th power of co - located samples of the first color component (e.g., S o n , where n is greater than 1), the product of a co - located sample and at least one adjacent sample, or the product of two or more different adjacent samples. The non - linear element 416 is the sum of a plurality of non - linear terms 414 and is represented in Equation (2).
[0095] In other words, in some embodiments, a plurality of samples of the first color component (e.g., a plurality of luminance samples 404) may be used to derive each non - linear term 414, the plurality of samples including one or more co - located samples of the first color component, one or more adjacent samples of the co - located samples, or their products. An example of the obtained non - linear element 416 is determined based on Equation (2), where P i denotes the non - linear term 414 indexed by i and corresponds to the model parameter w Pi .
[0096] In some embodiments, the predicted sample of the second color component (e.g., the first chrominance sample 402A) may be a combination of a plurality of non - linear terms 414(P i ) and an offset term B. For example, the plurality of non - linear terms 414 and the offset term B are combined in a weighted manner to generate the predicted sample of the second color component (e.g., the first chrominance sample 402A) based on Equation (3).
[0097] In some embodiments, the predicted sample of the second color component (e.g., the first chrominance sample 402A) may be a combination of a plurality of linear terms S i , a plurality of non - linear terms 414(P i ) and an offset term B. In one example, when the MH - CCP mode is enabled, the first chrominance sample 402A is predicted according to Equation (1).
[0098] In some embodiments, a first combination of the non - linear term 414(P i ) and the linear term S i is used for the first set of coded blocks of the current image frame 408. Additionally, in some embodiments, in the second set of coded blocks of the current image frame, each sample of the second color component is determined based on a second combination of a subset of all valid linear terms 518( Figure 5A ) and the offset term B (excluding any non - linear terms 414). Further, in some embodiments, in the third set of coded blocks of the current image frame 408, based on all valid predetermined non - linear terms 516( Figure 5A) subset of and the third combination of offset term B (excluding any linear term S i ) to determine each sample of the second color component. One or more syntax elements (e.g., the second syntax element 520) are written to indicate which one of the first combination, the second combination, and the third combination is used for each associated coding block. The one or more syntax elements can be written at the following levels: coding block level, superblock level, tile level, picture slice level, picture frame level, or picture sequence level.
[0099] In some embodiments, one or more non - linear functions are used to generate non - linear term 414 associated with the predicted sample of the second color component. For example, according to the luminance sample 404 (e.g., S i ) the sigmoid function (as shown in Equation (5)) determines non - linear term 414(P i ). In one example, according to the luminance sample 404 (e.g., S i ) the hyperbolic function (as shown in Equation (6)) determines non - linear term 414(P i ). In one example, at least one of a cosine function, a sine function, an exponential function, and a cubic function is used to determine non - linear term 414(P i ). In one example, the input of non - linear term 414(P i ) is the difference between the co - located sample and an adjacent sample (e.g., the difference between the first luminance sample 404A and the left - hand luminance sample 404W). In another example, a piece - wise function is used based on the range of the input (also referred to as the target luminance sample 404T( Figure 6 ).
[0100] In some embodiments, the model parameters 410 associated with different terms of the MH - CCP mode are written, derived from multiple predetermined options, or derived based on the reference region 412. At least the model parameter w i associated with non - linear term 414(P pi ) can be stored in the look - up table 502. The target weight index is used to identify one or more model parameter values stored in the look - up table 502. Multiple model parameters 410 are extracted from the look - up table 502 for determining the samples of the second color component. In one example, the look - up table 502 is constructed according to one or more non - linear functions and can be rounded (operation 508), shifted (operation 512), or clipped (operation 514) using a predetermined threshold.
[0101] In some embodiments, the first non - linear usage syntax element (e.g., Figure 4 the syntax element 418 in) indicates whether at least one non - linear term 414(P i)。The first non - linear usage syntax element can be written using advanced syntax, which includes but is not limited to: sequence - level flags, picture - level flags, sub - picture - level flags, slice - level flags, or tile - level flags. In some embodiments, the second non - linear usage syntax element (e.g., Figure 4 the syntax element 418 in) indicates whether more than one non - linear term 414 (P i ) can be used in the MH - CCP mode. The second non - linear usage syntax element can be written using advanced syntax, which includes but is not limited to: sequence - level flags, picture - level flags, sub - picture - level flags, slice - level flags, or tile - level flags.
[0102] Although Figure 7 a plurality of logical stages are shown in a particular order, the logical stages that are not order - dependent can be reordered, and other stages can be combined or decomposed. For those of ordinary skill in the art, some reorderings or other groupings not specifically mentioned will be obvious, so the orderings and groupings presented herein are not exhaustive. In addition, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0103] Now turning to some example embodiments.
[0104] (A1) In some implementations, a method 700 for decoding video data is implemented. Method 700 includes (at operation 702) receiving a video bitstream that includes a current encoded block of a current image frame, where the video bitstream includes a first syntax element for a multi - hypothesis cross - component prediction (MH - CCP) mode; (at operation 704) based on the first syntax element, determining to enable the MH - CCP mode to reconstruct a first chrominance sample of the current encoded block at least based on a first luminance sample and associated adjacent luminance samples, the first luminance sample being co - located with the first chrominance sample; (at operation 706) identifying the first luminance sample and one or more adjacent luminance samples in the current encoded block; (at operation 708) generating a plurality of non - linear terms based at least on the first luminance sample and a subset of one or more adjacent luminance samples; (at operation 710) predicting the first chrominance sample co - located with the first luminance sample in the current encoded block based on the plurality of non - linear terms; and (at operation 712) reconstructing the current image frame including the current encoded block.
[0105] (A2) In some embodiments of A1, each of the plurality of non - linear terms includes one of the following: the square of the first luminance sample; the square of each adjacent luminance sample; the M - th power of the first luminance sample, where M is an integer greater than 2; the M - th power of each adjacent luminance sample; the product of the first luminance sample and each subset of one or more adjacent luminance samples; and the product of each subset of two or more adjacent luminance samples.
[0106] (A3) In some embodiments of A1 or A2, predicting the first chrominance sample further includes combining a plurality of non - linear terms and an offset term to generate the first chrominance sample.
[0107] (A4) In some embodiments of any one of A1 to A3, predicting the first chrominance sample further includes: combining a plurality of linear terms and a plurality of non - linear terms to generate the first chrominance sample, each linear term corresponding to a respective luminance sample selected from a set of luminance samples including the first luminance sample and one or more adjacent luminance samples.
[0108] (A5) In some embodiments of A4, combining the plurality of linear terms and the plurality of non - linear terms is based on a plurality of model parameters, and the method further includes determining the plurality of model parameters based on a look - up table that includes one or more valid model parameter values for each of the linear terms and the non - linear terms.
[0109] (A6) In some embodiments of A5, the video bitstream includes a weight syntax element that identifies a target weight index, and determining the plurality of model parameters further includes: selecting the plurality of model parameters from the look - up table based on the weight syntax element, the look - up table mapping a plurality of weight indices to a plurality of sets of model parameters.
[0110] (A7) In some embodiments of A5, determining the plurality of model parameters further includes: identifying a reference region corresponding to the current coding block; and
[0111] selecting each of the plurality of model parameters from one or more valid model parameter values in the look - up table based on samples in the reference region.
[0112] (A8) In some embodiments of A7, selecting each of the plurality of model parameters further includes: determining a least mean square (LMS) value based on samples in the reference region; wherein the plurality of model parameters are determined iteratively to reduce the LMS value until the LMS value meets a predetermined criterion.
[0113] (A9) In some embodiments of any one of A5 to A8, method 700 further includes, after determining the plurality of model parameters based on the look - up table, applying a rounding, shifting, or clipping operation to one of the plurality of model parameters based on one or more predetermined weighted thresholds.
[0114] (A10) In some embodiments of any one of A1 to A9, generating the plurality of non - linear terms further includes: identifying a set of predetermined non - linear terms based on the first luminance sample and one or more adjacent luminance samples; and selecting the plurality of non - linear terms from the set of predetermined non - linear terms.
[0115] (A11)In some embodiments of any one of A1 to A10, predicting the first chrominance sample further includes: combining a plurality of non-linear terms and an offset term to generate the first chrominance sample; and excluding any linear terms associated with the first luminance sample and one or more adjacent luminance samples when generating the first chrominance sample.
[0116] (A12)In some embodiments of any one of A1 to A11, the current image frame includes an alternative coding block different from the current coding block, and the method further includes: when the MH-CCP mode is enabled for the alternative coding block: predicting an alternative chrominance sample in the alternative coding block based on a plurality of alternative linear terms associated with alternative luminance samples, the alternative luminance samples being co-located with the alternative chrominance sample, the prediction including excluding any non-linear terms associated with the alternative luminance samples.
[0117] (A13)In some embodiments of any one of A1 to A12, the method 700 further includes identifying a set of predetermined non-linear terms and a set of predetermined linear terms based on the first luminance sample and one or more adjacent luminance samples; wherein the video bitstream includes a second syntax element; wherein, based on the second syntax element, a plurality of linear terms are selected from the set of predetermined linear terms, and a plurality of non-linear terms are selected from the set of predetermined non-linear terms; and wherein predicting the first chrominance sample further includes combining the plurality of non-linear terms and the plurality of linear terms to generate the first chrominance sample.
[0118] (A14)In some embodiments of A13, the second syntax element is written at one of the following levels: coding block level, superblock level, tile level, slice level, frame level, or picture sequence level.
[0119] (A15)In some embodiments of any one of A1 to A14, generating the plurality of non-linear terms further includes, for each non-linear term: determining a corresponding target luminance sample based on the first luminance sample and one or more adjacent luminance samples; and applying at least one of a sigmoid function, a hyperbolic function, a cosine function, a sine function, an exponential function, and a cubic function to the corresponding target luminance sample.
[0120] (A16)In some embodiments of A15, determining the corresponding target luminance sample further includes: selecting the corresponding target luminance sample from the first luminance sample and one or more adjacent luminance samples.
[0121] (A17)In some embodiments of A15, determining the corresponding target luminance sample further includes: determining the corresponding target luminance sample based on a difference between the first luminance sample and one of the one or more adjacent luminance samples.
[0122] (A18)In some embodiments of any one of A1 to A17, the method further includes: dividing the full range of luminance samples of the current image frame into a plurality of regions; wherein, generating a plurality of non-linear terms further includes, for the first non-linear term: determining a corresponding target luminance sample based on the first luminance sample and one or more adjacent luminance samples, selecting a subset of a predetermined function based on the region of the corresponding target luminance sample, and applying the subset of the predetermined function to the corresponding target luminance sample to generate the first non-linear term.
[0123] (A19)In some embodiments of any one of A1 to A18, for a current coding block, in a video bitstream, a first non-linear usage syntax element is written at one of the following levels: block level, super-block level, image frame level, slice level, tile level, and image sequence level, the non-linear usage syntax element indicating whether at least one non-linear term is used in the MH-CCP mode.
[0124] (A20)In some embodiments of any one of A1 to A19, in addition to the first syntax element associated with the MH-CCP mode, the video bitstream further includes a second non-linear usage syntax element, the second non-linear usage syntax element indicating whether more than one non-linear term is used in the MH-CCP mode.
[0125] (A21)In some embodiments, a method includes: receiving video data, the video data including a current coding block of a current image frame; encoding the current image frame; transmitting the encoded current image frame via a video bitstream; and writing, via the video bitstream, a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct a first chrominance sample of the current coding block based on a first luminance sample and an associated adjacent luminance sample, the first luminance sample being co-located with the first chrominance sample; wherein, when the MH-CCP mode is enabled, a plurality of non-linear terms are determined based at least on the first luminance sample and a subset of one or more adjacent luminance samples of the first luminance sample, and the first chrominance sample co-located with the first luminance sample is predicted based on the plurality of non-linear terms.
[0126] (A22)In some embodiments of A21, the method is implemented to implement the features of any one of A2 to A20.
[0127] (A23) In some embodiments, a method includes: obtaining a source video sequence including a current image frame having a current coding block; and performing a conversion between the source video sequence and a video bitstream. The video bitstream includes: a current image frame having the current coding block; and a first syntax element for a multi-hypothesis cross-component prediction (MH-CCP) mode, the first syntax element indicating whether to reconstruct a first chrominance sample of the current coding block based on a first luma sample and associated neighboring luma samples, the first luma sample being co-located with the first chrominance sample. When the MH-CCP mode is enabled, a plurality of non-linear terms are determined based at least on the first luma sample and a subset of one or more neighboring luma samples of the first luma sample, and the first chrominance sample co-located with the first luma sample is predicted based on the plurality of non-linear terms.
[0128] (A24) In some embodiments of A23, the method is implemented to achieve the features of any one of A2 to A20.
[0129] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit. The memory stores one or more instruction sets configured to be executed by the control circuit, the one or more instruction sets including instructions for performing any one of the various methods described herein (e.g., A1 to A24 above).
[0130] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets for execution by a control circuit of a computing system, the one or more instruction sets including instructions for performing any one of the various methods described herein (e.g., A1 to A24 above).
[0131] Unless otherwise specified, any syntax element described herein may be high-level syntax (HLS). As used herein, HLS is written at a level higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements may be written in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, picture header, tile header, and / or CTU header.
[0132] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. It should also be understood that the term "and / or" as used herein refers to any and all possible combinations of one or more of the associated listed items, and encompasses any and all possible combinations of one or more of the associated listed items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specifically state the presence of the features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0133] As used herein, depending on the context, the term "if" can be interpreted to mean "when the precondition is true" or "while the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "in response to detecting that the precondition is true". Similarly, depending on the context, the expressions "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" can be interpreted to mean "when it is determined that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "when detecting that the precondition is true" or "in response to detecting that the precondition is true".
[0134] For purposes of explanation, the above description has been set forth with reference to specific embodiments. However, the foregoing illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best explain the working principles and practical applications, thereby enabling others skilled in the art to understand.
Claims
1. A method for video decoding, characterized in that, Comprising: Receiving a video bitstream, the video bitstream including a current coded block of a current image frame, wherein the video bitstream includes a first syntax element for a multi-hypothesis cross-component prediction mode; Based on the first syntax element, determining to enable the multi-hypothesis cross-component prediction mode to reconstruct a first chrominance sample of the current coded block based at least on a first luminance sample and associated neighboring luminance samples, the first luminance sample being collocated with the first chrominance sample; Identifying the first luminance sample and one or more neighboring luminance samples in the current coded block; Generating a plurality of non-linear terms based at least on the first luminance sample and a subset of the one or more neighboring luminance samples; Predicting the first chrominance sample collocated with the first luminance sample in the current coded block based on the plurality of non-linear terms; and Reconstructing the current image frame including the current coded block.
2. The method according to claim 1, wherein Each of the plurality of non-linear terms includes one of the following: The square of the first luminance sample; The square of each neighboring luminance sample; The M-th power of the first luminance sample, where M is an integer greater than 2; The M-th power of each neighboring luminance sample; The product of the first luminance sample and each subset of one or more neighboring luminance samples; And The product of each subset of two or more neighboring luminance samples.
3. The method according to claim 1, characterized in that, Predicting the first chrominance sample further includes: Combining the plurality of non-linear terms and an offset term to generate the first chrominance sample.
4. The method according to claim 1, wherein Predicting the first chrominance sample further includes: Combining a plurality of linear terms and the plurality of non-linear terms to generate the first chrominance sample, each linear term corresponding to a respective luminance sample selected from a set of luminance samples including the first luminance sample and the one or more neighboring luminance samples; Wherein, the plurality of linear terms and the plurality of non-linear terms are combined based on a plurality of model parameters, and the method further includes: determining the plurality of model parameters based on a look-up table, the look-up table including one or more valid model parameter values for each of the linear terms and the non-linear terms.
5. The method according to claim 4, wherein The video bitstream includes a weight syntax element identifying a target weight index, and determining the plurality of model parameters further includes: Based on the weight syntax element, selecting the plurality of model parameters from the look-up table, the look-up table mapping a plurality of weight indices to a plurality of sets of model parameters.
6. The method according to claim 4, wherein Determining the plurality of model parameters further includes: Identifying a reference region corresponding to the current coded block; and Based on the samples of the reference region, selecting each of the plurality of model parameters from one or more valid model parameter values in the look-up table; Wherein, selecting each of the plurality of model parameters further includes: Determining a minimum mean square value based on the samples of the reference region; Wherein, the plurality of model parameters are iteratively determined to reduce the minimum mean square value until the minimum mean square value meets a predetermined criterion.
7. The method according to claim 4, characterized in that, Further including, after determining the plurality of model parameters based on the look-up table, applying a rounding, shifting or clipping operation to one of the plurality of model parameters based on one or more predetermined weighting thresholds.
8. The method according to claim 1, characterized in that, Generating the plurality of non-linear terms further includes: Identifying a set of predetermined non-linear terms based on the first luminance sample and the one or more adjacent luminance samples; and Selecting the plurality of non-linear terms from the set of predetermined non-linear terms.
9. The method according to claim 1, characterized in that, Predicting the first chrominance sample further comprises:[[]] Combining the plurality of non-linear terms and an offset term to generate the first chrominance sample; and Excluding any linear terms associated with the first luminance sample and the one or more adjacent luminance samples when generating the first chrominance sample.
10. The method according to claim 1, wherein The current image frame includes an alternative coding block different from the current coding block, and the method further comprises: when the multi-hypothesis cross-component prediction mode is enabled for the alternative coding block: Predicting an alternative chrominance sample in the alternative coding block based on a plurality of alternative linear terms associated with alternative luminance samples, the alternative luminance samples being co-located with the alternative chrominance sample, the prediction including excluding any non-linear terms associated with the alternative luminance samples.
11. The method according to claim 1, characterized in that Further comprising: Identifying a set of predetermined non-linear terms and a set of predetermined linear terms based on the first luminance sample and the one or more adjacent luminance samples; wherein, the video bitstream includes a second syntax element; wherein, based on the second syntax element, a plurality of linear terms are selected from the set of predetermined linear terms, and the plurality of non-linear terms are selected from the set of predetermined non-linear terms; and wherein, predicting the first chrominance sample further comprises: combining the plurality of non-linear terms and the plurality of linear terms to generate the first chrominance sample; wherein, the second syntax element is written at one of the following levels: coding block level, super-block level, tile level, slice level, frame level or picture sequence level.
12. The method according to claim 1, wherein Generating the plurality of non-linear terms further comprises, for each non-linear term: Determining a corresponding target luminance sample based on the first luminance sample and the one or more adjacent luminance samples; and Applying at least one of a sigmoid function, a hyperbolic function, a cosine function, a sine function, an exponential function, and a cubic function to the corresponding target luminance sample.
13. The method according to claim 12, wherein Determining the corresponding target luminance sample further comprises: Selecting the corresponding target luminance sample from the first luminance sample and the one or more adjacent luminance samples; or Determining the corresponding target luminance sample based on a difference between the first luminance sample and one of the one or more adjacent luminance samples.
14. The method according to claim 1, wherein Further comprising dividing the full range of luminance samples of the current image frame into a plurality of regions, wherein generating the plurality of non-linear terms further comprises, for a first non-linear term: Determining a corresponding target luminance sample based on the first luminance sample and the one or more adjacent luminance samples; Selecting a subset of predetermined functions based on the region of the corresponding target luminance sample; and Applying the subset of predetermined functions to the corresponding target luminance sample to generate the first non-linear term.
15. The method according to any one of claims 1 to 14, characterized in that, For the current coding block, in the video bitstream, write a first non-linear usage syntax element at one of the following levels: block level, super-block level, picture level, slice level, tile level, and picture sequence level, where the first non-linear usage syntax element indicates whether at least one non-linear term is used in the multi-hypothesis cross-component prediction mode.
16. The method according to any one of claims 1 to 14, characterized in that, In addition to the first syntax element associated with the multi-hypothesis cross-component prediction mode, the video bitstream further includes a second non-linear usage syntax element, where the second non-linear usage syntax element indicates whether more than one non-linear term is used in the multi-hypothesis cross-component prediction mode.
17. A method for video encoding, characterized in that, Comprising: Receiving video data, where the video data includes a current coding block of a current picture; Encoding the current picture; Transmitting the encoded current picture via the video bitstream; And Writing, via the video bitstream, a first syntax element for the multi-hypothesis cross-component prediction mode, where the first syntax element indicates whether a first chroma sample of the current coding block is reconstructed based on a first luma sample and associated neighboring luma samples, and the first luma sample is co-located with the first chroma sample; Wherein, when the multi-hypothesis cross-component prediction mode is enabled, a plurality of non-linear terms are determined based at least on the first luma sample and a subset of one or more neighboring luma samples of the first luma sample, and the first chroma sample co-located with the first luma sample is predicted based on the plurality of non-linear terms.
18. A bitstream conversion method, characterized in that, Comprising: Obtaining a source video sequence including a current picture having a current coding block; And Performing a conversion between the source video sequence and the video bitstream, where the video bitstream includes: The current picture having the current coding block; and A first syntax element for the multi-hypothesis cross-component prediction mode, where the first syntax element indicates whether a first chroma sample of the current coding block is reconstructed based on a first luma sample and associated neighboring luma samples, and the first luma sample is co-located with the first chroma sample; Wherein, when the multi-hypothesis cross-component prediction mode is enabled, a plurality of non-linear terms are determined based at least on the first luma sample and a subset of one or more neighboring luma samples of the first luma sample, and the first chroma sample co-located with the first luma sample is predicted based on the plurality of non-linear terms.
19. A video decoding device, characterized in that, Comprising: A receiving module configured to receive a video bitstream, where the video bitstream includes a current coding block of a current picture, and where the video bitstream includes a first syntax element for the multi-hypothesis cross-component prediction mode; A determining module configured to determine, based on the first syntax element, that the multi-hypothesis cross-component prediction mode is enabled to reconstruct a first chroma sample of the current coding block based at least on a first luma sample and associated neighboring luma samples, and the first luma sample is co-located with the first chroma sample; An identifying module configured to identify the first luma sample and one or more neighboring luma samples in the current coding block; A generating module configured to generate a plurality of non-linear terms based at least on the first luma sample and a subset of the one or more neighboring luma samples; A prediction module, configured to predict the first chroma sample co-located with the first luma sample in the current coded block based on the plurality of non-linear terms; and A reconstruction module, configured to reconstruct the current picture frame including the current coded block.
20. A video encoding device, characterized in that, Comprising: A receiving module, configured to receive video data, the video data including a current coded block of a current picture frame; An encoding module, configured to encode the current picture frame; A transmission module, configured to transmit the encoded current picture frame via a video bitstream; And A writing module, configured to write, via the video bitstream, a first syntax element for a multi-hypothesis cross-component prediction mode, the first syntax element indicating whether to reconstruct the first chroma sample of the current coded block based on a first luma sample and associated neighboring luma samples, the first luma sample being co-located with the first chroma sample; Wherein, when the multi-hypothesis cross-component prediction mode is enabled, a plurality of non-linear terms are determined based on at least the first luma sample and a subset of one or more neighboring luma samples of the first luma sample, and the first chroma sample co-located with the first luma sample is predicted based on the plurality of non-linear terms.
21. A computing system, characterized in that, Comprising: A control circuit; And A memory storing one or more programs, the one or more programs being configured to be executed by the control circuit, the one or more programs further including instructions for performing the following operations: Receiving video data, the video data including a current coded block of a current picture frame; Encoding the current picture frame; Transmitting the encoded current picture frame via a video bitstream; And Writing, via the video bitstream, a first syntax element for a multi-hypothesis cross-component prediction mode, the first syntax element indicating whether to reconstruct the first chroma sample of the current coded block based on a first luma sample and associated neighboring luma samples, the first luma sample being co-located with the first chroma sample; Wherein, when the multi-hypothesis cross-component prediction mode is enabled, a plurality of non-linear terms are determined based on at least the first luma sample and a subset of one or more neighboring luma samples of the first luma sample, and the first chroma sample co-located with the first luma sample is predicted based on the plurality of non-linear terms.
22. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs being executed by at least one processor to perform the method according to any one of claims 1 to 16, or to perform the method according to claim 17, or to perform the method according to claim 18.