Apparatus, method and computer program for video encoding and decoding

By using the gradient and location values ​​of spatial samples in the cross-component prediction model, the correlation modeling of the luminance and chrominance channels is enhanced, which solves the coding efficiency problem of existing models when the correlation is insufficient, and improves the efficiency and quality of video coding.

CN120345247APending Publication Date: 2025-07-18NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082656.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-08-29
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing cross-component prediction models are inefficient in modeling the correlation between luminance and chrominance channels, especially when spatial correlation is insufficient, resulting in poor coding efficiency.

Method used

By incorporating the gradient and position values ​​of spatial samples into the cross-component prediction model to predict target samples for the chroma channel, the correlation modeling between luminance and chroma channels is enhanced.

Benefits of technology

It improves the efficiency and quality of video encoding, especially in cases where the correlation between the luminance and chrominance channels is insufficient, thereby enhancing encoding efficiency and compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345247A_ABST
    Figure CN120345247A_ABST
Patent Text Reader

Abstract

A method comprises: receiving an image block unit of a frame, the image block unit comprising samples in a color channel, where the color channel comprises at least one chroma channel and one luma channel (1200); reconstructing a sample of the luminance channel of an image block unit (1202); a reference region for predicting a target sample of at least one color channel of an image block unit is determined, where the reference region comprises one or more of reference samples in adjacent blocks in a current color channel / frame, in adjacent blocks of co-located blocks in a reference color channel / frame, and in adjacent blocks of co-located blocks in a reference color channel / frame, in a reference region of the target sample in the current color channel / frame. And / or within a collocated block in the reference color channel / frame (1204); determining one or more gradient values and / or one or more position values of spatial samples in the reference region used in a cross-component prediction model (1206); and predicting the target sample of at least one color channel of an image block unit using a cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of spatial samples in the reference region (1208).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus, method, and computer program for video encoding and decoding. Background Art

[0002] In video encoding, a color representation (such as YUV or YCbCr) consisting of one luma and two chroma channels is typically used to encode video and image samples. In these cases, the luma channel, which mainly represents scene illumination, is usually encoded at certain resolutions, while the chroma channels, which usually represent the differences between certain color components, are often encoded at a second resolution lower than the resolution of the luma signal. The purpose of this differential representation is to decorrelate the color components and enable more efficient data compression.

[0003] The Cross-Component Linear Model (CCLM) and the Convolutional Cross-Component Model (CCCM) are used to predict samples (e.g., Cb and Cr) in the chroma channels by cross-channel correlation (e.g., using luma samples). The model parameters are derived based on the reconstructed samples in the neighborhood of the chroma block, the neighboring samples at the same position in the luma block, and the reconstructed samples within the luma block at the same position.

[0004] Since cross-component prediction models (e.g., CCLM and CCCM) use spatial samples to derive the model parameters, they may not always perform optimally in some content when spatial correlation does not exist or it is not the main factor in cross-channel correlation. Summary of the Invention

[0005] Now, in order to at least mitigate the above problems, an enhanced method for better modeling cross-channel correlation in a cross-component model is introduced herein.

[0006] The scope of protection sought by various embodiments of the present invention is given by the independent claims. Embodiments and features (if any) described in this specification that do not fall within the scope of the independent claims will be construed as useful examples for understanding the various embodiments of the present invention.

[0007] The method according to the first aspect includes receiving a picture block unit of a frame, the picture block unit including samples in a color channel, where the color channel includes at least one chrominance channel and one luminance channel; reconstructing the samples of the luminance channel of the picture block unit; determining a reference region for predicting target samples of at least one color channel of the picture block unit, where the reference region includes one or more of the following reference samples, the reference samples being in adjacent blocks in the current color channel / frame, in adjacent blocks of the co-located blocks in the reference color channel / frame, and / or within the co-located blocks in the reference color channel / frame; determining one or more gradient values and / or one or more position values of the spatial samples in the reference region to be used in a cross-component prediction model; and predicting the target samples of at least one color channel of the picture block unit using the cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

[0008] The apparatus according to the second aspect includes means for receiving a picture block unit of a frame, the picture block unit including samples in a color channel, where the color channel includes at least one chrominance channel and one luminance channel; means for reconstructing the samples of the luminance channel of the picture block unit; means for determining a reference region for predicting target samples of at least one color channel of the picture block unit, where the reference region includes one or more of the following reference samples, the reference samples being in adjacent blocks in the current color channel / frame, in adjacent blocks of the co-located blocks in the reference color channel / frame, and / or within the co-located blocks in the reference color channel / frame; means for determining one or more gradient values and / or one or more position values of the spatial samples in the reference region to be used in a cross-component prediction model; and means for predicting the target samples of at least one color channel of the picture block unit using the cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

[0009] According to an embodiment, the apparatus includes means for using the gradient values and the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

[0010] According to an embodiment, the apparatus includes means for using a combination of the gradient values and the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

[0011] According to an embodiment, the apparatus includes components for signaling the use of the gradient value and / or the position value of the spatial samples in a bitstream including a predicted image block unit or together with the bitstream.

[0012] According to an embodiment, the apparatus includes means for pre-predicting the target samples of at least one color channel of an image block unit using the following cross-component prediction model:

[0013] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad hor + P n+1 ·grad ver + P n+2 ·grad diag +…+ P M-1 ·β

[0014] where pred is the predicted chrominance sample, P is the model coefficient for each tap of the M-tap filter, (C, N, S, W, E) are examples of spatial samples at the (center, north, south, west, east) positions, and grad hor , grad ver , grad diag are the directional gradient values in the horizontal, vertical, and diagonal directions respectively, and β is the offset value.

[0015] According to an embodiment, the apparatus includes components for using the sum or average of the gradient values in different directions as a parameter in the cross-component prediction model.

[0016] According to an embodiment, the apparatus includes components for using one or more of the gradient values in a non-linear function and using the output of the non-linear function as a parameter in the cross-component prediction model.

[0017] According to an embodiment, the apparatus includes components for using higher-order differentials of the gradient values as parameters in the cross-component prediction model.

[0018] According to an embodiment, the apparatus includes components for predicting the target samples of at least one color channel of an image block unit using the following cross-component prediction model:

[0019] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·X + P n+1 ·Y +…+ P M-1 ·β

[0020] where pred is the predicted chrominance sample, P are the model coefficients for each tap of the M-tap filter, (C, N, S, W, E) are examples of spatial samples for the (center, north, south, west, east) positions, and X and Y are the horizontal and vertical coordinates of the center sample, respectively, and β is the offset value.

[0021] According to an embodiment, the apparatus includes means for using the horizontal and / or vertical position of the center sample.

[0022] According to an embodiment, the apparatus includes means for using the horizontal and / or vertical position of (a plurality of) spatially adjacent samples together with the position information of the center sample or instead of the position information of the center sample.

[0023] According to an embodiment, the apparatus includes means for using one or more of the position values in a non-linear function and using the output of the non-linear function as a parameter in the cross-component prediction model.

[0024] According to an embodiment, the apparatus includes means for using a rate-distortion optimization (RDO) algorithm to determine a combination of the gradient value and / or the position value of the filter and the spatial samples, the combination of the gradient value and / or the position value of the filter and the spatial samples being used together with the spatial sample value in the cross-component prediction model or instead of the spatial sample value in the cross-component prediction model.

[0025] As described above, the apparatus and the computer-readable storage medium storing the code are thus arranged to implement one or more of the above-described methods and the related embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] For a better understanding of the present invention, reference will now be made, by way of example, to the accompanying drawings, in which:

[0027] Figure 1 An electronic device adopting an embodiment of the present invention is schematically shown;

[0028] Figure 2 A user equipment suitable for adopting an embodiment of the present invention is schematically shown;

[0029] Figure 3 An electronic device connected using wireless and wired network connections and adopting an embodiment of the present invention is further schematically shown;

[0030] Figure 4a and Figure 4b An encoder and a decoder suitable for implementing an embodiment of the present invention are schematically shown;

[0031] Figure 5Illustrates the positions of samples used for deriving parameters of a Cross-Component Linear Model (CCLM);

[0032] Figure 6a and Figure 6b respectively show examples of classifying luminance samples into two classes in the sampling domain and the spatial domain;

[0033] Figure 7 Shows an example of a co-located reference sample region composed of reconstructed luminance and chrominance samples defined for both luminance and chrominance for a Convolutional Cross-Component Model (CCCM);

[0034] Figure 8 Shows various examples of the dimensions of filter kernels in the CCCM;

[0035] Figure 9 Illustrates an example of four reference lines adjacent to a prediction block;

[0036] Figure 10 Illustrates the prediction process within a matrix-weighted intra-frame;

[0037] Figure 11 Shows an example of a Low-Frequency Non-Separable Transform (LFNST) process;

[0038] Figure 12 Shows a flowchart of a method for predicting samples in at least one color channel according to an embodiment of the present invention;

[0039] Figure 13 Shows a schematic diagram of an example multimedia communication system in which various embodiments can be implemented. Detailed Description of the Invention

[0040] Suitable apparatuses and possible mechanisms for chrominance sampling prediction are described in more detail below. In this regard, reference is first made to Figure 1 and Figure 2 , where Figure 1 shows a block diagram of a video coding system according to an example embodiment, as a schematic block diagram of an exemplary apparatus or electronic device 50 that can incorporate a codec according to an embodiment of the present invention. Figure 2 Shows the layout of an apparatus according to an example embodiment. The elements of Figure 1 and Figure 2 will be explained below.

[0041] The electronic device 50 can be, for example, a mobile terminal or a user equipment of a wireless communication system. However, it should be understood that embodiments of the present invention can be implemented within any electronic device or apparatus that may require encoding and decoding or encoding or decoding video images.

[0042] The device 50 may include a housing 30 for enclosing and protecting the device. The device 50 may also include a display 32 in the form of a liquid crystal display. In other embodiments of the present invention, the display may be any suitable display technology adapted to display images or videos. The device 50 may also include a keyboard 34. In other embodiments of the present invention, any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or a data input system that is part of a touch-sensitive display.

[0043] The device may include a microphone 36 or any suitable audio input that may be a digital or analog signal input. The device 50 may also include an audio output device, which in embodiments of the present invention may be any one of the following: headphones 38, a speaker, or an analog audio or digital audio output connection. The device 50 may also include a battery (or in other embodiments of the present invention, the device may be powered by any suitable mobile energy device, such as a solar cell, a fuel cell, or a clock generator). The device may also include a camera capable of recording or capturing images and / or videos. The device 50 may also include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, the device 50 may also include any suitable short-range communication solution, such as a Bluetooth wireless connection or a USB / FireWire wired connection.

[0044] The device 50 may include a controller 56, a processor, or a processor circuit for controlling the device 50. The controller 56 may be connected to a memory 58, which in embodiments of the present invention may store data in the form of image and audio data and / or may also store instructions for both implemented on the controller 56. The controller 56 may also be connected to a codec circuit 54, which is suitable for implementing the encoding and decoding of audio and / or video data or for assisting in the encoding and decoding implemented by the controller.

[0045] The device 50 may also include a card reader 48 and a smart card 46, such as a UICC and a UICC reader for providing user information and being suitable for providing authentication information for authenticating and authorizing the user on the network.

[0046] The device 50 may include a radio interface circuit 52, which is connected to the controller and is suitable for generating wireless communication signals, such as for communicating with a cellular communication network, a wireless communication system, or a wireless local area network. The device 50 may also include an antenna 44 connected to the radio interface circuit 52 for transmitting the radio frequency signals generated at the radio interface circuit 52 to other device(s) and for receiving radio frequency signals from other device(s).

[0047] The apparatus 50 may include a camera capable of recording or detecting individual frames, which are then passed to a codec 54 or a controller for processing. The apparatus may receive video image data for processing from another apparatus before transmission and / or storage. The apparatus 50 may also receive images for encoding / decoding wirelessly or via a wired connection. The structural elements of the apparatus 50 described above represent examples of components for performing the corresponding functions.

[0048] Regarding Figure 3 , an example of a system in which embodiments of the present invention may be utilized is shown. The system 10 includes a plurality of communication devices that may communicate via one or more networks. The system 10 may include any combination of wired or wireless networks, including but not limited to wireless cellular telephone networks (such as GSM, UMTS, CDMA networks, etc.), wireless local area networks (WLANs) defined by any IEEE802.x standard, Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet.

[0049] The system 10 may include suitable wired and wireless communication devices and / or the apparatus 50 for implementing embodiments of the present invention.

[0050] For example, Figure 3 the system shown shows a representation of a mobile telephone network 11 and the Internet 28. The connection to the Internet 28 may include but not limited to long-distance wireless connections, short-distance wireless connections, and various wired connections, including but not limited to telephone lines, cable lines, power lines, and similar communication paths.

[0051] The example communication devices shown in the system 10 may include but not limited to a combination of an electronic device or apparatus 50, a personal digital assistant (PDA), and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, and a notebook computer 22. When carried by a mobile individual, the apparatus 50 may be fixed or mobile. The apparatus 50 may also be in a transportation mode, including but not limited to an automobile, a truck, a taxi, a bus, a train, a ship, an airplane, a bicycle, a motorcycle, or any similar suitable transportation mode.

[0052] Embodiments may also be implemented in a set-top box; that is, in a flat or (laptop) personal computer (PC) having a combination of hardware or software or encoder / decoder implementations, in various operating systems, and in a chipset, a processor, a DSP, and / or an embedded system providing hardware / software-based encoding, a digital TV receiver that may or may not have a display or wireless capabilities.

[0053] Some or additional devices can send and receive calls and messages and communicate with a service provider via a wireless connection 25 to a base station 24. The base station 24 can be connected to a network server 26, which allows communication between the mobile phone network 11 and the Internet 28. The system can include additional communication devices and various types of communication devices.

[0054] The communication devices can communicate using various transmission technologies, including but not limited to Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol - Internet Protocol (TCP-IP), Short Message Service (SMS), Multimedia Message Service (MMS), email, Instant Messaging Service (IMS), Bluetooth, IEEE 802.11, and any similar wireless communication technology. The communication devices involved in various embodiments of the present invention can communicate using various media, which includes but is not limited to radio, infrared, laser, cable connection, and any suitable connection.

[0055] In telecommunication and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wire, while a logical channel can refer to a logical connection on a multiplexed medium capable of conveying several logical channels. A channel can be used to convey an information signal, such as a bit stream, from one or several transmitters (or emitters) to one or several receivers.

[0056] The MPEG-2 Transport Stream (TS) specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0 is a format for carrying audio, video, and other media, as well as program metadata or other metadata, in a multiplexed stream. Packet Identifiers (PIDs) are used to identify the elementary streams (also known as packetized elementary streams) within the TS. Thus, the logical channels within the MPEG-2 TS can be considered to correspond to specific PID values.

[0057] Available media file format standards include the ISO Base Media File Format (ISO / IEC 14496-12, which can be abbreviated as ISOBMFF) and the file format for NAL unit structured video derived from ISOBMFF (ISO / IEC 14496-15).

[0058] A video codec includes an encoder that transforms an input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a visual form. The video encoder and / or the video decoder can also be separated from each other, i.e., there is no requirement to form a codec. Typically, the encoder discards some information in the original video sequence to represent the video in a more compact form (i.e., at a lower bit rate).

[0059] Typical hybrid video encoders (e.g., multiple encoder implementations of ITU-T H.263 and H.264) encode video information in two stages. First, pixel values in certain picture regions (or "blocks") are predicted, for example, by a motion compensation component (finding and indicating a region in one of the previously encoded video frames that closely corresponds to the block being encoded) or by a spatial component (using pixel values around the block encoded in a specified manner). Second, the prediction error, i.e., the difference between the predicted pixel block and the original pixel block, is determined. This is typically done by transforming the difference in pixel values using a specified transform (e.g., the discrete cosine transform (DCT) or a variant thereof), quantifying the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate).

[0060] In temporal prediction, the prediction source is a previously decoded picture (also known as a reference picture). In intra block copy (IBC; also known as intra block copy prediction), the application of prediction is similar to temporal prediction, but the reference picture is the current picture and only previously decoded samples can be referred to during the prediction process. Inter-layer or inter-view prediction can be applied similarly to temporal prediction, but the reference pictures are decoded pictures from another scalable layer or from another view, respectively. In some cases, inter-frame prediction may refer only to temporal prediction, while in other cases, inter-frame prediction may refer collectively to temporal prediction and any intra block copy, inter-layer prediction, and provided inter-view prediction, as long as they are performed in the same or a similar process as temporal prediction. Inter-frame prediction or temporal prediction is sometimes referred to as motion compensation or motion compensation prediction.

[0061] Motion compensation can be performed with full-sampling or sub-sampling precision. In the case of full-sampling accurate motion compensation, motion can be represented as a motion vector with integer values for horizontal and vertical displacements and the motion compensation process uses those displacements to efficiently copy samples from a reference picture. In the case of sub-sampling accurate motion compensation, the motion vector is represented by fractional or decimal values for the horizontal and vertical components of the motion vector. In the case where the motion vector refers to a non-integer position in the reference picture, a sub-sampling interpolation process is typically invoked to calculate a predicted sample value based on the reference samples and the selected sub-sampling position. The sub-sampling interpolation process typically consists of a horizontal filtering compensating for the horizontal offset relative to the full-sampling position, followed by a vertical filtering compensating for the vertical offset relative to the full-sampling position. However, in some environments, the vertical processing can also be done before the horizontal processing.

[0062] Inter-frame prediction (which may also be referred to as temporal prediction, motion compensation or motion-compensated prediction) reduces temporal redundancy. In inter-frame prediction, the prediction source is a previously decoded picture. Intra-frame prediction exploits the fact that adjacent pixels within the same picture may be correlated. Intra-frame prediction can be performed in the spatial or transform domain, i.e., it can be either the predicted sample values or the transform coefficients. Intra-frame prediction is typically exploited in intra-frame coding where inter-frame prediction is not applied.

[0063] One result of the encoding process is a set of encoded parameters such as motion vectors and quantized transform coefficients. If multiple parameters are predicted first from spatially or temporally adjacent parameters, they can be entropy-encoded more efficiently. For example, a motion vector can be predicted based on spatially adjacent motion vectors and only the difference relative to the predicted value of the motion vector can be encoded. The prediction of encoded parameters and intra-frame prediction can be collectively referred to as intra-picture prediction.

[0064] Figure 4a and Figure 4b Illustrates suitable encoders and decoders for adopting embodiments of the present invention. A video codec consists of an encoder that transforms an input video into a suitable compressed representation for storage / transmission and a decoder that can decompress the compressed video representation back into a visual form. Typically, the encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate). Figure 4a Illustrates an example of the encoding process. Figure 4a Illustrates the encoded image (I n ); the predicted representation of the image block (P'n); the prediction error signal (D n ); the reconstructed prediction error signal (D'n); the preliminary reconstructed image (I' n ); the final reconstructed image (R' n ); the transform (T) and the inverse transform (T -1 ); the quantization (Q) and the inverse quantization (Q-1 );Entropy encoding (E); Reference frame memory (RFM); Inter-frame prediction (P inter );Intra-frame prediction (P intra );Mode selection (MS) and filter (F).

[0065] Figure 4b The figure illustrates an example of the decoding process. Figure 4b The figure illustrates the predicted representation of an image block (P'n); the reconstructed prediction error signal (D'n); the preliminary reconstructed image (I'n); the final reconstructed image (R'n); the inverse transform (T -1 );Inverse quantization (Q -1 );Entropy decoding (E -1 );Reference frame memory (RFM); Prediction (inter-frame or intra-frame) (P); and filter (F).

[0066] Multiple hybrid video encoders encode video information in two stages. First, the pixel values in certain picture regions (or "blocks") are predicted, for example, by a motion compensation component (finding and indicating a region in one of the previously encoded video frames that closely corresponds to the block being encoded) or by a spatial component (using the pixel values around the block encoded in a specified manner). Second, the prediction error, i.e., the difference between the predicted pixel block and the original pixel block. This is typically done by transforming the difference in pixel values using a specified transform (such as the discrete cosine transform (DCT) or its variant), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate). The video codec can also provide a transform skip mode that the encoder can choose to use. In the transform skip mode, the prediction error is encoded in the sample domain, for example, by deriving the sample-by-sample differences relative to certain adjacent samples and using an entropy encoder to encode the sample-by-sample differences.

[0067] Entropy encoding / decoding can be performed in various ways. For example, context-based encoding / decoding can be applied, where both the encoder and the decoder modify the context state of the encoding parameters based on the previously encoded / decoded encoding parameters. Context-based encoding can be, for example, context-adaptive binary arithmetic coding (CABAC) or context-based variable length coding (CAVLC) or any similar entropy coding. Entropy encoding / decoding can alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding / decoding or Exp-Golomb coding / decoding. The decoding of the encoding parameters from the bitstream or codewords of the entropy encoding can be referred to as parsing.

[0068] A phrase along a bitstream (e.g., an indication along a bitstream) can be defined to refer to out-of-band transmission, signaling, or storage in a manner associated with the bitstream by out-of-band data. Decoding of a phrase along a bitstream, etc., can refer to decoding the out-of-band data referred to that is associated with the bitstream (which can be obtained from out-of-band transmission, signaling, or storage). For example, an indication along a bitstream can refer to metadata in a container file encapsulating the bitstream.

[0069] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard was published by both of the parent standardization organizations and is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, integrating new extensions or features into the specification. These extensions include Scalable Video Coding (SVC) and Multi-View Video Coding (MVC).

[0070] Version 1 of High Efficiency Video Coding (H.265 / HEVC, also known as HEVC) was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by both of the parent standardization organizations and is known as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Later versions of H.265 / HEVC include scalable, multi-view, range-extended, 3D, and screen content coding extensions, which can be abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.

[0071] Versatile Video Coding (VVC) (MPEG-1 Part 3), also known as ITU-T H.266, is a video compression standard developed by the Joint Video Exploration Team (JVET) of the Moving Picture Experts Group (MPEG) (officially ISO / IEC JTC1 SC29 WG11) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU), succeeding HEVC / H.265.

[0072] In this chapter, some key definitions, bitstreams, and coding structures, as well as concepts of H.264 / AVC and HEVC, are described as examples of a video encoder, decoder, encoding method, decoding method, and bitstream structure, where embodiments can be implemented. Some key definitions, bitstreams, and coding structures, as well as concepts of H.264 / AVC are the same as those in HEVC, and therefore, they are described together below. Aspects of the present invention are not limited to H.264 / AVC or HEVC, but an embodiment is given as a possible basis that can partially or fully implement the present invention.

[0073] Similar to a number of earlier video coding standards, bitstream syntax and semantics, as well as the decoding process for an error-free bitstream, are specified in H.264 / AVC and HEVC. The encoding process is not specified, but the encoder must generate a consistent bitstream. Bitstream and decoder consistency can be verified by assuming a reference decoder (HRD). The standard includes coding tools that help to cope with transmission errors and losses, but the use of this tool in encoding is optional and the decoding process for an error bitstream is not specified.

[0074] The basic unit for the input of an H.264 / AVC or HEVC encoder and the output of an H.264 / AVC or HEVC decoder is a picture. A picture given as the input of an encoder can also be referred to as a source picture, and a picture decoded by decoding can be referred to as a decoded picture.

[0075] Each of the source picture and the decoded picture includes one or more sampling arrays, such as one sampling array from the following set of sampling arrays:

[0076] - Luminance only (Y) (monochrome).

[0077] - Luminance and two chrominances (YCbCr or YCgCo).

[0078] - Green, blue, and red (GBR, also known as RGB).

[0079] - An array representing other unspecified monochrome or trichromatic stimulus color samplings (e.g., YZX, also known as XYZ).

[0080] In H.264 / AVC and HEVC, a picture can be a frame or a field. A frame includes a matrix of luminance samplings and possibly corresponding chrominance samplings. A field is an alternating set of sampling rows of a frame and can be used as the encoder input when the source signal is interlaced. The chrominance sampling array may be absent (and thus monochrome sampling can be used) or the chrominance sampling array may be subsampled when compared with the luminance sampling array. The chrominance formats can be summarized as follows:

[0081] - In monochrome sampling, there is only one sampling array, which can be nominally considered as the luminance array.

[0082] - In 4:2:0 sampling, each of the two chrominance arrays has half the height and half the width of the luminance array.

[0083] - In 4:2:2 sampling, each of the two chrominance arrays has the same height and half the width of the luminance array.

[0084] - In 4:4:4 sampling, when not using separate color planes, each of the two chrominance arrays has the same height and width as the luminance array.

[0085] In H.264 / AVC and HEVC, it is possible to encode the sample arrays as separate color planes into the bitstream and decode the separately encoded color planes from the bitstream. When using separate color planes, each of them is separately processed (by the encoder and / or decoder) as a picture with monochrome sampling.

[0086] A partition can be defined as dividing a set into subsets such that each element of the set is exactly in one of the subsets.

[0087] When describing the operations of HEVC encoding and / or decoding, the following terms can be used. An encoding block can be defined as an N×N sampling block for some N value, such that dividing the coding tree block into encoding blocks is a partition. A coding tree block (CTB) can be defined as an N×N sampling block for some N value such that dividing the component into coding tree blocks is a partition. A coding tree unit (CTU) can be defined as a coding tree block of luminance samples, two corresponding coding tree blocks of chrominance samples of a picture with three sample arrays, or a coding tree block of samples of a monochrome picture or a picture encoded using three separate color planes and syntax structures for encoding samples. A coding unit (CU) can be defined as a coding block of luminance samples, two corresponding coding blocks of chrominance samples of a picture with three sample arrays, or a coding block of samples of a monochrome picture or a picture encoded using three separate color planes and syntax structures for encoding samples. The CU with the maximum allowed size can be named as LCU (largest coding unit) or coding tree unit (CTU) and the video picture is divided into non-overlapping LCUs.

[0088] A CU consists of one or more prediction units (PUs) that define a prediction process for samples within the CU and one or more transform units (TUs) that define a prediction error encoding process for the samples in the CU. Generally, a CU consists of a square block of samples having a size that can be selected from a predefined set of possible CU sizes. Each PU and TU can be further divided into smaller PUs and TUs to increase the granularity of the prediction and prediction error encoding processes, respectively. Each PU has prediction information associated with it that defines what type of prediction (e.g., motion vector information for an inter-prediction PU and intra-prediction directionality information for an intra-prediction PU) will be applied to the pixels within the PU.

[0089] Each TU can be associated with information (including, for example, DCT coefficient information) that describes a prediction error decoding process for sampling within the TU. In the absence of a prediction error residual associated with the CU, it can be considered that there are no TUs for the CU. Generally, in the bitstream, the partitioning of the image into CUs and the partitioning of the CUs into PUs and TUs are signaled, allowing the decoder to reproduce the expected structure of these units.

[0090] In HEVC, a picture can be partitioned into rectangles and contain an integer number of tiles of LCU. In HEVC, the partitioning of tiles forms a regular grid where the height and width of tiles differ by at most one LCU from each other. In HEVC, a slice is defined as an integer number of coding tree units within the same access unit and all subsequent dependent slice segments (if any) before the next independent slice segment (if any). In HEVC, a slice segment is defined as an integer number of coding tree units that are consecutively ordered in a tile scan and contained within a single NAL unit. The partitioning of each picture into slice segments is a partition. In HEVC, an independent slice segment is defined as a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values for the previous slice segment, and a dependent slice segment is defined as a slice segment for which the values of some of the syntax elements of the slice segment header are inferred in decoding order from the values for the previous independent slice segment. In HEVC, a slice header is defined as the slice segment header of an independent slice segment that is the current slice segment or an independent slice segment before the current dependent slice segment, and a slice segment header is defined as the part of the coded slice segment that contains data elements belonging to the first or all coding tree units represented in the slice segment. If tiles are not used, the CUs are scanned in raster scan order of the LCUs within the tile or picture. Within an LCU, the CUs have a specific scan order.

[0091] The decoder reconstructs the output video by applying a prediction component similar to the encoder to form a predictive representation of pixel blocks (using motion or spatial information created by the encoder and stored in the compressed representation) and predictive error decoding (the inverse operation of predictive error encoding to recover the quantized predictive error signal in the spatial pixel domain). After applying the prediction and predictive error decoding components, the decoder sums the prediction and predictive error signals (pixel values) to form the output video frame. The decoder (and encoder) may also apply additional filter components to improve the quality of the output video before passing the output video for display and / or storing it as a predictive reference for upcoming frames in a video sequence.

[0092] The filter may include, for example, one or more of the following: deblocking, sample adaptive offset (SAO), and / or adaptive loop filter (ALF). H.264 / AVC includes deblocking, while HEVC includes both deblocking and SAO.

[0093] In a typical video codec, motion information, such as prediction units, is indicated by motion vectors associated with each motion-compensated image block. Each of these motion vectors represents the displacement of an image block in the encoded (on the encoder side) or decoded (on the decoder side) picture and a predicted source block in one of the previously encoded or decoded pictures. To facilitate efficient representation of the motion vectors, those motion vectors are typically differentially encoded relative to a block-specific predicted motion vector. In a typical video codec, the predicted motion vector is created in a predetermined manner, such as by computing the median of the encoded or decoded motion vectors of adjacent blocks. Another way to create a motion vector prediction is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in a temporal reference picture and signal the selected candidate as the motion vector prediction value. In addition to the predicted motion vector value, it can also be predicted which (multiple) reference pictures are used for motion-compensated prediction and this prediction information can be represented, for example, by the reference index of the previously encoded / decoded picture. The reference index is typically predicted based on adjacent blocks and / or co-located blocks in the temporal reference picture. Furthermore, typical high-efficiency video codecs employ an additional motion information encoding / decoding mechanism often referred to as merge / merge mode, where all motion field information including the motion vectors and the corresponding reference picture indices for each available reference picture list is predicted and used without any modification / correction. Similarly, the motion field information is predicted using the motion field information of adjacent blocks and / or co-located blocks in the temporal reference picture and the used motion field information is signaled in a list of motion field candidates filled with the motion field information of available adjacent / co-located blocks.

[0094] In a typical video codec, first, the prediction residual after motion compensation is transformed using a transform kernel (e.g., DCT), and then it is encoded. The reason is that there are often still some correlations in the residuals, and the transform can help reduce this correlation and provide more efficient encoding in many cases.

[0095] Video coding standards and specifications may allow the encoder to divide the coded picture into coding slices, etc. Intra-picture prediction is typically disabled across slice boundaries. Thus, slices can be considered a way to divide the coded picture into independently decodable tiles. In H.264 / AVC and HEVC, intra-picture prediction can be disabled across slice boundaries. Thus, slices can be considered a way to divide the coded picture into independently decodable tiles, and thus slices are often considered the basic unit for transmission. In many cases, the encoder can indicate in the bitstream which types of intra-picture prediction are turned off across slice boundaries, and the decoder operation takes this information into account, for example, when inferring which prediction sources are available. For example, if adjacent CUs reside in different slices, then samples from adjacent CUs can be considered unavailable for intra-frame prediction.

[0096] The basic unit for the output of an H.264 / AVC or HEVC encoder and the input of an H.264 / AVC or HEVC decoder is the Network Abstraction Layer (NAL) unit. For transport over a packet-oriented network or storage in a structured file, the NAL unit can be encapsulated into a packet or a similar structure. Byte-stream formats have been specified in H.264 / AVC and HEVC for transport or storage environments that do not provide a framing structure. The byte-stream format separates NAL units from each other by appending a start code before each NAL unit. To avoid error detection at NAL unit boundaries, the encoder runs a byte-oriented start code emulation prevention algorithm that adds emulation prevention bytes to the NAL unit payload if the start code would otherwise occur. To facilitate direct gateway operation between packet-oriented and stream-oriented systems, start code emulation prevention can always be performed regardless of whether the byte-stream format is used. The NAL unit can be defined as a syntax structure that contains an indication of the data type that follows and bytes containing data in the form of a Raw Byte Sequence Payload (RBSP) that may be scattered with emulation prevention bytes as needed. The Raw Byte Sequence Payload (RBSP) can be defined as a syntax structure that contains an integral number of bytes encapsulated in the NAL unit. The RBSP is either empty or has the form of a data bitstring containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.

[0097] The NAL unit consists of a header and a payload. In H.264 / AVC and HEVC, the NAL unit header indicates the type of the NAL unit.

[0098] In HEVC, a two-byte NAL unit header is used for all specified NAL unit types. The NAL unit header contains a reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plus1 indication for the temporal level (which may be required to be greater than or equal to 1), and a six-bit nuh_layer_id syntax element. The temporal_id_plus1 syntax element can be regarded as the temporal identifier for the NAL unit, and the zero-based TemporalId variable can be derived as follows: TemporalId = temporal_id_plus1 - 1. The abbreviation TID can be used interchangeably with the TemporalId variable. A TemporalId equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 is required to be non-zero in order to avoid start code emulation involving two NAL unit header bytes. The bitstream created by excluding all VCL NAL units with a TemporalId greater than or equal to the selected value and including all other VCL NAL units still conforms. Therefore, a picture with a TemporalId equal to tid_value does not use any picture with a TemporalId greater than tid_value as an inter-prediction reference. A sublayer or temporal sublayer can be defined as a temporal scalability layer (or temporal layer TL) of a temporally scalable bitstream, which consists of VCL NAL units with a specific value of the TemporalId variable and associated non-VCL NAL units. The nuh_layer_id can be understood as the scalability layer identifier.

[0099] NAL units can be classified into video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units. In HEVC, VCL NAL units contain syntax elements representing one or more CUs.

[0100] Non-VCL NAL units can be, for example, one of the following types: sequence parameter set, picture parameter set, supplementary enhancement information (SEI) NAL unit, access unit delimiter, sequence end NAL unit, bitstream end NAL unit, or filler data NAL unit. Parameter sets may be required for reconstructing decoded pictures, while multiple other non-VCL NAL units are not necessary for reconstructing the sample values of decoded pictures.

[0101] Parameters that remain constant in an encoded video sequence can be included in the sequence parameter set. In addition to the parameters that may be required for the decoding process, the sequence parameter set may optionally contain video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. In HEVC, the sequence parameter set RBSP includes parameters that can be referenced by one or more picture parameter set RBSPs or one or more SEI NAL units containing buffering period SEI messages. The picture parameter set contains these parameters that may not change in several encoded pictures. The picture parameter set RBSP may include parameters that can be referenced by the coded slice NAL units of one or more encoded pictures.

[0102] In HEVC, the video parameter set (VPS) can be defined as a syntax structure containing syntax elements that apply to the content of the syntax elements found in the SPS, which are referenced by the syntax elements found in the PPS, which are referenced by the syntax elements found in each slice segment header, for zero or more entire encoded video sequences.

[0103] The video parameter set RBSP may include parameters that can be referenced by one or more sequence parameter set RBSPs.

[0104] The relationship and hierarchy among the video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) can be described as follows. The VPS resides in the parameter set hierarchy and at a level above the SPS in the context of scalability and / or 3D video. The VPS may include parameters that are common across all slices for all (scalability or view) layers in the entire encoded video sequence. The SPS includes parameters that are common to all slices in a specific (scalability or view) layer in the entire encoded video sequence and may be shared by multiple (scalability or view) layers. The PPS includes parameters that are common to all slices in a specific layer representation (the representation of a scalability or view layer in an access unit) and may be shared by all slices in multiple layer representations.

[0105] The VPS can provide information about the layer dependencies in the bitstream, as well as several other information that can be applied across all slices for all (scalability or view) layers in the entire encoded video sequence. The VPS can be considered to include two parts, the base VPS and the VPS extension, where the VPS extension may be optionally present.

[0106] Out-of-band transmission, signaling, or memory can be additionally or alternatively used for other purposes besides tolerating transmission errors, such as ease of access or session negotiation. For example, the sample entry of a track in a file conforming to the ISO base media file format can include a parameter set, while the encoded data in the bitstream is stored elsewhere in the file or in another file. Phrases along the bitstream (e.g., along the bitstream indication) or along the coded units of the bitstream (e.g., along the coded tile indication) can be used in the claims and the described embodiments to refer to out-of-band transmission, signaling, or memory in such a way that the out-of-band data is associated with the bitstream or the coded unit, respectively. Decoding of phrases such as along the bitstream or along the coded units of the bitstream, etc. can refer to decoding the referenced out-of-band data (which can be obtained from out-of-band transmission, signaling, or memory) that is associated with the bitstream or the coded unit, respectively.

[0107] An SEI NAL unit can contain one or more SEI messages that are not required for decoding of the output picture but can assist related processes such as picture output timing, rendering, error detection, error concealment, and resource reservation.

[0108] An encoded picture is an encoded representation of a picture.

[0109] In HEVC, an encoded picture can be defined as an encoded representation of a picture that contains all the coded tree units of the picture. In HEVC, an access unit (AU) can be defined as a set of NAL units that are related to each other according to specified classification rules, are consecutive in decoding order, and contain at most one picture with any particular value of nuh_layer_id. Besides the VCL NAL units that contain the encoded picture, the access unit can also contain non-VCL NAL units. The specified classification rules can, for example, associate pictures with the same output time or picture output count value into the same access unit.

[0110] A bitstream can be defined as a sequence of bits in the form of a NAL unit stream or a byte stream that forms the representation of an encoded picture and the related data of one or more encoded video sequences. A first bitstream can be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. A elementary stream (in the context of video coding) can be defined as a sequence of one or more bitstreams. The end of the first bitstream can be indicated by a specific NAL unit, which can be called the end-of-bitstream (EOB) NAL unit and which is the last NAL unit of the bitstream. In HEVC and its current draft extensions, the EOB NAL unit needs to have a nuh_layer_id equal to 0.

[0111] In H.264 / AVC, an encoded video sequence is defined as a sequence of consecutive access units in decoding order from an IDR access unit (including that point) to the next IDR access unit (excluding that point) or the end of the bitstream, whichever comes earlier.

[0112] In HEVC, for example, an encoded video sequence (CVS) can be defined as a sequence of access units that, in decoding order, consists of an IRAP access unit with a NoRaslOutputFlag equal to 1, followed by zero or more access units that are not IRAP access units with a NoRaslOutputFlag equal to 1, and that includes all subsequent access units until but not including any subsequent access unit that is an IRAP access unit with a NoRaslOutputFlag equal to 1. An IRAP access unit can be defined as an access unit in which the base layer picture is an IRAP picture. For each IDR picture, each BLA picture, and each IRAP picture, the value of NoRaslOutputFlag is equal to 1, where each IRAP picture is the first picture in that particular layer in the bitstream in decoding order and is the first IRAP picture after the end of the sequence of NAL units with the same nuh_layer_id value in decoding order. There can be a component that provides a HandleCraAsBlaFlag value from an external entity such as a player or receiver that can control the decoder. For example, a player that locates a new position in the bitstream or tunes to a broadcast and starts decoding, and then starts decoding from a CRA picture, can set the HandleCraAsBlaFlag to 1. When the HandleCraAsBlaFlag for a CRA picture is equal to 1, the CRA picture is processed and decoded as if it were a BLA picture.

[0113] In HEVC, when a particular NAL unit (which can be referred to as an end-of-sequence (EOS) NAL unit) appears in the bitstream and has a nuh_layer_id equal to 0, the end of the encoded video sequence can be additionally or alternatively (in accordance with the above specification) specified.

[0114] A Group of Pictures (GOP) and its characteristics can be defined as follows. A GOP can be decoded regardless of whether any previous pictures have been decoded. An open GOP is a group of pictures where, when decoding starts from an intra picture of the initial frame of the open GOP, pictures before the initial intra picture in the output order may not be decoded correctly. In other words, pictures of an open GOP can refer to pictures (in inter prediction) belonging to the previous GOP. An HEVC decoder can identify the intra picture that starts an open GOP because a specific NAL unit type (CRA NAL unit type) can be used for its coding segment. A closed GOP is a group of pictures where, when decoding starts from an intra picture of the initial frame of the closed GOP, all pictures can be decoded correctly. In other words, no pictures in a closed GOP refer to any pictures in the previous GOP. In H.264 / AVC and HEVC, a closed GOP can start from an IDR picture. In HEVC, a closed GOP can also start from a BLA_W_RADL or BLA_N_LP picture. Compared with the closed GOP coding structure, the open GOP coding structure may be more efficient in terms of compression because of greater flexibility in selecting reference pictures.

[0115] A Decoded Picture Buffer (DPB) can be used in the encoder and / or decoder. There are two reasons for buffering decoded pictures, for reference in inter prediction and for reordering the decoded pictures into the output order. Since H.264 / AVC and HEVC provide great flexibility for both reference picture marking and output reordering, separate buffering for reference picture buffering and output picture buffering may waste memory resources. Therefore, the DPB can include a unified decoded picture buffering process for reference picture and output reordering. When a decoded picture is no longer used as a reference and is not needed for output, the decoded picture can be removed from the DPB.

[0116] In multiple coding modes of H.264 / AVC and HEVC, reference pictures for inter prediction are indicated by indices of reference picture lists. The indices can be encoded by variable length coding, which generally makes smaller indices have shorter values for the corresponding syntax elements. In H.264 / AVC and HEVC, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi - directional prediction (B) slice, and one reference picture list (reference picture list 0) is formed for each inter - coded (P) slice.

[0117] Multiple coding standards, including H.264 / AVC and HEVC, can have a decoding process to derive a reference picture index for a reference picture list, which can be used to indicate which reference picture among multiple reference pictures is used for inter-frame prediction of a specific block. In some inter-frame coding modes, the reference picture index can be encoded into the bitstream by the encoder, or can be derived (by the encoder and decoder) using adjacent blocks, for example, in some other inter-frame coding modes.

[0118] The motion parameter type or motion information can include, but is not limited to, one or more of the following types:

[0119] - Indication of the prediction type (e.g., intra prediction, single prediction, dual prediction) and / or the number of reference pictures;

[0120] - Indication of the prediction direction, such as inter-frame (also known as temporal) prediction, inter-layer prediction, view-interpolation prediction, view synthesis prediction (VSP), and inter-component prediction (which can be indicated by reference pictures and / or prediction types, and where in some embodiments, view-interpolation and view synthesis prediction can be regarded as one prediction direction) and / or

[0121] - Indication of the reference picture type, such as short-term reference picture and / or long-term reference picture and / or inter-layer reference picture (which can be indicated, for example, for each reference picture)

[0122] - The reference index of the reference picture list and / or any other identifier of the reference picture (which can be indicated, for example, for each reference picture and whose type can depend on the prediction direction and / or reference picture type and which can be accompanied by other relevant information, such as the reference picture list or similar information for applying the reference index);

[0123] - Horizontal motion vector component (which can be indicated, for example, for each prediction block or each reference index, etc.);

[0124] - Vertical motion vector component (which can be indicated, for example, for each prediction block or each reference index, etc.);

[0125] - One or more parameters, such as the picture order count difference and / or the relative camera separation between the picture containing the motion parameter or associated with the motion parameter and its reference picture, which can be used to scale the horizontal motion vector component and / or the vertical motion vector component in one or more motion vector prediction processes (where the one or more parameters can be indicated, for example, for each reference picture or each reference index, etc.);

[0126] - The coordinates of the block to which the motion parameter and / or motion information is applied, for example, the coordinates of the upper left sample of the block in units of luminance sampling;

[0127] - The range of blocks (e.g., width and height) to which the motion parameters and / or motion information are applied.

[0128] Compared with previous video coding standards, the Versatile Video Codec (H.266 / VVC) introduces multiple new coding tools, such as the following:

[0129] · Intra prediction

[0130] - Intra modes with wide-angle mode extension

[0131] - 4-tap interpolation filter related to block size and mode

[0132] - Position-dependent intra prediction combination (PDPC)

[0133] - Cross-component linear model intra prediction (CCLM)

[0134] - Multi-reference line intra prediction

[0135] - Intra sub-partitioning

[0136] - Weighted intra prediction with matrix multiplication

[0137] · Inter-picture prediction

[0138] - Block motion copy with spatial, temporal, history-based, and pairwise average merge candidates

[0139] - Affine motion inter prediction

[0140] - Sub-block-based temporal motion vector prediction

[0141] - Adaptive motion vector resolution

[0142] - 8×8 block-based motion compression for temporal motion prediction

[0143] - High-precision (1 / 16 pixel) motion vector memory and motion compensation using an 8-tap interpolation filter for the luminance component and a 4-tap interpolation filter for the chrominance component

[0144] - Triangle partitioning

[0145] - Joint intra and inter prediction

[0146] - Merge with MVD (MMVD)

[0147] - Symmetric MVD coding

[0148] - Bidirectional optical flow

[0149] - Decoder-side motion vector refinement

[0150] - Bi - directional prediction with CU - level weighting

[0151] · Transformation, quantization, and coefficient coding

[0152] - Multiple primary transform selection with DCT2, DST7, and DCT8

[0153] - Secondary transform for low - frequency bands

[0154] - Sub - block transform for inter - prediction residuals

[0155] - Corresponding quantization of the maximum QP is increased from 51 to 63

[0156] - Transform coefficient coding with sign data hiding

[0157] - Transform skip residual coding

[0158] · Entropy coding

[0159] - Arithmetic coding engine with adaptive dual - window probability update

[0160] · In - loop filter

[0161] - In - loop shaping

[0162] - Deblocking filter with strong and longer filters

[0163] - Sample adaptive offset

[0164] - Adaptive loop filter

[0165] · Screen content coding:

[0166] - Current picture reference with reference region restriction

[0167] · 360 - degree video coding

[0168] - Horizontal wrap - around motion compensation

[0169] · Advanced syntax and parallel processing

[0170] - Reference picture management with direct reference picture list signaling

[0171] - Tile groups with rectangular - shaped tile groups

[0172] Implementing partitions in VVC similar to HEVC, i.e., dividing each picture into Coding Tree Units (CTUs). A picture can also be divided into slices, tiles, rectangular blocks, and sub-pictures. A CTU can be split into smaller CUs using a quadtree structure. Each CU can be divided using a quadtree and a nested multi-type tree including ternary splitting and binary splitting. However, there are specific rules for inferring partitions at picture boundaries, and redundant splitting patterns are not allowed in nested multi-type partitioning.

[0173] Among the new coding tools listed above, the Cross-Component Linear Model (CCLM) prediction mode is used in VVC to reduce cross-component redundancy. Among them, chrominance samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model:

[0174] pred C (i,j) = α·rec L′ (i,j) + β (Equation 1a)

[0175] where pred C (i,j) represents the predicted chrominance sample in the CU and rec L '(i,j) represents the downsampled reconstructed luma sample of the same CU.

[0176] Alternatively, the following equation can be used for CCLM:

[0177]

[0178] where the >> operation represents shifting the value k bits to the right.

[0179] CCLM parameters (α and β) are derived from at most four adjacent chrominance samples and their corresponding downsampled luma samples. Assuming the current chrominance block dimension is W×H, then W’ and H’ are set to

[0180] - When applying the LM mode, W’ = W, H’ = H;

[0181] - When applying the LM-A mode, W’ = W + H;

[0182] - When applying the LM-L mode, H’ = H + W;

[0183] In this paper, the LM-A mode refers to the linear model above, where only the template above (i.e., the sample values from adjacent positions above the CU) is used to calculate the linear model coefficients. To obtain more samples, the template above is extended to (W + H). The LM-L mode refers to the linear model_left, where only the left template (i.e., the sample values from adjacent positions from left to the CU) is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H + W). For non-square blocks, the template above is extended to W + W, and the left template is extended to H + H.

[0184] The adjacent positions above are represented as S[0, -1]... S[W’ - 1, -1], and the left adjacent positions are represented as S[-1, 0]... S[-1, H’ - 1]. Then four samples are selected as

[0185] - When the LM mode is adopted and both the above and left adjacent samples are available, S[W' / 4, -1], S[3*W′ / 4, -1], S[-1, H′ / 4], S[-1, 3*H′ / 4];

[0186] - When the LM-A mode is applied or only the adjacent samples above are available, S[W' / 8, -1], S[3*W′ / 8, -1], S[5*W′ / 8, -1], S[7*W′ / 8, -1];

[0187] - When the LM-L mode is applied or only the left adjacent samples are available, S[-1, H′ / 8], S[-1, 3*H' / 8], S[-1, 5*H' / 8], S[-1, 7*H' / 8];

[0188] The four adjacent luma samples at the selected positions are downsampled and four comparisons are made to find two smaller values: x0A and x1A, and two larger values: x0B and x1B. Their corresponding chroma sampled values are represented as y0A, y1A, y0B, and y1B. Then xA, xB, yA, and yB are derived as:

[0189] X a =(x 0 A +x 1 A +1)>>1; X b =(x 0 B +x 1 B +1)>>1; Y a =(y 0 A +y 1 A +1)>>1; Yb = (y 0 B + y 1 B + 1) >> 1 (Equation 2)

[0190] Finally, the linear model parameters are obtained according to the following equation:

[0191]

[0192] β = Y b - α · X b (Equation 4)

[0193] Figure 5 Shows examples of the positions of the left and upper samples involved in the CCLM mode and the samples of the current block.

[0194] The division operation for calculating the parameter α is implemented through a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented by exponential notation. For example, diff is approximated by 4 significant bits and an exponent. Thus, the table for 1 / diff is reduced to 16 elements for 16 values of the significant part as follows:

[0195] DivTable[] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0} (Equation 5)

[0196] This provides the benefit of reducing both the computational complexity and the memory size required to store the required table.

[0197] To match the chrominance sampling positions for 4:2:0 video sequences, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type - O" and "Type - 2" content respectively:

[0198]

[0199] It should be noted that when the upper reference line is at the CTU boundary, only one luma line (the common line buffer in intra prediction) is used to produce the downsampled luma samples.

[0200] This parameter calculation is performed as part of the decoding process and not just as an encoder search operation. As a result, no syntax is used to convey the α and β values to the decoder.

[0201] For chrominance intra mode coding, a total of 8 intra modes are allowed for chrominance intra mode coding. These modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). The chrominance mode signaling and derivation process are shown in Table 1. The chrominance mode coding directly depends on the intra prediction mode of the corresponding luma block. Due to the block partitioning structure that enables separation of the luma and chrominance components in the I slice, a chrominance block can correspond to multiple luma blocks. Therefore, for the chrominance DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chrominance block is directly inherited.

[0202]

[0203]

[0204] Table 1

[0205] As shown in Table 2, regardless of the value of sps_cclm_enabled_flag, a single binarization table is used.

[0206] Value of intra_chroma_pred_mode Binary string 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111

[0207] Table 2

[0208] In Table 2, the first bin indicates whether it is a regular (0) mode or an LM mode (1). If it is an LM mode, the next bin indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next bin indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first bin of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first bin is inferred to be 0 and thus not encoded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two bins in Table 2 are context-coded through their own context models, and the remaining bins are bypass-coded.

[0209] In addition, in order to reduce the luma-chroma latency in the dual tree, when a 64×64 luma coding tree node is partitioned into non-split (and intra sub-partition (ISP) is not used for 64×64 CU) or QT, the chroma CUs in the 32×32 / 32×16 chroma coding tree nodes are allowed to use CCLM in the following way:

[0210] - If the 32×32 chroma node is not split or partitioned by QT split, all chroma CUs in the 32×32 node can use CCLM.

[0211] - If the 32×32 chroma nodes are partitioned horizontally by BT and the 32×16 child nodes are not partitioned or use vertical BT splitting, all chroma CUs in the 32×16 chroma nodes can use CCLM.

[0212] In all other luma and chroma coding tree splitting conditions, CCLM is not allowed for chroma CUs.

[0213] Multi-model LM (MMLM)

[0214] CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are classified into two classes using a threshold that is the average of the luma reconstructed neighboring samples. A linear model for each class is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Figure 6a and Figure 6b Illustrates two luma-chroma models obtained for luma (Y) thresholds of 17 in the sampling domain and the spatial domain, respectively. Each luma-chroma model has its own linear model parameters α and β. As Figure 6b seen, each luma-chroma model corresponds to a spatial segmentation of the content (i.e., they correspond to different objects or textures in the scene).

[0215] Convolutional cross-component model (CCCM)

[0216] An improved version of cross-component prediction known as CCCM uses a 2D filter kernel to derive the luma-chroma model. Filter coefficients are derived on the decoder side using the reconstructed set of the input data and chroma samples. For filter coefficient derivation, a co-located reference sampling region (consisting of the reconstructed luma and chroma samples) is defined for both luma and chroma, as Figure 7 shown, where the commonly used 4:2:0 chroma subsampling is applied. For example, the reference sampling region for a given block can be the six lines above and to the left as Figure 7 shown, but any number of reference lines can be used (which can be implemented by both the encoder and the decoder). Generally, the reference samples can include any chroma and luma samples that have been reconstructed by both the encoder and the decoder. Once the reference samples are determined, filter coefficients can be derived using, for example, different types of linear regression tools such as ordinary least squares estimation, orthogonal matching pursuit, optimized orthogonal matching pursuit, ridge regression, or least absolute shrinkage and selection operator.

[0217] The dimensions of the filter core can be, for example, 1×3 (1D vertical), 3×1 (1D horizontal), 3×3, 7×7, or any dimension, and can be shaped (by selecting only a subset of all possible core positions) into a cross or diamond (as Figure 8 shown) or any given shape. When referring to samples within the filter kernel, the following notations are used: north (up), east (right), south (down), west (left), and center, as Figure 8 illustrated, using the letters N, E, S, W, C.

[0218] Here, the overall method of reconstructing chrominance sampling using the convolution between the filter kernel obtained on the decoder side and the input data set is referred to as the Convolutional Cross-Component Model (CCCM). The following steps can be applied to perform CCCM operations:

[0219] 1) Define co-located reference regions on the luminance component and the chrominance component.

[0220] 2) Downsample the luminance samples to match the chrominance grid (optionally).

[0221] 3) Scan the luminance and chrominance samples of the reference regions and collect available statistics (such as autocorrelation matrices and cross-correlation vectors) based on the filter shape.

[0222] 4) Solve for the filter coefficients by minimizing the squared error (or any other metric) based on the available statistics (such as autocorrelation matrices and cross-correlation vectors).

[0223] 5) Calculate the predicted chrominance block by convolving the downsampled luminance samples with the filter kernel.

[0224] Let us define the (possibly downsampled) luminance samples as a 2D array Y(x, y) indexed by horizontal x coordinates and vertical y coordinates. Let us also define the co-located chrominance samples as a 2D array C(x, y) and the filter kernel (i.e., the coefficients) as a 3×3 array F(i, j). At the sample level, we define the convolution between Y and F as,

[0225]

[0226] When using other data terms such as non-linear square root terms, the additional convolution becomes,

[0227]

[0228] where the filter coefficients residing outside the 2D filter kernel have been obtained as part of the linear equations used to solve for the 2D filter coefficients in step 4 above. Similarly, we can add a bias term to the convolution as,

[0229]

[0230] Multi-reference line (MRL) intra prediction

[0231] Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. In Figure 9 an example with 4 reference lines is depicted, where the samples of segment A and segment F are not retrieved from the reconstructed neighboring samples, but are filled with the closest samples from segment B and segment E respectively. HEVC picture intra prediction uses the closest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used.

[0232] The index of the selected reference line (mrl_idx) is signaled and used to generate the intra predictor. For reference line idx greater than 0, only the additional reference line modes are included in the MPM list and only the mpm index is signaled without the residual modes. The reference line index is signaled before the intra prediction mode, and the planar mode is excluded from the intra prediction mode in case a non-zero reference line index is signaled.

[0233] MRL for the first row of blocks within a CTU is disabled to prevent the use of extended reference samples outside the current CTU line. Additionally, PDPC is disabled when additional lines are used. For the MRL mode, the derivation of the DC value in the DC intra prediction mode for a non-zero reference line index is aligned with the derivation for reference line index 0. MRL requires storing 3 adjacent luma reference lines of the CTU to generate the prediction. The CCLM tool also requires 3 adjacent luma reference lines for its downsampling filter. The definition of using the same 3 lines of MLR is aligned with CCLM to reduce the storage requirements for the decoder.

[0234] Intra sub - partitioning (ISP)

[0235] The Intra Sub - Partitioning (ISP) divides the luma intra - prediction block vertically or horizontally into 2 or 4 sub - partitions according to the block size. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided by 4 sub - partitions. It has been noted that M×128 (M≤64) and 128×N (N≤64) ISP blocks may generate potential problems for the 64×64 VDPU. For example, in the single - tree case, an M×128 CU has an M×128 luma TB and two corresponding M / 2×64 chroma TBs. If the CU uses ISP, the luma TB will be divided into four M×32 TBs (only horizontal splitting is possible), and each TB is smaller than a 64×64 block. However, in the current design of ISP, the chroma blocks are not divided. Therefore, both chroma components will have a size larger than 32×32 blocks. Approximately, a similar situation can be created with 128×N CUs using ISP. Thus, these two cases are problems for the 64×64 decoder pipeline. For this reason, the CU size used by ISP can be limited to a maximum of 64×64. All sub - partitions satisfy the condition of having at least 16 samples.

[0236] Matrix weighted intra prediction (MIP)

[0237] The Matrix - weighted Intra - Prediction (MIP) method is an intra - prediction technique newly added to VVC. To predict the samples of a rectangular block with width W and height H, the Matrix - weighted Intra - Prediction (MIP) takes as input a row of H reconstructed adjacent boundary samples on the left side of the block and a row of W reconstructed adjacent boundary samples above the block. If the reconstructed samples are not available, they are generated as in the conventional intra - prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix - vector multiplication, and linear interpolation, as Figure 10 shown.

[0238] Decoder-side intra mode derivation (DIMD)

[0239] When applying DIMD, two intra - modes are derived from the reconstructed adjacent samples, and as described in JVET - 00449, these two predictors are combined with a planar - mode predictor with weights derived from the gradient. The division operation in the weight derivation is performed by the integerization scheme based on the same Look - Up Table (LUT) used by CCLM. For example, the division operation in the orientation calculation

[0240] Orient = G y / G x

[0241] is calculated by the following LUT - based scheme:

[0242] x = Floor(Log2(Gx))

[0243] normDiff = ((Gx << 4) >> x) & 15

[0244] x += (3 + (normDiff != 0 ? 1 : 0))

[0245] Orient = (Gy * (DivSigTable[normDiff] | 8) + (1 << (x - 1))) >> x

[0246] where

[0247] DivSigTable

[16] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}

[0248] The derived intra mode is included in the main list of the most probable intra modes (MPM), so the DIMD process is performed before constructing the MPM list. The main derived intra mode of the DIMD block is stored with the block and is used for the construction of the MPM list of adjacent blocks.

[0249] Fusion for template-based intra mode derivation (TIMD)

[0250] For each intra prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The top two intra prediction modes with the minimum SATD are selected as the TIMD modes. These two TIMD modes are weighted and fused after applying the PDPC process, and this weighted intra prediction is used for encoding the current CU. The position-dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.

[0251] The costs of the two selected modes are compared with a threshold. In the test, the cost factor 2 is applied as follows: costMode2 < 2 * costMode1.

[0252] If this condition is true, fusion is applied; otherwise, the unique mode 1 is used.

[0253] The weighting of the modes is calculated based on their SATD costs as follows:

[0254] weight1 = costMode2 / (costMode1 + costMode2)

[0255] weight2 = 1 - weight1

[0256] The division operation is implemented by the same lookup table (LUT)-based integerization scheme used by the CCLM.

[0257] Low-frequency non-separable transform (LFNST)

[0258] As Figure 11 shown in Figure 11 , in VVC, LFNST is applied between the forward main transform and quantization (at the encoder) and between dequantization and the inverse main transform (at the decoder side). In LFNST, according to the block size, a 4×4 non-separable transform or an 8×8 non-separable transform is applied. For example, 4×4 LFNST is applied to smaller blocks (i.e., min(width, height) < 8), while 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4).

[0259] The application of the non-separable transform used in LFNST is described as follows using the input as an example. The 4x4 LFNST will be applied to a 4x4 input unit X

[0260]

[0261] First, it is represented as a vector

[0262]

[0263] The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16×16 transform matrix. Subsequently, the 16×1 coefficient vector is reorganized into a 4×4 block using the scan order (horizontal, vertical, or diagonal) for that block. Coefficients with smaller indices will be placed in the 4×4 coefficient block together with smaller scan indices.

[0264] Reduce non-separable transform

[0265] LFNST (Low Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method, enabling it to be implemented in one pass without multiple iterations. However, it is necessary to reduce the dimension of the non-separable transform matrix to minimize the computational complexity and memory space for storing the transform coefficients. Therefore, the Reduced Non-Separable Transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (for 8×8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix as follows:

[0266]

[0267] where the R rows of the transform are the R bases of the N-dimensional space.

[0268] The inverse transform matrix for RT is the transpose of its forward transform. For 8×8 LFNST, a reduction factor of 4 is applied, and the 64×64 direct matrix of the conventional 8×8 non-separable transform matrix size is reduced to a 16×48 direct matrix. Thus, a 48×16 inverse RST matrix is used at the decoder side to generate the core (primary) transform coefficients in the 8×8 upper left region. When applying the 16×48 matrix instead of the 16×64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4×4 blocks in the 8×8 upper left block except for the lower right 4×4 block. Contributing to the dimensionality reduction, the memory usage for storing all LFNST matrices is reduced from 10 KB to 8 KB with a reasonable performance degradation. To facilitate complexity reduction, LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are unimportant. Thus, when LFNST is applied, all only the primary transform coefficients must be zero. This allows for the adjustment of the LFNST index signaling at the last valid position and thus avoids the additional coefficient scan in the current LFNST design, which is only needed at specific positions to check for valid coefficients. The worst-case processing of LFNST (in terms of multiplications per pixel) restricts the non-separable transforms of 4×4 and 8×8 blocks to 8×16 and 8×48 transforms, respectively. In these cases, when LFNST is applied, for other sizes less than 16, the last valid scan position must be less than 8. For blocks with shapes of 4×N and N×4 where N > 8, the proposed restriction means that LFNST is now applied only once and only to the upper left 4×4 region. Since all only the primary coefficients are zero when LFNST is applied, the number of operations required for the primary transform is reduced in this case. From the encoder's perspective, when testing the LFNST transform, the quantization of coefficients is significantly simplified. For the first 16 coefficients (in scan order), rate-distortion optimized quantization must be performed with the maximum value, and the remaining coefficients are forced to zero.

[0269] LFNST transform selection

[0270] In LFNST, there are a total of 4 transform sets, and each transform set uses 2 non-separable transform matrices (kernels). As shown in the table below, the mapping from the intra prediction mode to the transform set is pre-defined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 ≤ predModeIntra ≤ 83), then transform set 0 is selected for the current chrominance block. For each transform set, the selected non-separable sub-transform candidate is further specified by the explicitly signaled LFNST index. After the transform coefficients, each intra CU is signaled the index once in the bitstream.

[0271] IntraPredMode Tr. set index intrapredmode<0 1 0 <= intrapredmode <= 1 0 2 <= intrapredmode <= 12 1 13 <= intrapredmode <= 23 2 24 <= intrapredmode <= 44 3 45 <= intrapredmode <= 55 2 56 <= intrapredmode <= 80 1 81 <= intrapredmode <= 83 0

[0272] Transform selection table

[0273] LFNST index signaling and interaction with other tools

[0274] Since LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are unimportant, the LFNST index coding depends on the position of the last important coefficient. In addition, the LFNST index is context-coded but does not depend on the intra prediction mode, and only the first bin is context-coded. Further, LFNST is applied to the intra CUs in both intra and inter slices, and is applied to both luminance and chrominance. If dual-tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter-slice (dual-tree disabled), a single LFNST index is signaled and used for both luminance and chrominance.

[0275] Considering that large CUs larger than 64×64 are implicitly partitioned (TU tiling) due to the existing maximum transform size limit (64×64), the LFNST index search can quadruple the data buffering for some number of decoding pipeline stages. Therefore, the maximum size allowed by LFNST is limited to 64x64. Note that LFNST is enabled only by DCT2. The LFNST index signaling is placed before the MTS index signaling.

[0276] The use of the scaling matrix for perceptual quantization is not obvious, and the scaling matrix specified for the primary matrix can be useful for LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For the single-tree partitioning mode, chrominance LFNST is not applied.

[0277] The Convolutional Cross-Component Model (CCCM) uses spatial luminance samples for cross-component prediction to model the luminance-chrominance relationship. Although this method is able to model the cross-channel relationship of samples, it is still sub-optimal when the spatial samples in the luminance and chrominance channels are not well correlated.

[0278] Now an improved method for implementing cross-component prediction is introduced, which provides enhanced modeling of the cross-channel relationship of samples.

[0279] Figure 12A method according to one aspect is shown, wherein the method includes receiving (1200) a block unit of an image frame, the block unit of the image including samples in color channels, wherein the color channels include at least one chrominance channel and one luminance channel; reconstructing (1202) the samples of the luminance channel of the block unit of the image; determining (1204) a reference region for predicting a target sample of at least one color channel of the block unit of the image, wherein the reference region includes one or more of the following reference samples: reference samples in adjacent blocks in the current color channel / frame, reference samples in adjacent blocks of co-located blocks in a reference color channel / frame, and / or reference samples within co-located blocks in the reference color channel / frame; determining (1206) one or more gradient values and / or one or more position values of spatial samples in the reference region to be used in a cross-component prediction model; and predicting (1208) the target sample of at least one color channel of the block unit of the image using the cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

[0280] Thus, in the method, the gradient values and / or position values of the spatial sampling in the reference region are determined to be used in a cross-component prediction model, such as a cross-component linear model (CCLM) and a convolutional cross-component model (CCCM), so as to provide enhanced correlation between the reference and target samplings.

[0281] According to an embodiment, the method includes using the gradient values and / or the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

[0282] According to an embodiment, the method includes using a combination of the gradient values and the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

[0283] Thus, the CCCM model may include one or more gradients and / or position values calculated from spatial samples. The gradients and / or one or more position values may be used together with or instead of one or more spatial samples used in the CCCM model.

[0284] According to an embodiment, the method includes signaling the use of the gradient values and / or the position values of the spatial samples in a bitstream including or associated with the predicted block unit of the image.

[0285] Thus, the use of the gradient value and / or the position value may be signaled in a bitstream including encoded image data or along a bitstream including encoded image data. Alternatively, the use of the gradient value and / or the position value may be predefined, for example, in a standard specification. Yet alternatively, the use of the gradient value and / or the position value may be derived at the decoder side.

[0286] In the following, various embodiments are mainly described in the context of enhancing the performance of a convolutional cross-component model (CCCM). It should be understood that the CCCM method is given as an example in the present invention, and the proposed methods and embodiments may be used in any other method with a similar concept, such as local illumination compensation (LIC) and other intra- or inter-frame prediction methods. It should also be noted that cross-channel prediction may be performed from one chrominance channel to another chrominance channel (e.g., Cb to Cr or vice versa), or it may be from a chrominance channel to luminance (Cb / Cr to Y) or vice versa.

[0287] Various embodiments may be implemented in various filter kernels of an M-tap filter, such as Figure 8 those shown, Figure 8 illustrating 3-tap vertical filtering (1×3 (1D vertical)), 3-tap horizontal filtering (3×1 (1D horizontal)), 5-tap cross filter, and 25-tap diamond filter using North (N, up), East (E, right), South (S, down), West (W, left), and Center (C) symbols.

[0288] Gradient value

[0289] According to an embodiment, the method includes predicting the target sample of at least one color channel of an image block unit using the following cross-component prediction model:

[0290] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad hor + P n+1 ·grad ver + P n+2 ·grad diag +…+ P M-1 ·β

[0291] where pred is the predicted chrominance sample, P is the model coefficient for each tap of the M-tap filter, (C, N, S, W, E) are examples of spatial samples for (center, north, south, west, east) positions, and grad hor , grad ver , grad diagThey are the directional gradient values in the horizontal, vertical, and diagonal directions, respectively, and β is an offset or bias value.

[0292] In this document, the gradient values can be calculated in different ways. For example, the gradient values can be calculated in each direction:

[0293] grad hor = W - E;

[0294] grad ver = N - S;

[0295] grad diag1 = NW - SE;

[0296] grad diag2 = NE - SW;

[0297] where NW, SW, NE, and SE indicate samples at the northwest, southwest, northeast, and southeast positions, respectively.

[0298] It is also possible to use more than two samples to calculate the directional gradient:

[0299] grad hor = (2 × W + NW + SW) - (2 × E + NE + SE);

[0300] grad ver = (2 × N + NW + NE) - (2 × S + SW + SE);

[0301] grad diag1 = (2 × NW + N + W) - (2 × SE + S + E);

[0302] grad diag2 = (2 × NE + N + E) - (2 × SW + S + W);

[0303] In the above examples, the weighting value 2 is given to the sample center in each direction. However, for example, different weighting values can be used for each sample. The value(s) of the weighting parameter(s) used can be fixed, or can be signaled for each block (e.g., as a weighting index). Alternatively, the weighting parameter can be determined on the encoder and decoder sides based on certain factors such as block size, availability of reference samples, coding mode of one or more adjacent blocks in the neighboring blocks, coding mode of the reference block in the reference channel, etc.

[0304] According to an embodiment, the method includes using the absolute value of the gradient value instead of the signature value.

[0305] According to an embodiment, the method includes using the sum or average of the gradient values in different directions as a parameter in the cross-component prediction model.

[0306] For example, this document may use the average, weighted average, or sum of horizontal and vertical gradients, or the sum or average of different diagonal gradients, such as follows:

[0307] grad1 = abs(grad hor ) + abs(grad ver );

[0308] grad2 = abs(grad diag1 ) + abs(grad diag2 );

[0309] grad3 = abs(grad hor ) + abs(grad ver ) + abs(grad diag1 ) + abs(grad diag2 );

[0310] Where abs() indicates the absolute value of the gradient.

[0311] Another example with the average value of the gradient:

[0312] grad1 = (abs(grad hor ) + abs(grad ver )) / 2;

[0313] grad2 = (abs(grad diag1 ) + abs(grad diag2 )) / 2;

[0314] grad3 = (abs(grad hor ) + abs(grad ver ) + abs(grad diag1 ) + abs(grad diag2 )) / 4.

[0315] According to an embodiment, the method includes using both the absolute and signed versions of the gradient values in the cross-component prediction model. For example, the absolute value of the horizontal gradient and the signed version of the vertical gradient can be used in the model.

[0316] According to an embodiment, the method includes using one or more gradient values in a non-linear function and using the output of the non-linear function as a parameter in a cross-component prediction model. For example, the square root or power of two gradients among one or more gradients can be used. In another example, the sum, weighted sum, or average of non-linear functions having one or more gradient information can be used as a parameter in a filter. In yet another example, one or more outputs of the non-linear function using one or more gradient information among the gradient information as inputs can be combined with one or more outputs of the one or more non-linear functions in the non-linear function using one or more spatial samples among the spatial samples as inputs. The non-linear functions for each input type can be the same or different. For example, for spatial samples, the non-linear function can be a square root function, while for gradient information, it can be an N-th power function, and vice versa. In another example, one or more gradient information among the gradient information can be combined with one or more spatial samples among the spatial samples, and then the combined version is used as an input to the non-linear function.

[0317] According to an embodiment, the gradient information to be used in a filter is determined based on a consistency metric. For example, a cross-component model can be applied in a specific region of a reference sample and this metric can be calculated by determining how the presence of gradient information affects prediction accuracy.

[0318] According to an embodiment, a decision regarding the input type of the non-linear function used in a cross-component model can be made based on gradient values. For example, the sum, average, or weighted average of one or more gradient information among the gradient information can be calculated using a reference region sample and if the value is greater than / less than or equal to a certain threshold, the input to the non-linear function can be one or more gradient information among the gradient information or one or more spatial samples among the spatial samples.

[0319] According to an embodiment, when chroma subsampling is enabled, gradient values are obtained on downsampled luminance samples or using a full luminance grid.

[0320] According to an embodiment, the method includes combining luminance downsampling and a gradient filter into a single filter operation.

[0321] According to an embodiment, the method includes using higher-order differentials of gradient values as parameters in a cross-component prediction model.

[0322] In other words, for example, a first-order gradient is obtained iteratively, and then a second-order (or even higher) gradient is obtained. For example, a higher-order gradient can be obtained using one iteration. For example, the following second-order gradient is generated in the horizontal and vertical directions:

[0323] grad hor2 =(WW - W)-(W - E)=WW - 2W + E;

[0324] grad ver2 =(NN - N)-(N - S)=NN - 2N + S;

[0325] Or in some cases,

[0326] grad hor2 =(W - E)-(E - EE)=E - 2E + EE;

[0327] grad ver2 =(N - S)-(S - SS)=N - 2S + SS.

[0328] The three - step gradients for horizontal and vertical directions can be defined as

[0329] grad hor3 =(WW - W)-(W - E)-(W - E)+(E - EE)=WW - 3W + 3E - EE;

[0330] grad ver3 =(NN - N)-(N - S)-(N - S)+(S - SS)=NN - 3N + 3S - SS.

[0331] Note that higher - order versions follow binary coefficients. Higher - order gradients can be calculated in any direction (i.e., horizontal, vertical, or diagonal).

[0332] According to an embodiment, any angle instead of the fixed 45 - degree diagonal can be used to obtain the diagonal gradient.

[0333] According to an embodiment, the gradient direction can correspond to the direction of one or more intra - angle prediction modes in the angle frame. For example, the intra - angle prediction mode from one or more of adjacent blocks, co - located luminance blocks can be used. In another example, texture analysis methods (such as DIMD and TIMD methods) applied to some or all of the reconstructed reference samples and / or some or all of the reconstructed samples in the reference channels can be used to obtain the intra - angle prediction method for gradient direction decision.

[0334] Position value

[0335] According to an embodiment, the method includes using the following cross - component prediction model to predict the target sample of at least one color channel of an image block unit:

[0336] pred = P0·C+P1·N+P2·S+P3·W+P4·E+...+P n ·X+P n+1 ·Y+…+P M-1 ·β

[0337] Where pred is the predicted chrominance sample, P are the model coefficients for each tap in the model, (C, N, S, W, E) are examples of spatial samples for the (center, north, south, west, east) positions, and X and Y are the horizontal and vertical coordinates of the center sample, respectively, and β is the offset or bias value.

[0338] In this document, the position value can be calculated in different ways. For example, the coordinates of the sample relative to the upper left corner of the block, or the coordinates of the sample relative to the upper left corner of the reference region used to derive the filter parameters, or the coordinates of the sample relative to the upper left corner of the picture, or the coordinates of the sample relative to the upper left corner of the CTU / tile / fragment / sub-picture can be used.

[0339] According to an embodiment, the method includes using the horizontal and / or vertical position of the center sample.

[0340] According to an embodiment, the method includes using the horizontal and / or vertical positions of the (multiple) spatially adjacent samples together with or instead of the position information of the center sample.

[0341] Therefore, the position information (horizontal and / or vertical positions) of the (multiple) spatially adjacent samples (e.g., north, south, west, and east) can be used together with or instead of the position information of the center tap sample. The use of the position information in the filter can be determined based on the availability of the reference samples from the adjacent blocks, the block size (height and width), the coding mode of one or more of the adjacent blocks, the coding mode of the reference block in the reference channel, etc.

[0342] According to an embodiment, the use of the position information in the filter is determined based on the magnitude or absolute value of the model coefficients related to the sample coordinates.

[0343] According to an embodiment, the use of the position information in the filter is determined based on a consistency metric. For example, a cross-component model can be applied in a specific region of the reference sample and this metric can be calculated by determining how the presence of the position information affects the prediction accuracy.

[0344] According to an embodiment, the method includes using one or more of the position values in a non-linear function and using the output of the non-linear function as a parameter in the cross-component prediction model.

[0345] For example, a polynomial version of the position information can be used. For example, the CCCM can be an M-tap filter as follows:

[0346] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P m ·X n +Pm+1 ·Y n +P m+2 ·X n-1 +P m+3 ·Y n-1 +…+P k ·X+P k+1 ·Y+…+P M-1 ·β

[0347] According to the embodiment, a filter is selected from two or more candidate filter sets, where at least one candidate filter includes position information and at least one other candidate filter does not include position information. For example, it can be determined that the first filter includes at least model coefficients related to horizontal gradient, vertical gradient, x coordinate, and y coordinate; the second filter can be determined to include at least model coefficients related to horizontal gradient, vertical gradient, first diagonal gradient, and second diagonal gradient. In the case of this example, if at least one of the model coefficients related to the coordinates has an absolute value greater than the threshold, the first filter can be selected, otherwise the second filter can be selected.

[0348] As an example of using both one or more pieces of position information and one or more pieces of gradient information in the position information together with one or more spatial sample values in the prediction model or using them instead of one or more spatial sample values in the prediction model, CCCM can be implemented as an M - tap filter, for example, as follows:

[0349] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad hor +P n+1 ·grad ver +P n+2 ·grad diag +…+P k ·X+P k+1 ·Y+…+P M-1 ·β;

[0350] An alternative implementation for the M - tap filter can be as follows:

[0351] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad hor +P n+1 ·grad ver +P n+2 ·grad diag +…+P k ·X 2 +P k+1 ·Y2 +P k+2 ·X·Y + P k+3 ·X + P k+4 ·Y + … + P M-1 ·β

[0352] According to an embodiment, the method includes using a rate - distortion optimization (RDO) algorithm to determine a combination of the gradient values and / or the position values of the filter and the spatial samples, wherein the gradient values and / or the position values of the spatial samples will be used with or in place of the spatial sample values in a cross - component prediction model.

[0353] Thus, different filters with different combinations of the above - mentioned embodiments can be defined and the best performing version can be decided based on a rate - distortion optimization (RDO) algorithm on the encoder and / or decoder side. In the case where the selection is made on the encoder side, the index or flag of the selected version can be signaled to the decoder in the bitstream.

[0354] According to an embodiment, a combination of (multiple) gradients and position information is decided based on a consistency metric. For example, a cross - component model can be applied in a specific region of a reference sample and the metric can be calculated by determining how the presence of the combination of (multiple) gradients and position information affects the prediction accuracy.

[0355] According to an embodiment, the final prediction of a block can be obtained by combining two or more different versions of the above - mentioned embodiments. For example, two models can be derived: a first model can be derived from the left - hand side samples of a reference sample using one or more position and / or gradient information; and a second model can be derived from the above - mentioned reference sample using one or more of the position and gradient information. The position and gradient information in the first and second models can be the same or different or they can have one or more identical position and / or gradient information.

[0356] It can be considered that the spatial component represents the low - frequency sub - band of a luminance - chrominance model, while the gradient component corresponds to the high - frequency sub - band. Therefore, different training regions can be considered for the two different sub - bands.

[0357] According to an embodiment, a larger training region can be used to obtain coefficients for the spatial component. Then the obtained spatial coefficients are applied to a smaller training region and the resulting prediction is subtracted from the smaller training region before solving for the gradient coefficients.

[0358] According to an embodiment, the spatial component can be obtained using only the left - hand side or the upper - hand side of a training region. Then the obtained spatial coefficients are applied to the full training region and the resulting prediction is subtracted from the full training region before solving for the gradient coefficients.

[0359] According to an embodiment, only the left or upper side of the training region can be used to obtain gradient components. A spatial coefficient is obtained over the entire training region and subtracted from other training regions before calculating the gradient coefficient.

[0360] According to an embodiment, when using a larger training region, the spatial component can use a further subsampled luminance and chrominance grid. As a result, low-frequency and high-frequency components can be separated.

[0361] According to an embodiment, higher-order gradients can be regarded as partitions of high-frequency subbands. Thus, further partitioning of the training region can be applied to different orders of gradients.

[0362] According to an embodiment, in the presence of more than one cross-component model for predicting a block, classification parameters can be determined based on one or more of gradient and / or position information. For example, different models can be derived and applied to samples where the gradient is greater than or less than a specific threshold. In another example, different models can be derived and applied according to the magnitude of the gradient in each direction.

[0363] According to an embodiment, a transform selection, such as LFNST, applied to a block predicted using one or more of the described embodiments can be determined based on one or more of the gradient values. For example, the LFNST transform type can be selected based on the magnitude and direction of one or more of the gradient information used in the prediction model.

[0364] An apparatus according to one aspect includes components for receiving an image block unit of a frame, the image block unit including samples in color channels, where the color channels include at least one chrominance channel and one luminance channel; components for reconstructing the samples of the luminance channel of the image block unit; components for determining a reference region for target samples for predicting at least one color channel of the image block unit, where the reference region includes one or more of the following reference samples, the reference samples in adjacent blocks in the current color channel / frame, in adjacent blocks of corresponding blocks in a reference color channel / frame, and / or within corresponding blocks in a reference color channel / frame; components for determining one or more gradient values and / or one or more position values of spatial samples in the reference region used in a cross-component prediction model; and components for using the cross-component prediction model to predict the target samples of at least one color channel of the image block unit based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

[0365] According to an embodiment, the apparatus includes components for using the gradient values and / or the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

[0366] According to an embodiment, the apparatus includes components for using or substituting for the spatial sample values in the cross-component prediction model a combination of the gradient values and the position values of the spatial samples.

[0367] According to an embodiment, the apparatus includes components for signaling the use of the gradient values and / or the position values of the spatial samples, which are indicated in or with a bitstream including a predicted image block unit.

[0368] According to an embodiment, the apparatus includes components for predicting a target sample of at least one color channel of an image block unit using the following cross-component prediction model:

[0369] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad hor + P n+1 ·grad ver + P n+2 ·grad diag +…+ P M-1 ·β

[0370] where pred is the predicted chrominance sample, P are the model coefficients for each tap of an M-tap filter, (C, N, S, W, E) are examples of spatial samples at (center, north, south, west, east) positions, and grad hor , grad ver , grad diag are the directional gradient values in the horizontal, vertical, and diagonal directions respectively, and β is an offset value.

[0371] According to an embodiment, the apparatus includes components for using the sum or average of the gradient values in different directions as a parameter in the cross-component prediction model.

[0372] According to an embodiment, the apparatus includes components for using one or more of the gradient values in a non-linear function and using the output of the non-linear function as a parameter in the cross-component prediction model.

[0373] According to an embodiment, the apparatus includes components for using higher-order differentials of the gradient values as parameters in the cross-component prediction model.

[0374] According to an embodiment, the apparatus includes components for predicting a target sample of at least one color channel of an image block unit using the following cross-component prediction model:

[0375] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·X + P n+1 ·Y + … + P M-1 ·β

[0376] where pred is the predicted chrominance sample, P are the model coefficients for each tap of the M - tap filter, (C, N, S, W, E) are examples of spatial samples for the (center, north, south, west, east) positions, and X and Y are the horizontal and vertical coordinates of the center sample respectively, and β is the offset value.

[0377] According to an embodiment, the apparatus includes means for using the horizontal and / or vertical position of the center sample.

[0378] According to an embodiment, the apparatus includes means for using the horizontal and / or vertical positions of the (multiple) spatially adjacent samples together with or instead of the position information of the center sample.

[0379] According to an embodiment, the apparatus includes means for using one or more of the position values in a non - linear function and using the output of the non - linear function as a parameter in a cross - component prediction model.

[0380] According to an embodiment, the apparatus includes means for using a rate - distortion optimization (RDO) algorithm to determine a combination of the gradient value and / or the position value of the filter and the spatial samples, the combination of the gradient value and / or the position value of the spatial samples being used together with or instead of the spatial sample values in a cross - component prediction model.

[0381] As another aspect, there is provided an apparatus including: at least one processor and at least one memory, the at least one memory storing code which, when executed by the at least one processor, causes the apparatus to at least perform: receiving an image block unit of a frame, the image block unit including samples in color channels, where the color channels include at least one chrominance channel and one luminance channel; reconstructing the samples of the luminance channel of the image block unit; determining a reference region for a target sample for predicting at least one color channel of the image block unit, where the reference region includes one or more of the following reference samples, the reference samples being in adjacent blocks in the current color channel / frame, in adjacent blocks of the corresponding block in the reference color channel / frame, and / or within the corresponding block in the reference color channel / frame; determining one or more gradient values and / or one or more position values of the spatial samples in the reference region to be used in a cross-component prediction model; and predicting the target sample of at least one color channel of the image block unit using the cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

[0382] According to an embodiment, the apparatus includes code which causes the apparatus to use the gradient value and / or the position value of the spatial sample together with or instead of the spatial sample value in the cross-component prediction model.

[0383] According to an embodiment, the apparatus includes code which causes the apparatus to use a combination of the gradient value and the position value of the spatial sample together with or instead of the spatial sample value in the cross-component prediction model.

[0384] According to an embodiment, the apparatus includes code which causes the apparatus to indicate the use of the gradient value and / or the position value of the spatial sample when included in or signaled along a bitstream including the predicted image block unit.

[0385] According to an embodiment, the apparatus includes code which causes the apparatus to predict the target sample of at least one color channel of the image block unit using the following cross-component prediction model:

[0386] pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad hor + P n+1 ·grad ver + P n+2 ·grad diag +…+ P M-1 ·β

[0387] where is the predicted chrominance sample, P are the model coefficients for each tap of the M - tap filter, (C, N, S, W, E) are examples of spatial samples at the (center, north, south, west, east) positions, and are the directional gradient values in the horizontal, vertical, and diagonal directions respectively, and is the offset value.

[0388] According to an embodiment, the apparatus includes code that causes the apparatus to use the sum or average of gradient values in different directions as a parameter in the cross - component prediction model.

[0389] According to an embodiment, the apparatus includes code that causes the apparatus to use one or more of the gradient values in a non - linear function and use the output of the non - linear function as a parameter in the cross - component prediction model.

[0390] According to an embodiment, the apparatus includes code that causes the apparatus to use a higher - order differential of the gradient value as a parameter in the cross - component prediction model.

[0391] According to an embodiment, the apparatus includes code that causes the apparatus to predict the target sample of at least one color channel of the image block unit using the following cross - component prediction model:

[0392] pred = P0·C + P1·N + P2·S + P3·W + P4·E+...+P n ·X + P n+1 ·Y + …+P M-1 ·β

[0393] where pred is the predicted chrominance sample, P are the model coefficients for each tap of the M - tap filter, (C, N, S, W, E) are examples of spatial samples at the (center, north, south, west, east) positions, and X and Y are the horizontal and vertical coordinates of the center sample respectively, and β is the offset value.

[0394] According to an embodiment, the apparatus includes code that causes the apparatus to use the horizontal and / or vertical positions of the center sample.

[0395] According to an embodiment, the apparatus includes code that causes the apparatus to use the horizontal and / or vertical positions of (multiple) spatially adjacent samples together with the position information of the center sample or in place of the position information of the center sample.

[0396] According to an embodiment, the apparatus includes code that causes the apparatus to use one or more of the position values in a non - linear function and use the output of the non - linear function as a parameter in the cross - component prediction model.

[0397] According to an embodiment, the apparatus includes code that causes the apparatus to use a rate distortion optimization (RDO) algorithm to determine a combination of the gradient value and / or the position value of a filter and a spatial sample, the combination of the gradient value and / or the position value of the filter and the spatial sample to be used with or in place of the spatial sample value in a cross-component prediction model.

[0398] Such an apparatus may include, for example, any Figure 1 , Figure 2 , Figure 4a and Figure 4b of the functional units disclosed for implementing the embodiments.

[0399] The apparatus further includes code stored in the at least one memory that, when executed by the at least one processor, causes the apparatus to perform one or more of the embodiments disclosed herein.

[0400] Figure 13 is a graphical representation of an example multimedia communication system in which various embodiments may be implemented. A data source 1510 provides a source signal in an analog, uncompressed digital, or compressed digital format or any combination of these formats. The encoder 1520 may include or be coupled with preprocessing, such as data format conversion and / or filtering of the source signal. The encoder 1520 encodes the source signal into an encoded media bitstream. It should be noted that the bitstream to be decoded may be received directly or indirectly from a remote device located within virtually any type of network. Additionally, the bitstream may be received from local hardware or software. The encoder 1520 is capable of encoding more than one media type, such as audio and video, or may require more than one encoder 1520 to encode different media types of the source signal. The encoder 1520 may also obtain synthetically generated inputs, such as graphics and text, or it may be capable of generating an encoded bitstream of synthetic media. In the following, only the processing of one encoded media bitstream of one media type is considered to simplify the implementation. However, it should be noted that typically a live broadcast service includes several streams (usually at least one audio, video, and text caption stream). It should also be noted that the system may include multiple encoders, but only one encoder 1520 is shown in the figure to simplify the implementation without loss of generality. It should be further understood that while the text and examples contained herein may specifically describe the encoding process, those skilled in the art will understand that the same concepts and principles also apply to the corresponding decoding process and vice versa.

[0401] The encoded media bitstream can be transmitted to the storage device 1530. The storage device 1530 can include any type of mass storage to store the encoded media bitstream. The format of the encoded media bitstream in the storage device 1530 can be a basic self - contained bitstream format, or one or more encoded media bitstreams can be encapsulated into a container file, or the encoded media bitstream can be encapsulated into a suitable segment format for DASH (or a similar streaming system) and stored as a sequence of segments. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figure) can be used to store one or more media bitstreams in the file and create file - format metadata, which can also be stored in the file. The encoder 1520 or the storage device 1530 can include a file generator, or the file generator is operatively attached to the encoder 1520 or the storage device 1530. Some systems operate "in real - time", i.e., omit storage and transmit the encoded media bitstream directly from the encoder 1520 to the transmitter 1540. Then the encoded media bitstream can be transmitted to the transmitter 1540 (also referred to as the server) as needed. The format used in transmission can be a basic self - contained bitstream format, a packet - stream format, a suitable segment format for DASH (or a similar streaming system), or one or more encoded media bitstreams can be encapsulated into a container file. The encoder 1520, the storage device 1530, and the server 1540 can reside in the same physical device, or they can be included in separate devices. The encoder 1520 and the server 1540 can operate with live real - time content, in which case the encoded media bitstream is generally not permanently stored, but is buffered in the content encoder 1520 and / or the server 1540 for a short period of time to smooth out variations in processing delay, transmission delay, and the encoded media bit rate.

[0402] The server 1540 sends the encoded media bitstream using a communication protocol stack. The stack can include, but is not limited to, one or more of the Real - Time Transport Protocol (RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet - oriented, the server 1540 encapsulates the encoded media bitstream as packets. For example, when using RTP, the server 1540 encapsulates the encoded media bitstream into RTP packets according to the RTP payload format. Generally, each media type has a dedicated RTP payload format. It should be noted again that the system can include multiple servers 1540, but for simplicity, the following embodiments only consider one server 1540.

[0403] If the media content is encapsulated in a container file for the storage device 1530 or for inputting data to the transmitter 1540, the transmitter 1540 may include or be operatively attached to a "transmission file parser" (not shown in the figure). In particular, if the container file is not transmitted in such a way, but at least one of the encoded media bitstreams contained therein is encapsulated for transmission via a communication protocol, the transmission file parser locates the appropriate portion of the encoded media bitstream for transmission via the communication protocol. The transmission file parser may also assist in creating the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may contain encapsulation instructions, such as hint tracks in ISOBMFF, for encapsulating at least one of the media bitstreams contained in the media bitstream over the communication protocol.

[0404] The server 1540 may or may not be connected to the gateway 1550 via a communication network, which may be, for example, a CDN, the Internet, and / or a combination of one or more access networks. The gateway may also or alternatively be referred to as a middlebox. For DASH, the gateway may be an edge server (of a CDN) or a web proxy. It should be noted that the system may generally include any number of gateways or the like, but for simplicity, only one gateway 1550 is considered in the following embodiments. The gateway 1550 may perform different types of functions, such as converting a packet flow from one communication protocol stack to another, merging and splitting data streams, and manipulating the data stream according to the downlink and / or receiver capabilities, such as controlling the bit rate of the forwarded stream according to the main downlink network conditions. In various embodiments, the gateway 1550 may be a server entity.

[0405] The system includes one or more receivers 1560, which are typically capable of receiving, demodulating the transmitted signal, and de-encapsulating it into an encoded media bitstream. The encoded media bitstream can be transmitted to a recording storage device 1570. The recording storage device 1570 can include any type of mass storage to store the encoded media bitstream. The recording storage device 1570 can alternatively or additionally include computational memory, such as random access memory. The format of the encoded media bitstream in the recording storage device 1570 can be a basic self-contained bitstream format, or one or more encoded media bitstreams can be encapsulated into a container file. If there are multiple encoded media bitstreams associated with each other (such as an audio stream and a video stream), then a container file is typically used and the receiver 1560 includes or is attached to a container file generator that generates the container file from the input streams. Some systems operate "in real time", i.e., the recording storage device 1570 is omitted and the encoded media bitstream is transmitted directly from the receiver 1560 to the decoder 1580. In some systems, only the most recent portion of the stream is recorded, for example, a 10-minute excerpt of the stream is retained in the recording storage device 1570, while any earlier recorded data is discarded from the recording storage device 1570.

[0406] The encoded media bitstream can be transmitted from the recording storage device 1570 to the decoder 1580. If there are multiple encoded media bitstreams, such as an audio stream and a video stream, that are associated with each other and encapsulated into a container file, or a single media bitstream is encapsulated in a container file, for example, for easier access, a file parser (not shown in the figure) is used to de-encapsulate each encoded media bitstream from the container file. The recording storage device 1570 or the decoder 1580 can include the file parser, or the file parser is attached to the recording storage device 1570 or the decoder 1580. It should also be noted that the system can include multiple decoders, but only one decoder 1570 is discussed here to simplify the implementation without loss of generality.

[0407] The encoded media bitstream can be further processed by the decoder 1570, and the output of the decoder 1570 is one or more uncompressed media streams. Finally, the renderer 1590 can reproduce the uncompressed media stream using, for example, a speaker or a display. The receiver 1560, the recording storage device 1570, the decoder 1570, and the renderer 1590 can reside in the same physical device, or they can be included in separate devices.

[0408] The transmitter 1540 and / or the gateway 1550 may be configured to perform a switch between different representations, such as a switch between different viewports for 360-degree video content, a view switch, bitrate adaptation, and / or fast start, and / or the transmitter 1540 and / or the gateway 1550 may be configured to select the transmitted (multiple) representations. The switch between different representations may occur for a variety of reasons, such as in response to a request from the receiver 1560 or the main conditions of the network through which the bitstream is transmitted, such as throughput. In other words, the receiver 1560 may initiate a switch between representations. A request from the receiver may be, for example, a request for a segment or sub-segment from a different representation than the previous one, a request for a change in the transmitted scalability layer and / or sub-layer, or a request for a change in a rendering device with different capabilities than the previous one. A request for a segment may be an HTTP GET request. A request for a sub-segment may be an HTTP GET request with a byte range. Additionally or alternatively, bitrate adjustment or bitrate adaptation may be used, for example, in a streaming service to provide a so-called fast start, where the bitrate of the transmitted stream is lower than the channel bitrate after starting or randomly accessing the stream, in order to facilitate immediate start of playback and achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions. Bitrate adaptation may include multiple operations of switching up and down representations or layers that occur in various orders.

[0409] The decoder 1580 may be configured to perform a switch between different representations, such as a switch between different viewports for 360-degree video content, a view switch, bitrate adaptation, and / or fast start, and / or the decoder 1580 may be configured to select the transmitted (multiple) representations. The switch between different representations may occur for a variety of reasons, such as to achieve faster decoding operations or to adapt the transmitted bitstream (e.g., in terms of bitrate) to the main conditions of the network through which the bitstream is transmitted (such as throughput). For example, if the device including the decoder 1580 is multitasking and uses computing resources for other purposes than decoding the video bitstream, then faster decoding operations may be required. In another example, when playing back content at a speed faster than the normal playback speed (e.g., twice or three times faster than the conventional real-time playback speed), faster decoding operations may be required.

[0410] In the foregoing, some embodiments have been described with reference to and / or using the terms of HEVC and / or VVC. It should be understood that the embodiments can be similarly implemented with any video encoder and / or video decoder.

[0411] In the case where example embodiments have been described above with reference to an encoder, it should be understood that the resulting bitstream and decoder can have corresponding elements therein. Similarly, in the case where example embodiments have been described with reference to a decoder, it should be understood that the encoder can have a structure and / or computer program for generating the bitstream to be decoded by the decoder. For example, some embodiments have been described that involve generating a prediction block as part of encoding. Embodiments can be similarly implemented by generating a prediction block as part of decoding, except that encoding parameters different from those determined by the encoder (such as horizontal and vertical offsets) are decoded from the bitstream.

[0412] The embodiments of the present invention described above describe the codec according to separate encoder and decoder devices to facilitate understanding of the processes involved. However, it should be understood that the device, structure, and operation can be implemented as a single encoder-decoder device / structure / operation. In addition, the encoder and decoder can share some or all common elements.

[0413] Although the examples above describe embodiments of the present invention operating within a codec in an electronic device, it should be understood that the present invention as defined in the claims can be implemented as part of any video codec. Thus, for example, embodiments of the present invention can be implemented in a video codec that can achieve video encoding via a fixed or wired communication path.

[0414] Therefore, a user equipment can include a video codec such as that described above in the embodiments of the present invention. It should be understood that the term user equipment is intended to cover any suitable type of wireless user equipment, such as a mobile phone, a portable data processing device, or a portable web browser.

[0415] In addition, elements of a public land mobile network (PLMN) can also include a video codec as described above.

[0416] Generally, various embodiments of the present invention can be implemented using hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software that can be executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well understood that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0417] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Further in this regard, it should be noted that any box of the logical flow in the figures may represent a program step, or an interconnected logical circuit, box, and function, or a combination of program steps and logical circuits, boxes, and functions. The software may be stored on a physical medium such as a memory chip or a memory block implemented within the processor, a magnetic medium such as a hard disk or a floppy disk, and an optical medium such as, for example, a DVD and its data variants, a CD.

[0418] The memory may be of any type adapted to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. As a non-limiting example, the data processor may be of any type adapted to the local technical environment and may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture.

[0419] Embodiments of the present invention may be implemented in various components such as integrated circuit modules. The design of an integrated circuit is generally a highly automated process. Sophisticated and powerful software tools can be used to transform a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0420] Programs such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design of San Jose, California use well-established design rules and pre-stored design module libraries to automatically route conductors and place components on a semiconductor chip. Once the design of the semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transferred to a semiconductor manufacturing facility or "fab" for fabrication.

[0421] The above embodiments have provided a complete and informative description of exemplary embodiments of the present invention by way of exemplary and non-limiting examples. However, various modifications and adaptations may become apparent to those skilled in the relevant art in view of the above description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications taught by the present invention will still fall within the scope of the present invention.

Claims

1. An apparatus, comprising: means for receiving a picture block unit of a frame, the picture block unit including samples in color channels, wherein the color channels include at least one chrominance channel and one luminance channel; means for reconstructing samples of the luminance channel of the picture block unit; means for determining a reference region for target samples for predicting at least one color channel of the picture block unit, wherein the reference region includes one or more of the following reference samples, the reference samples being in adjacent blocks in the current color channel / frame, and / or in adjacent blocks of the corresponding block in the reference color channel / frame, and / or within the corresponding block in the reference color channel / frame; means for determining one or more gradient values and / or one or more position values of spatial samples in the reference region to be used in a cross-component prediction model; and means for predicting the target samples of at least one color channel of the picture block unit using the cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

2. The apparatus according to claim 1, comprising: means for using the gradient values and / or the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

3. The apparatus according to claim 1 or 2, comprising: means for using a combination of the gradient values and the position values of the spatial samples together with or instead of the spatial sample values in the cross-component prediction model.

4. The apparatus according to any one of the preceding claims, comprising: means for signaling in a bitstream including the predicted picture block unit or together with the bitstream to indicate the use of the gradient values and / or the position values of the spatial samples.

5. The apparatus according to any one of the preceding claims, comprising: means for predicting the target samples of at least one color channel of the picture block unit using the following cross-component prediction model: pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·grad nor + P n+1 ·grad ver + P n+2 ·grad diag +…+ P M-1 ·β where pred is the predicted chrominance sample, P is the model coefficient for each tap in the M-tap filter, (C, N, S, W, E) are examples of spatial samples for the (center, north, south, west, east) positions, and grad hor , grad ver , grad diag are the directional gradient values in the horizontal, vertical, and diagonal directions respectively, and β is the offset or bias value.

6. The apparatus according to any one of the preceding claims, comprising: means for using the sum or average of the gradient values in different directions as a parameter in the cross-component prediction model.

7. The apparatus according to any one of the preceding claims, comprising: means for using one or more of the gradient values in a non-linear function and using the output of the non-linear function as a parameter in the cross-component prediction model.

8. The apparatus according to claim 7, comprising: means for using higher-order differentials of the gradient values as the parameter in the cross-component prediction model.

9. The apparatus according to any one of the preceding claims, comprising: means for predicting the target samples of at least one color channel of the picture block unit using the following cross-component prediction model: pred = P0·C + P1·N + P2·S + P3·W + P4·E +... + P n ·X + P n+1 ·Y + … + P M-1 ·β where pred is the predicted chrominance sample, P is the model coefficient for each tap in the M-tap filter, (C, N, S, W, E) are examples of spatial samples for the (center, north, south, west, east) positions, and X and Y are the horizontal and vertical coordinates of the center sample respectively, and β is an offset or bias value.

10. The apparatus according to any one of the preceding claims, comprising: means for using the horizontal position and / or the vertical position of the center sample.

11. The apparatus according to claim 10, comprising: means for using the horizontal position and / or the vertical position of one or more spatially adjacent samples together with the position information of the center sample or in place of the position information of the center sample.

12. The apparatus according to any one of the preceding claims, comprising: means for using one or more of the position values in a non-linear function and using the output of the non-linear function as the parameter in the cross-component prediction model.

13. The apparatus according to any one of the preceding claims, comprising: means for using a rate-distortion optimization RDO algorithm to determine a combination of the gradient value and / or the position value of the filter and the spatial sample, the combination of the gradient value and / or the position value of the filter and the spatial sample being used together with the spatial sample value in the cross-component prediction model or in place of the spatial sample value in the cross-component prediction model.

14. A method, comprising: receiving an image block unit of a frame, the image block unit including samples in a color channel, wherein the color channel includes at least one chrominance channel and one luminance channel; reconstructing the samples of the luminance channel of the image block unit; determining a reference region for predicting a target sample of at least one color channel of the image block unit, wherein the reference region includes one or more of the following reference samples, the reference samples being in adjacent blocks in the current color channel / frame, and / or in adjacent blocks of the co-located block in the reference color channel / frame, and / or within the co-located block in the reference color channel / frame; determining one or more gradient values and / or one or more position values of the spatial samples in the reference region to be used in a cross-component prediction model; and predicting the target sample of at least one color channel of the image block unit using the cross-component prediction model based at least on the one or more gradient values and / or the one or more position values of the spatial samples in the reference region.

15. The method according to claim 14, comprising: using the gradient value and / or the position value of the spatial sample together with the spatial sample value in the cross-component prediction model or in place of the spatial sample value in the cross-component prediction model.