Apparatus, method, and computer program for video encoding and decoding
By mapping luminance residuals to chrominance prediction using co-located reference areas and functions, the method addresses the inefficiencies in existing video coding standards, enhancing inter-component prediction and improving compression efficiency for both luminance and chrominance channels.
Patent Information
- Application Number
- JP2025501366
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-13
- Filing Date
- 2023-05-26
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-05-26
AI Technical Summary
Existing video coding standards prioritize luminance blocks over chrominance blocks in inter-coding, leading to less effective chrominance-specific tools and inefficient compression due to the lack of a robust method to model chrominance using co-located motion compensated luminance and chrominance.
A method is introduced to directly map luminance residuals to chrominance prediction, utilizing co-located reference areas and mapping functions to improve inter-component prediction, applicable in single-tree or dual-tree blocks, with additional methods for enhancing mapping robustness and alternative configurations.
This approach enhances the efficiency of video encoding by improving the inter-component residual model, leading to better compression and de-correlation of color components, thus optimizing the encoding process for both luminance and chrominance channels.
Smart Images

Figure 2025525513000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus, a method and a computer program for video encoding and decoding. [Background technology]
[0002] In video coding, video and image samples are typically encoded using a color representation such as YUV or YCbCr, which consists of one luminance (luma) channel and two chrominance (chroma) channels. In such cases, the luminance channel, which primarily represents the illumination of the scene, is typically coded at a particular resolution, while the chrominance channels, which typically represent the differences between specific color components, are coded at a second resolution, often lower than that of the luminance signal. The intent of this type of differential representation is to de-correlate the color components, allowing the data to be compressed more efficiently.
[0003] The Versatile Video Coding (VVC / H.266) standard uses a cross-component linear model (CCLM) as a linear model for predicting samples in chroma channels (e.g., Cb and Cr). Model parameters are derived based on reconstructed samples in the neighborhood of a chroma block, adjacent samples at the same position within a luma block, and reconstructed samples within the same position within the luma block. An improved version of cross-component prediction, known as CCLM, uses a two-dimensional (2D) filter kernel to derive a luma-to-chroma model. Filter coefficients are derived at the decoder side using the reconstructed input data and a set of chroma samples.
[0004] Inter-coding in modern video coding standards (e.g., H.266) uses single-tree motion-compensated prediction, in which areas from one or more reference frames are copied to a target frame using subpixel interpolation. To account for variations in illumination and / or color differences between the reference and target frames, additional tools such as weighted prediction (WP), bi-prediction with CU-level weight (BCW), combined inter and intra prediction (CIIP), or local illumination compensation (LIC) are used. Single-tree coding signals one set of partition information per block, and this partition is shared between the luminance and chrominance components. Due to the paucity of chrominance-specific coding tools, inter-coding often implicitly prioritizes luminance blocks compared to chrominance. Even when chrominance-specific tools such as CCCM or CCLM are used, they are less effective because chroma blocks in the reference frame often have higher quality compared to intra-predicted inter-component blocks. To benefit from inter-component tools in inter prediction, there needs to be a method to model chrominance using primarily co-located motion compensated luminance and chrominance. Summary of the Invention
[0005] Herein, an improved method is introduced to achieve a better inter-component residual model for prediction in order to at least alleviate the above problems.
[0006] The scope of protection claimed for various embodiments of the invention is set out in the independent claims. The embodiments and features described herein that do not fall within the scope of the independent claims, if any, should be interpreted as examples that help to understand various embodiments of the invention.
[0007] According to some embodiments, a method is provided for directly mapping luminance residual to chrominance prediction. In particular, this mapping may be applied in single-tree inter-coded blocks where the same partitioning is shared between luminance and chrominance blocks. An intuitive method for performing this mapping after decoding the luminance residual between single-tree blocks is also disclosed. This method can also be extended to dual-tree blocks or any other partitioning schemes. Additional methods for improving the robustness of the mapping procedure are also provided. Alternative configurations, such as chrominance-to-chrominance prediction, are also described.
[0008] A method according to a first aspect includes receiving an image block unit of a frame, the image block unit including samples in a first chrominance channel, a second chrominance channel, and one luminance channel; defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel; performing motion compensation prediction of the luminance component, the first chrominance component, and the second chrominance component using the reference frame and the motion vector; obtaining a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component using the co-located luminance prediction and chrominance prediction; reconstructing a luminance residual; and obtaining a first chrominance prediction by using the first function and the luminance residual and obtaining a second chrominance prediction by using the second function and the luminance residual.
[0009] an image processing unit for processing a frame including a first chrominance channel, a second chrominance channel, and a luminance channel; a reference frame for defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel; a motion compensation prediction unit for performing motion compensation on the luminance component, the first chrominance component, and the second chrominance component using the reference frame and the motion vector; a co-located luminance prediction unit for obtaining a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component using the co-located luminance prediction unit and the chrominance residual; a reconstructing unit for reconstructing a luminance residual; and a second chrominance prediction unit for obtaining a first chrominance prediction by using the first function and the luminance residual and a second chrominance prediction by using the second function and the luminance residual.
[0010] According to an embodiment, the device comprises means for downsampling the co-located luminance component to match the sub-sampled chrominance grid.
[0011] According to an embodiment, the device comprises means for filtering the co-located luminance components using a one-dimensional or two-dimensional filter.
[0012] According to an embodiment, the apparatus comprises means for using either a first mapping function or a second mapping function to map the luminance component to the first chrominance component or the second chrominance component.
[0013] According to an embodiment, the apparatus comprises means for using either a first mapping function or a second mapping function to map a first chrominance component to a second chrominance component.
[0014] According to an embodiment, the apparatus comprises means for signaling the use of a cross-component residual model (CCRM) from an encoder to a decoder at the level of a transform unit, a prediction unit, a coding unit, a coding tree unit, a slice, a frame, or a sequence.
[0015] According to an embodiment, the device comprises means for clipping or spatially filtering the luminance residual before applying the inter-component model.
[0016] According to an embodiment, the apparatus comprises means for determining the type and amount of filtering to be applied to the luminance residual separately for the first chrominance component and the second chrominance component.
[0017] According to an embodiment, the apparatus comprises means for determining the type and amount of filtering to be applied to the luminance residual based on the output of the first mapping function or the output of the second mapping function.
[0018] According to an embodiment, the device comprises means for applying different types and amounts of filtering to the luma residual at a sub-block level or at a sample level.
[0019] According to an embodiment, the apparatus comprises means for using different mapping functions in block prediction based on one or more thresholds derived for luminance, luminance residual, or chrominance.
[0020] According to an embodiment, the device comprises means for mapping chrominance components to luminance components.
[0021] According to an embodiment, the apparatus comprises means for mapping a first chrominance component to a second chrominance component.
[0022] According to an embodiment, the device comprises means for upsampling the input of the mapping function to match the luminance grid.
[0023] According to an embodiment, the device comprises means for upsampling the output of the mapping function to match the luminance grid.
[0024] According to an embodiment, the apparatus comprises means for mapping a first chrominance component to a second chrominance component.
[0025] Thus, the apparatus and computer readable storage medium having code stored thereon is arranged to perform one or more of the above-described methods and related embodiments, as described above.
[0026] For a better understanding of the present invention, reference will now be made, by way of example, to the accompanying drawings in which: [Brief explanation of the drawings]
[0027] [Figure 1] 1 is a diagram illustrating a schematic diagram of an electronic device incorporating an embodiment of the present invention; [Figure 2] FIG. 1 illustrates schematically user equipment suitable for employing embodiments of the present invention; [Figure 3] FIG. 2 further illustrates, in schematic form, electronic devices employing embodiments of the present invention connected using wireless and wired network connections. [Figure 4a] 1 illustrates schematically an encoder and decoder suitable for implementing embodiments of the present invention; [Figure 4b] 1 illustrates schematically an encoder and decoder suitable for implementing embodiments of the present invention; [Figure 5] FIG. 10 shows the locations of samples used to derive the parameters of the inter-component linear model. [Figure 6a]1A and 1B show examples of classification of luma samples into two classes in the sample domain and in the spatial domain, respectively. [Figure 6b] 1A and 1B show examples of classification of luma samples into two classes in the sample domain and in the spatial domain, respectively. [Figure 7] FIG. 10 is a diagram illustrating an example of four reference lines adjacent to a prediction block. [Figure 8] FIG. 1 illustrates a matrix weighted intra prediction process. [Figure 9] FIG. 2 illustrates a flowchart of a method for predicting samples in at least one color channel according to an embodiment of the present invention. [Figure 10] FIG. 10 illustrates an example of a co-located reference sample area consisting of reconstructed luma and chroma samples, defined with respect to both luma and chroma, in accordance with an embodiment of the present invention. [Figure 11] 1A-1C illustrate various example filter kernel dimensions according to one approach. [Figure 12] FIG. 10 illustrates upper and left neighboring blocks used in joint inter and intra prediction (CIIP) weight derivation, according to one approach. [Figure 13] FIG. 1 shows a schematic diagram of an exemplary multimedia communication system in which various embodiments may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0028] In the following, suitable apparatus and possible mechanisms for predicting chroma samples will be described in more detail. In this regard, reference is first made to Figures 1 and 2, where Figure 1 shows a block diagram of a video encoding system according to an example embodiment as a schematic block diagram of an exemplary apparatus or electronic device 50 that can incorporate a codec according to an embodiment of the present invention. Figure 2 shows the layout of an apparatus according to an example embodiment. Next, the elements of Figures 1 and 2 will be described.
[0029] The electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communications system, but it will be understood that embodiments of the present invention may be implemented in any electronic device or apparatus that may require encoding and / or decoding of video images.
[0030] The device 50 may include a housing 30 for housing and protecting the device. The device 50 may further include a display 32 in the form of a liquid crystal display. In other embodiments of the invention, the display may be any suitable display technology suitable for displaying images or video. The device 50 may further include a keypad 34. In other embodiments of the invention, any suitable data or user interface mechanism may be employed. For example, the user interface may be embodied as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0031] The device may include a microphone 36 or any suitable audio input, which may be a digital or analog signal input. The device 50 may further include an audio output device, which in embodiments of the present invention may be earphones 38, a speaker, or any one of an analog or digital audio output connection. The device 50 may also include a battery (or in other embodiments of the present invention, the device may be powered by any suitable mobile energy device, such as a solar cell, a fuel cell, or a clockwork generator). The device may further include a camera capable of recording or capturing images and / or video. The device 50 may further include an infrared port for short-range line-of-sight communication to other devices. In other embodiments, the device 50 may further include any suitable short-range communication solution, such as, for example, a Bluetooth wireless connection or a USB / Firewire wired connection.
[0032] The device 50 may include a controller 56, processor, or processor circuitry for controlling the device 50. The controller 56 may be connected to a memory 58, which in embodiments of the present invention may store both data in the form of image data and audio data and / or may also store instructions for execution by the controller 56. The controller 56 may further be connected to a codec circuit 54 suitable for performing encoding and decoding of the audio and video data or aiding in the encoding and decoding performed by the controller.
[0033] The device 50 may further comprise a card reader 48 and a smart card 46, e.g. a UICC and a UICC reader, suitable for providing user information and authentication information for authentication and authorization of the user on the network.
[0034] The device 50 may include a radio interface circuit 52 connected to the controller and suitable for generating wireless communication signals for communication with, for example, a cellular communication network, a wireless communication system, or a wireless local area network. The device 50 may further include an antenna 44 connected to the radio interface circuit 52 for transmitting radio frequency signals generated by the radio interface circuit 52 to and receiving radio frequency signals from other devices.
[0035] Apparatus 50 may include a camera capable of recording or detecting individual frames, which are then passed to codec 54 or a controller for processing. The apparatus may receive video image data for processing from another device before transmission and / or storage. Apparatus 50 may also receive images for encoding / decoding wirelessly or via a wired connection. The structural elements of apparatus 50 described above represent examples of means for performing the corresponding functions.
[0036] 3, an example of a system in which embodiments of the present invention may be utilized is shown. System 10 comprises a plurality of communication devices capable of communicating over one or more networks. System 10 may comprise any combination of wired or wireless networks, including, but not limited to, wireless cellular telephone networks (such as GSM, UMTS, CDMA networks), wireless local area networks (WLANs) as defined by any of the IEEE 802.x standards, Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet.
[0037] System 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the present invention.
[0038] For example, the system depicted in Figure 3 shows a representation of a cellular network 11 and the Internet 28. Connections to the Internet 28 may include, but are not limited to, long-range wireless connections, short-range wireless connections, and various wired connections, including, but not limited to, telephone lines, cable lines, power lines, and similar communication paths.
[0039] Exemplary communication devices shown in system 10 may include, but are not limited to, electronic devices or apparatus 50, a combination personal digital assistant (PDA) and mobile phone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, and a notebook computer 22. The apparatus 50 may be stationary or mobile when carried by a moving individual. The apparatus 50 may also be located in a vehicle, including, but not limited to, a car, truck, taxi, bus, train, boat, airplane, bicycle, motorcycle, or any similar suitable vehicle.
[0040] Embodiments may also be implemented in set-top boxes, i.e., digital television receivers, which may or may not have display or wireless capabilities, in tablets or (laptop) personal computers (PCs) including hardware or software or encoder / decoder combination implementations, in various operating systems, and in chipsets, processors, DSPs, and / or embedded systems that provide hardware / software based encoding.
[0041] Some or further devices may send and receive calls and messages and communicate with a service provider via a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that enables communication between the cellular network 11 and the Internet 28. The system may include additional and different types of communication devices.
[0042] Communication devices may communicate using various transmission technologies, including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time division multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, and any similar wireless communication technology. Communication devices involved in implementing various embodiments of the present invention may communicate using various mediums, including, but not limited to, radio, infrared, laser, cabled connections, and any suitable connection.
[0043] In telecommunications and data networks, a channel may refer to either a physical channel or a logical channel. A physical channel may refer to a physical transmission medium such as a wire, while a logical channel may refer to a logical connection on a multiplexed medium that can carry multiple logical channels. A channel may be used to convey an information signal, e.g., a bit stream, from one or more senders (or transmitters) to one or more receivers.
[0044] The MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format for carrying audio, video, and other media, as well as program or other metadata, in a multiplexed stream. Packet identifiers (PIDs) are used to identify elementary streams (also known as packetized elementary streams) within a TS. Thus, logical channels within an MPEG-2 TS may be considered to correspond to specific PID values.
[0045] Available media file format standards include the ISO Base Media File Format (ISO / IEC 14496-12, sometimes abbreviated as ISOBMFF) and the NAL unit structured video file format (ISO / IEC 14496-15), which is derived from ISOBMFF.
[0046] A video codec consists of an encoder that converts input video into a compressed representation suitable for storage / transmission, and a decoder that can convert the compressed video representation back into a displayable form. The video encoder and / or video decoder may also be separate from each other, i.e., they may not need to form a codec. Typically, an encoder discards some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate).
[0047] Typical hybrid video encoders, e.g., many implementations of ITU-T H.263 and H.264, encode video information in two stages. First, pixel values within a specific image area (or "block") are predicted, e.g., by motion compensation (locating and indicating an area in one of the previously coded video frames that closely corresponds to the block being coded) or by spatial means (using pixel values surrounding the block being coded in a specified manner). Next, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g., the Discrete Cosine Transform (DCT) or a variant thereof), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (image quality) and the size of the resulting coded video representation (file size or transmission bit rate).
[0048] In temporal prediction, the source of prediction is a previously decoded image (also known as a reference image). In intra block copy (IBC, also known as intra block copy prediction), prediction is applied similarly to temporal prediction, except that the reference image is the current image and only previously decoded samples may be referenced in the prediction process. Inter-layer prediction or inter-view prediction may be applied similarly to temporal prediction, except that the reference image is an image decoded from another scalable layer or another view, respectively. In some cases, inter-prediction may refer only to temporal prediction, but in other cases, inter-prediction may collectively refer to any of intra-block copy, inter-layer prediction, and inter-view prediction, provided that they are performed in the same or similar process as temporal prediction. Inter-prediction or temporal prediction may also be referred to as motion compensation or motion-compensated prediction.
[0049] Motion compensation can be performed with either full-sample accuracy or sub-sample accuracy. For full-sample accuracy motion compensation, motion can be represented as a motion vector containing integer values for horizontal and vertical displacement, and the motion compensation process effectively copies samples from a reference image using those displacements. For sub-sample accuracy motion compensation, the motion vector is represented by fractional or decimal values for the horizontal and vertical components of the motion vector. When a motion vector references a non-integer position in the reference image, a sub-sample interpolation process is typically invoked to calculate predicted sample values based on the reference sample and a selected sub-sample position. The sub-sample interpolation process typically consists of horizontal filtering, which compensates for horizontal offsets with respect to full sample positions, followed by vertical filtering, which compensates for vertical offsets with respect to full sample positions. However, in some environments, vertical processing can occur before horizontal processing.
[0050] Inter-prediction, sometimes called temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter-prediction, the source of the prediction is a previously decoded image. Intra-prediction exploits the fact that adjacent pixels in the same image are likely to be correlated. Intra-prediction can be performed in the spatial domain or the transform domain, i.e., either sample values or transform coefficients can be predicted. Intra-prediction is usually used in intra-coding, where inter-prediction is not applied.
[0051] One result of the encoding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be more efficiently entropy coded if they are first predicted from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from a spatially neighboring motion vector, and only the difference relative to this motion vector predictor may be coded. Prediction of coding parameters and intra-prediction are sometimes collectively referred to as in-picture prediction.
[0052] Figures 4a and 4b show an encoder and decoder suitable for employing embodiments of the present invention. A video codec consists of an encoder that converts the input video into a compressed representation suitable for storage / transmission, and a decoder that can convert the compressed video representation back into a displayable form. Typically, the encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (i.e., lower bit rate). An example of the encoding process is shown in Figure 4a. Figure 4a shows the image (I) to be encoded. n ), the predicted representation of the image block (P ’n ), prediction error signal (D n ), the reconstructed prediction error signal (D ’n ), the preliminary reconstructed image (I ’n ), the final reconstructed image (R ’n ), transformation (T) and inverse transformation (T -1 ), quantization (Q) and dequantization (Q -1 ), entropy encoding (E), reference frame memory (RFM), inter prediction (P inter ), intra prediction (P intra ), mode selection (MS) and filtering (F).
[0053] An example of the decoding process is shown in Figure 4b. Figure 4b shows a predicted representation (P ’n ), the reconstructed prediction error signal (D ’n ), the preliminary reconstructed image (I ’n ), the final reconstructed image (R ’n ), inverse transformation (T -1 ), inverse quantization (Q -1 ), entropy decoding (E -1 ), Reference Frame Memory (RFM), Prediction (either Inter or Intra) (P), and Filtering (F).
[0054] Many hybrid video encoders encode video information in two stages. First, pixel values within a particular image area (or "block") are predicted, for example, by motion compensation means (locating and indicating an area in one of the previously coded video frames that corresponds exactly to the block being coded) or by spatial means (using pixel values surrounding the block being coded in a specified way). In the first stage, predictive coding may be applied, for example, as so-called sample prediction and / or so-called syntactic prediction.
[0055] In sample prediction, pixel or sample values within a particular image area or "block" are predicted. These pixel or sample values may be predicted using, for example, one or more of motion compensation or intra-prediction mechanisms.
[0056] Motion compensation mechanisms (sometimes called inter-prediction, temporal prediction, or motion-compensated temporal prediction or motion-compensated prediction or MCP) involve finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being coded. Inter-prediction can reduce temporal redundancy.
[0057] Intra-prediction, where pixel or sample values can be predicted by spatial mechanisms, involves finding and representing spatial domain relationships. Intra-prediction takes advantage of the fact that adjacent pixels in the same image are likely to be correlated. Intra-prediction can be performed in the spatial domain or the transform domain, i.e., either sample values or transform coefficients can be predicted. Intra-prediction is typically used in intra-coding, where inter-prediction is not applicable.
[0058] In syntax prediction, sometimes called parameter prediction, syntax elements and / or syntax element values and / or variables derived from syntax elements are predicted from previously encoded (decoded) syntax elements and / or previously derived variables. Non-limiting examples of syntax prediction are given below.
[0059] In motion vector prediction, for example, motion vectors for inter prediction and / or inter-view prediction may be differentially coded with respect to a block-specific predicted motion vector. In many video codecs, predicted motion vectors are created in a predefined manner, for example, by calculating the median of the encoded or decoded motion vectors of neighboring blocks. Another method for creating motion vector predictions, sometimes called advanced motion vector prediction (AMVP), is to generate a list of candidate predictions from neighboring and / or co-located blocks in a temporal reference picture and signal a selected candidate as the motion vector predictor. In addition to predicting motion vector values, reference indices of previously coded / decoded pictures may be predicted. The reference indices are typically predicted from neighboring and / or co-located blocks in a temporal reference picture. Differential coding of motion vectors is typically disabled across slice boundaries.
[0060] For example, block division from CTU to CU and to PU may be predicted.
[0061] In filter parameter prediction, for example, the filtering parameters of the sample adaptive offset may be predicted.
[0062] Prediction techniques that use image information from previously coded images may also be called inter-prediction methods, and are sometimes called temporal prediction and motion compensation. Prediction techniques that use image information within the same image may also be called intra-prediction methods.
[0063] Next, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the differences in pixel values using a specified transform (e.g., the discrete cosine transform (DCT) or a variant thereof), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, an encoder can control the balance between the precision of the pixel representation (image quality) and the size of the resulting coded video representation (file size or transmission bit rate). Video codecs may also provide a transform skip mode that the encoder can choose to use. In transform skip mode, the prediction error is coded in the sample domain, for example, by deriving a per-sample difference value relative to a particular neighboring sample and encoding the per-sample difference value using an entropy coder.
[0064] Entropy encoding / decoding may be performed in many ways. For example, both the encoder and the decoder may apply context-based encoding / decoding, which modifies the context state of coding parameters based on previously encoded / decoded coding parameters. The context-based coding may be, for example, context adaptive binary arithmetic coding (CABAC) or context-based variable length coding (CAVLC), or any similar entropy coding. Alternatively or additionally, entropy encoding / decoding may be performed using a variable-length coding scheme, such as Huffman coding / decoding or Exp-Golomb coding / decoding. Decoding coding parameters from an entropy-coded bitstream or codeword is sometimes referred to as parsing.
[0065] The phrase "along the bitstream" (e.g., indicating that it is along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that associates the out-of-band data with the bitstream. The phrase "decoding along the bitstream" or similar phrases may refer to decoding referenced out-of-band data associated with the bitstream (which may be obtained from out-of-band transmission, signaling, or storage). For example, instructions along the bitstream may refer to metadata within a container file that encapsulates the bitstream.
[0066] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of the International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard was published by both parent standards organizations and is known as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10 (also known as MPEG-4 Part 10 Advanced Video Coding (AVC)). There have been multiple versions of the H.264 / AVC standard that incorporate new enhancements or features into the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0067] Version 1 of the H.265 / HEVC standard, also known as High Efficiency Video Coding (HEVC), was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standards organizations and is known as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2 (also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC)). Later versions of H.265 / HEVC include extensions for scalable, multiview, range-fidelity, 3D, and screen content coding, which are sometimes abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0068] Versatile Video Coding (VVC) (MPEG-I Part 3), also known as ITU-T H.266, is a video compression standard developed by the Joint Video Experts Team (JVET) of the Moving Picture Experts Group (MPEG) (formally ISO / IEC JTC1 SC29 WG11) and the Video Coding Experts Group (VCEG) of the International Telecommunication Union (ITU) as the successor to HEVC / H.265.
[0069] Some important definitions, bitstreams, and coding structures, and concepts of H.264 / AVC and HEVC are described in this section as examples of video encoders, decoders, encoding methods, decoding methods, and bitstream structures on which embodiments may be implemented. Some important definitions, bitstreams, and coding structures, and concepts of H.264 / AVC are the same as those of HEVC and are therefore described together below. Aspects of the present invention are not limited to either H.264 / AVC or HEVC, and a description is given on one possible basis on which the present invention may be partially or fully realized.
[0070] Like many previous video coding standards, H.264 / AVC and HEVC specify bitstream syntax and semantics, as well as the decoding process for error-free bitstreams. The encoding process is not specified, but the encoder must produce a conforming bitstream. The conformance of the bitstream and decoder can be verified using a Hypothetical Reference Decoder (HRD). These standards include coding tools that help address transmission errors and losses, but the use of these tools in encoding is optional, and for erroneous bitstreams, the decoding process is not specified.
[0071] The basic unit for input to an H.264 / AVC or HEVC encoder and output of an H.264 / AVC or HEVC decoder, respectively, is an image. An image given as input to an encoder is sometimes called a source image, and an image decoded by a decoder is sometimes called a decoded image.
[0072] The source image and the decoded image each consist of one or more sample sequences, such as one of the following sets of sample sequences: - Luma(Y) only (monochrome). - Luma and two chroma (YCbCr or YCgCo). - Green, Blue, and Red (GBR, also known as RGB). - Arrays representing other unspecified monochrome or tristimulus color samplings (e.g., YZX, also known as XYZ).
[0073] In H.264 / AVC and HEVC, an image may be either a frame or a field. A frame contains a matrix of luma samples and, possibly, corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as an encoder input if the source signal is interlaced. The chroma sample array may be absent (thus, monochrome sampling may be used), or the chroma sample array may be subsampled when compared to the luma sample array. Chroma formats may be summarized as follows: - In monochrome sampling, there is only one sample array that can nominally be considered the luma array. - In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. - In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array. - In 4:4:4 sampling, if separate color planes are not used, each of the two chroma arrays has the same height and width as the luma array.
[0074] In H.264 / AVC and HEVC, the sample array can be encoded into the bitstream as separate color planes, and each coded color plane can be decoded separately from the bitstream. When separate color planes are used, each one of the color planes is processed separately (by the encoder and / or decoder) as an image using monochrome sampling.
[0075] A partition may be defined as the division of a set into subsets such that each element of the set is contained in exactly one of the subsets.
[0076] The following terminology may be used when describing HEVC encoding and / or decoding operations: A coding block may be defined as an N×N block of samples, for some value N, such that the division of a coding tree block into coding blocks is a partition. A coding tree block (CTB) may be defined as an N×N block of samples, for some value N, such that the division of a component into coding tree blocks is a partition. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples for an image containing a three sample array, or a coding tree block of samples for a monochrome image, or an image coded using a syntax structure used to code three separate color planes and samples. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding tree blocks of chroma samples for an image containing a three sample array, or a coding block of samples for a monochrome image, or an image coded using a syntax structure used to code three separate color planes and samples. A CU with the largest allowed size may be named an LCU (largest coding unit) or a coding tree unit (CTU), and a video image is divided into non-overlapping LCUs.
[0077] A CU consists of one or more prediction units (PUs), which define the prediction process for samples within the CU, and one or more transform units (TUs), which define the prediction error coding process for samples within the CU. Typically, a CU consists of a square block of samples with a size selectable from a predefined set of possible CU sizes. Each PU and TU may be further divided into smaller PUs and TUs, respectively, to increase the granularity of the prediction and prediction error coding processes. Each PU has associated prediction information (e.g., motion vector information for an inter-predicted PU and intra-prediction direction information for an intra-predicted PU) that defines the type of prediction applied to pixels within that PU.
[0078] Each TU may be associated with information (e.g., including DCT coefficient information) describing the prediction error decoding process of the samples within that TU. Typically, whether prediction error coding is applied to each CU is signaled at the CU level. If there is no prediction error residual associated with a CU, the TU for that CU may be considered to be absent. The division of an image into CUs and the division of CUs into PUs and TUs is typically signaled in the bitstream, allowing a decoder to reproduce the intended structure of these units.
[0079] In HEVC, an image may be divided into tiles, which are rectangular and contain an integer number of LCUs. In HEVC, the division into tiles forms a regular grid, and the heights and widths of the tiles differ from each other by at most one LCU. In HEVC, a slice is defined as an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit. In HEVC, a slice segment is defined as an integer number of coding tree units contained in a single NAL unit that is consecutively ordered in a tile scan. The division of each image into slice segments is a partition. In HEVC, an independent slice segment is defined as a slice segment in which values of syntax elements in its slice segment header are not inferred from values of preceding slice segments, and a dependent slice segment is defined as a slice segment in which values of some syntax elements in its slice segment header are inferred from values of preceding independent slice segments in decoding order. In HEVC, a slice header is defined as the slice segment header of an independent slice segment that is either the current slice segment or an independent slice segment preceding the current dependent slice segment, and a slice segment header is defined as the part of a coded slice segment that contains data elements for the first or all coding tree units represented in the slice segment. CUs are scanned in the raster scan order of LCUs within a tile, or, if tiles are not used, within an image. Within an LCU, CUs have a specific scan order.
[0080] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predictive representation of pixel blocks (using motion or spatial information created by the encoder and stored in the compressed representation), and prediction error decoding (the inverse operation of prediction error encoding, which recovers the quantized prediction error signal in the spatial pixel domain). After applying the prediction and prediction error decoding means, the decoder puts together the prediction and prediction error signals (pixel values) to form an output video frame. The decoder (and encoder) may also apply additional filtering means to improve the quality of the output video before passing the output video for display and / or storing it as a predictive reference for the next frame in the video sequence.
[0081] The filtering may include, for example, one or more of deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF). H.264 / AVC includes deblocking, while HEVC includes both deblocking and SAO.
[0082] In typical video codecs, motion information is indicated using motion vectors associated with each motion-compensated image block, such as a prediction unit. Each of these motion vectors represents the displacement of an image block in an image being coded (at the encoder side) or decoded (at the decoder side) and a predicted source block in one of the previously coded or decoded images. To efficiently represent motion vectors, they are usually differentially coded with respect to a block-specific predicted motion vector. In typical video codecs, predicted motion vectors are created in a predefined manner, for example, by calculating the median of the encoded or decoded motion vectors of neighboring blocks. Another method for creating motion vector predictions is to generate a list of candidate predictions from neighboring and / or co-located blocks in a temporal reference image and signal a selected candidate as a motion vector predictor. In addition to predicting motion vector values, it is possible to predict which reference image will be used for motion-compensated prediction, and this prediction information may be represented, for example, by a reference index of a previously coded / decoded image. Reference indices are typically predicted from neighboring and / or co-located blocks in temporal reference pictures. Furthermore, typical high-efficiency video codecs employ an additional motion information encoding / decoding mechanism, often called merge / merge mode, in which all motion field information, including motion vectors and corresponding reference picture indices for each available reference picture list, is predicted and used without modification. Similarly, predicting motion field information is performed using motion field information from neighboring and / or co-located blocks in temporal reference pictures, and the motion field information to be used is signaled in a list of motion field candidate lists filled with the motion field information from available neighboring / co-located blocks.
[0083] In typical video codecs, the motion-compensated prediction residual is first transformed using a transform kernel (such as a DCT) and then encoded. This is because there is often still some correlation between the residual and the transform, and it can often help to reduce this correlation, resulting in more efficient encoding.
[0084] Video coding standards and specifications may allow an encoder to divide an encoded image into coded slices or the like. In-picture prediction is typically disabled across slice boundaries. Thus, slices may be considered a method for dividing an encoded image into independently decodable elements. In H.264 / AVC and HEVC, in-picture prediction may be disabled across slice boundaries. Thus, slices may be considered a method for dividing an encoded image into independently decodable elements, and thus slices are often considered the basic unit of transmission. Often, an encoder can indicate in the bitstream what type of in-picture prediction is turned off across slice boundaries, and the decoder operation takes this information into account, for example, when concluding which prediction sources are available. For example, samples from neighboring CUs may be considered unavailable for intra prediction if the neighboring CUs reside in different slices.
[0085] The basic unit for the output of an H.264 / AVC or HEVC encoder and the input of an H.264 / AVC or HEVC decoder, respectively, is the Network Abstraction Layer (NAL) unit. For transport over a packet-oriented network or storage in a structured file, NAL units may be encapsulated in packets or similar structures. H.264 / AVC and HEVC specify a byte-stream format for transmission or storage environments that do not provide a framing structure. The byte-stream format separates NAL units from each other by prepending a start code to each NAL unit. To prevent false detection of NAL unit boundaries, the encoder implements a byte-oriented start code emulation prevention algorithm and adds an emulation prevention byte to the NAL unit payload if a start code would otherwise occur. To enable easy gateway operation between packet-oriented and stream-oriented systems, start code emulation prevention may always be performed, regardless of whether the byte-stream format is used. A NAL unit may be defined as a syntactic structure that contains an indication of the type of data it tracks and the bytes containing that data in the form of RBSPs interspersed with emulation prevention bytes, if necessary. A raw byte sequence payload (RBSP) may be defined as a syntactic structure that contains an integer number of bytes encapsulated in a NAL unit. An RBSP is either empty or has the form of a data bit string containing syntax elements, followed by an RBSP stop bit, followed by zero or more subsequent bits equal to 0.
[0086] A NAL unit consists of a header and a payload. In H.264 / AVC and HEVC, the header of a NAL unit indicates the type of the NAL unit.
[0087] HEVC uses a two-byte NAL unit header for all specified NAL unit types. The NAL unit header contains one reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plus1 indication of temporal level (which may be required to be greater than or equal to one), and a six-bit nuh_layer_id syntax element. The temporal_id_plus1 syntax element may be considered a temporal identifier for the NAL unit, and a zero-based TemporalId variable may be derived as TemporalId = temporal_id_plus1 - 1. The abbreviation TID may be used interchangeably with the TemporalId variable. A TemporalId equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 must be non-zero to prevent start code emulation involving the two bytes of the NAL unit header. A bitstream created by excluding all VCL NAL units with a TemporalId greater than or equal to a selected value and including all other VCL NAL units remains conformant. Thus, a picture with TemporalId equal to tid_value does not use a picture with TemporalId greater than tid_value as an inter-prediction reference. A sub-layer or temporal sub-layer may be defined as a temporal scalable layer (or temporal layer TL) of a temporal scalable bitstream, consisting of VCL NAL units and associated non-VCL NAL units with a particular value of the TemporalId variable. nuh_layer_id may be understood as a scalability layer identifier.
[0088] NAL units can be classified into Video Coding Layer (VCL) NAL units and non-VCL NAL units. A VCL NAL unit is typically a coded slice NAL unit. In HEVC, a VCL NAL unit contains syntax elements that represent one or more CUs.
[0089] A non-VCL NAL unit may be, for example, one of the following types: sequence parameter set, picture parameter set, supplemental enhancement information (SEI) NAL unit, access unit delimiter, end of sequence NAL unit, end of bitstream NAL unit, or filler data NAL unit. Parameter sets may be required for the reconstruction of a decoded picture, while many of the other non-VCL NAL units are not required for the reconstruction of decoded sample values.
[0090] Parameters that do not change throughout a coded video sequence may be included in a sequence parameter set. In addition to parameters that may be needed by the decoding process, a sequence parameter set may optionally include video usability information (VUI), including parameters that may be important for buffering, picture output timing, rendering, and resource reservation. In HEVC, a sequence parameter set RBSP includes parameters that may be referenced by one or more picture parameter sets RBSP or one or more SEI NAL units containing a buffering period SEI message. A picture parameter set includes parameters that are likely not to change across multiple coded pictures. A picture parameter set RBSP may include parameters that may be referenced by coded slice NAL units of one or more coded pictures.
[0091] In HEVC, a video parameter set (VPS) may be defined as a syntax structure containing syntax elements that apply to zero or more entire coded video sequences, determined by the content of syntax elements found in the SPS referenced by syntax elements found in the PPS referenced by syntax elements found in each slice segment header.
[0092] A video parameter set RBSP may contain parameters that may be referenced by one or more sequence parameter sets RBSP.
[0093] The relationships and hierarchy among video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) may be described as follows: A VPS exists within the parameter set hierarchy and, in the context of scalability and / or 3D video, one level above an SPS. A VPS may contain parameters that are common to all slices across all (scalability or view) layers in the entire coded video sequence. An SPS contains parameters that are common to all slices in a particular (scalability or view) layer in the entire coded video sequence and may be shared by multiple (scalability or view) layers. A PPS contains parameters that are common to all slices in a particular layer representation (a representation of one scalability or view layer in one access unit) and are likely shared by all slices in multiple layer representations.
[0094] The VPS may provide information about layer dependencies in the bitstream, as well as many other pieces of information applicable to all slices across all (scalability or view) layers in the entire coded video sequence. The VPS may be considered to include two parts: a base VPS and a VPS extension, which may optionally be present.
[0095] Out-of-band transmission, signaling, or storage may additionally or alternatively be used for purposes other than resilience to transmission errors, such as ease of access or session negotiation. For example, a sample entry for a track in a file conforming to the ISO Base Media File Format may include a parameter set, while the encoded data in the bitstream is stored elsewhere in the file or in a separate file. The phrases "along the bitstream" (e.g., indicating along the bitstream) or "along the coded unit of the bitstream" (e.g., indicating along the coded tile) may be used in the claims and described embodiments to refer to out-of-band transmission, signaling, or storage in such a way that the out-of-band data is associated with the bitstream or coded unit, respectively. The phrases "decoding along the bitstream" or "along the coded unit of the bitstream," or similar phrases, may refer to decoding the referenced out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) associated with the bitstream or coded unit, respectively.
[0096] An SEI NAL unit may contain one or more SEI messages that are not required for decoding of the output image, but may be useful for related processes such as image output timing, rendering, error detection, error concealment, and resource reservation.
[0097] An encoded image is a coded representation of an image.
[0098] In HEVC, a coded picture may be defined as a coded representation of a picture, including all coding tree units of the picture. In HEVC, an access unit (AU) may be defined as a set of NAL units associated with each other according to specified classification rules, consecutive in decoding order, and including at most one picture with any particular value of nuh_layer_id. In addition to containing VCL NAL units of a coded picture, an access unit may also contain non-VCL NAL units. Such specified classification rules may, for example, associate pictures with the same output time or picture output count value with the same access unit.
[0099] A bitstream may be defined as a sequence of bits in the form of a NAL unit stream or byte stream that forms a representation of coded pictures and associated data forming one or more coded video sequences. A first bitstream may be followed by a second bitstream within the same logical channel, such as within the same file or the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams. The end of the first bitstream may be indicated by a specific NAL unit, sometimes called the end of bitstream (EOB) NAL unit, which is the last NAL unit of the bitstream. HEVC and its current draft extensions require that the EOB NAL unit have nuh_layer_id equal to 0.
[0100] In H.264 / AVC, a coded video sequence is defined as a sequence of consecutive access units in decoding order from one IDR access unit (inclusive) to the next IDR access unit (exclusive) or the end of the bitstream, whichever occurs first.
[0101] In HEVC, a coded video sequence (CVS) may be defined as a sequence of access units, consisting of, for example, an IRAP access unit with NoRaslOutputFlag equal to 1 in decoding order, followed by zero or more access units that are not IRAP access units with NoRaslOutputFlag equal to 1, including all subsequent access units, but not including any subsequent access units that are IRAP access units with NoRaslOutputFlag equal to 1. An IRAP access unit may be defined as an access unit whose base layer picture is an IRAP picture. The value of NoRaslOutputFlag is equal to 1 for every IDR picture, every BLA picture, and every IRAP picture, where the IRAP picture is the first picture of that particular layer in the bitstream in decoding order and the first IRAP picture following the end-of-sequence NAL unit with the same value of nuh_layer_id in decoding order. There may be means for providing the value of HandleCraAsBlaFlag to the decoder from an external entity, such as a player or receiver, that can control the decoder. HandleCraAsBlaFlag may be set to 1, for example, by a player that seeks to a new position in the bitstream or tunes to a broadcast and begins decoding, and then begins decoding from the CRA picture. If a CRA picture's HandleCraAsBlaFlag is equal to 1, the CRA picture is treated and decoded as if it were a BLA picture.
[0102] In HEVC, a coded video sequence may additionally or alternatively (to the above specification) be specified to end when a specific NAL unit, sometimes called an end of sequence (EOS) NAL unit, appears in the bitstream and has a nuh_layer_id equal to 0.
[0103] A group of pictures (GOP) and its characteristics may be defined as follows: A GOP can be decoded regardless of whether previous pictures have been decoded. An open GOP is a group of pictures such that when decoding starts from the first intra picture of the open GOP, pictures preceding the first intra picture in output order may not be correctly decoded. In other words, pictures of an open GOP may reference pictures belonging to a previous GOP (in inter prediction). An HEVC decoder can recognize an intra picture that starts an open GOP because a specific NAL unit type, the CRA NAL unit type, may be used for its coded slice. A closed GOP is a group of pictures such that all pictures can be correctly decoded when decoding starts from the first intra picture of the closed GOP. In other words, pictures in a closed GOP do not reference pictures in a previous GOP. In H.264 / AVC and HEVC, a closed GOP may start with an IDR picture. In HEVC, a closed GOP may also start with a BLA_W_RADL or BLA_N_LP picture. An open GOP coding structure may also be more efficient in compression compared to a closed GOP coding structure due to greater flexibility in the selection of reference pictures.
[0104] A decoded picture buffer (DPB) may be used in an encoder and / or a decoder. There are two reasons for buffering decoded pictures: for reference in inter-prediction and for reordering decoded pictures into output order. Because H.264 / AVC and HEVC provide great flexibility in both marking reference pictures and reordering outputs, separate buffers for buffering reference pictures and buffering output pictures may waste memory resources. Therefore, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. Decoded pictures may be removed from the DPB when they are no longer used as references and are no longer needed for output.
[0105] In many coding modes of H.264 / AVC and HEVC, reference pictures for inter prediction are indicated using indices into reference picture lists. The indices may be coded using variable length coding, which typically causes smaller indices to have shorter values for the corresponding syntax elements. In H.264 / AVC and HEVC, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi-predictive (B) slice, and one reference picture list (reference picture list 0) is formed for each inter-coded (P) slice.
[0106] Many encoding standards, including H.264 / AVC and HEVC, may have a decoding process for deriving a reference image index into a reference image list, which may be used to indicate which of multiple reference images is used for inter-prediction of a particular block. The reference image index may be coded into the bitstream by the encoder in some inter-coding mode, or may be derived (by the encoder and decoder) using neighboring blocks in some other inter-coding mode, for example.
[0107] The types of motion parameters or motion information may include, but are not limited to, one or more of the following types: - an indication of the type of prediction (e.g., intra-prediction, uni-prediction, bi-prediction) and / or the number of reference images, an indication of the prediction direction, such as inter-prediction (also known as temporal prediction), inter-layer prediction, inter-view prediction, view synthesis prediction (VSP), and inter-component prediction (this may be indicated per reference picture and / or per prediction type; in some embodiments, inter-view prediction and view synthesis prediction may be considered together as one prediction direction); and / or - indication of the type of reference image, such as short-term reference image and / or long-term reference image and / or inter-layer reference image (this may for example be indicated for each reference image); a reference index to a reference picture list and / or any other identifier of the reference picture (which may for example be indicated for each reference picture, the type of which may depend on the prediction direction and / or the type of reference picture, and may be accompanied by other relevant information, such as the reference picture list to which the reference index applies), - horizontal motion vector component (which may be indicated for example per prediction block or per reference index or similar); - a vertical motion vector component (which may be indicated for example per prediction block, or per reference index, or similar); one or more parameters, such as the difference in picture order count and / or the relative camera separation between a picture containing or associated with a motion parameter and its reference picture, that may be used for scaling the horizontal and / or vertical motion vector components in one or more motion vector prediction processes (the one or more parameters may, for example, be indicated for each reference picture or each reference index, or similar); - coordinates of the block to which the motion parameters and / or motion information apply, e.g., the coordinate of the top-left sample of the block in luma samples, - The extent of the block (e.g. width and height) to which the motion parameters and / or motion information apply.
[0108] Compared to previous video coding standards, the versatile video codec (H.266 / VVC) introduces several new coding tools, such as: Intra prediction - 67 Intra modes with wide-angle mode extension - Block size and mode dependent 4-tap interpolation filter - Position dependent intra prediction combination (PDPC) - Component-to-component linear model intra-prediction (CCLM) - Multiple reference line intra prediction - Intra sub-partitions - Weighted intra prediction with matrix multiplication Inter-image prediction - Copy block movements using spatial, temporal, history-based, and pairwise average merge candidates - Affine motion inter-prediction - Sub-block based temporal motion vector prediction - Adaptive motion vector resolution - 8x8 block-based motion compression for temporal motion prediction - High precision (1 / 16 pixel) motion vector storage and motion compensation with an 8-tap interpolation filter for the luma component and a 4-tap interpolation filter for the chrominance component - Triangulation - Joint intra and inter prediction - Merge with MVD (MMVD) - Symmetric MVD coding - Bidirectional optical flow - Improved decoder-side motion vectors - CU-level weighted bi-prediction Transformation, quantization, and coding of coefficients - Multiple primary transform selections using DCT2, DST7, and DCT8 - Secondary transformation of the low frequency zone - Sub-block transform of inter-predicted residuals - Dependency quantization with max QP increased from 51 to 63 - Transform coefficient coding with code data hiding - Transform skip residual coding Entropy coding - Arithmetic coding engine with adaptive double windows probability update In-loop filter - In-loop reshape - Deblocking filter using a stronger longer filter - Sample adaptive offset - Adaptive Loop Filter Screen content encoding - Current image reference with reference area constraints 360-degree video encoding - Horizontal wraparound motion compensation High-level syntax and parallelism - Reference image management with direct reference image list signaling - Tile groups containing rectangular tile groups
[0109] Partitioning in VVC is performed similarly to HEVC, i.e., each image is divided into coding tree units (CTUs). Images may also be divided into slices, tiles, bricks, and sub-images. CTUs may be divided into smaller CUs using a quadtree structure. Each CU may be divided using quadtrees and nested multi-type trees, including 3-way and 2-way partitions. However, there are specific rules for estimating partitions at image boundaries, and nested multi-type partitions do not allow redundant partitioning patterns.
[0110] In the new coding tool mentioned above, to reduce the redundancy between components, a cross-component linear model (CCLM) prediction mode is used in VVC, where chroma samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model: pred C (i,j)=α·rec L '(i,j)+β (Equation 1a)
[0111] where pred C (i,j) represents the predicted chroma sample in the CU, and rec L '(i,j) represents the downsampled and reconstructed luma sample of the same CU.
[0112] Alternatively, the following equation may be used for CCLM: pred C (i,j)=α·rec L ’(i,j) >>k+β (Equation 1b) Here, the >> operation indicates a bit shift to the right by the value k.
[0113] The CCLM parameters (α and β) are derived using up to four adjacent chroma samples and their corresponding downsampled luma samples. Assuming the dimensions of the current chroma block are W × H, W' and H' are set as follows: - When LM mode is applied, W'=W, H'=H. - When LM-A mode is applied, W'=W+H. - If LM-L mode is applied, H'=H+W.
[0114] In this specification, LM-A mode refers to linear model-top, where only the top template (i.e., sample values from adjacent positions above the CU) is used to calculate the linear model coefficients. To obtain more samples, the top template is extended to (W+H). Next, LM-L mode refers to linear model-left, where only the left template (i.e., sample values from adjacent positions to the left of the CU) is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H+W). For non-square blocks, the top template is extended to W+W, and the left template is extended to H+H.
[0115] The top adjacent positions are denoted as S[0,-1]...S[W'-1,-1], and the left adjacent positions are denoted as S[-1,0]...S[-1,H'-1]. Then, four samples are selected as follows: - S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4], if LM mode is applied and both upper and left neighboring samples are available. - If LM-A mode is applied or only upper adjacent samples are available, S[W' / 8,-1], S[3*W' / 8,-1], S[5*W' / 8,-1], S[7*W' / 8,-1]. If the -LM-L mode is applied or only the left adjacent sample is available, then S[-1,H' / 8], S[-1,3*H' / 8], S[-1,5*H' / 8], S[-1,7*H' / 8].
[0116] The four adjacent luma samples at the selected position are downsampled and compared four times to find two smaller values, x0A and x1A, and two larger values, x0B and x1B. Their corresponding chroma sample values are denoted as y0A, y1A, y0B, and y1B. Then, xA, xB, yA, and yB are derived as follows: X a =(x 0A +x 1 A +1)>>1, X b =(x 0 B +x 1 B +1)>>1, Y a =(y 0 A +y 1 A +1)>>1, Y b =(y 0 B +y 1 B +1)>>1 (equation 2)
[0117] Finally, the linear model parameters α and β are obtained according to the following equations:
number
[0118] FIG. 5 shows an example of the positions of the left and top samples and the current block samples involved in CCLM mode.
[0119] The division operation to calculate the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are expressed in exponential notation. For example, the diff is approximated using a 4-bit significant part and exponent. As a result, the table for 1 / diff is reduced to 16 elements for 16 values of the significant part, as follows: DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} (Equation 5)
[0120] This has the advantage of reducing the computational complexity as well as the memory size required to store the necessary tables.
[0121] In addition to the top and left templates being used together to calculate the linear model coefficients, they can also be used alternatively in two other LM modes, referred to as the LM_A and LM_L modes.
[0122] In LM_A mode, only the top template is used to calculate the linear model coefficients. To obtain more samples, the top template is extended to (W+H). In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H+W).
[0123] For non-square blocks, the top template is dilated to W+W and the left template is dilated to H+H.
[0124] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by a flag at the SPS level. The two downsampling filters are as follows, corresponding to "Type 0" and "Type 2" content, respectively:
[0125]
number
[0126] Note that when the upper reference line is at a CTU boundary, only one luma line (a common line buffer in intra prediction) is used to create the downsampled luma samples.
[0127] This parameter calculation is performed as part of the decoding process, not just as a search operation in the encoder. As a result, no syntax is used to communicate the values of α and β to the decoder.
[0128] For chroma intra mode coding, a total of eight intra modes are allowed for chroma intra mode coding. These modes include five conventional intra modes and three inter-component linear model modes (CCLM, LM_A, and LM_L). The chroma mode signaling and derivation process is shown in Table 1. Chroma mode coding directly depends on the intra prediction mode of the corresponding luma block. In an I slice, separate block partition structures for luma and chroma components are enabled, so one chroma block may correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.
[0129] [Table 1]
[0130] As shown in Table 2, a single binarization table is used regardless of the value of sps_cclm_enabled_flag. [Table 2]
[0131] In Table 2, the first binary number indicates whether it is standard mode (0) or LM mode (1). If it is LM mode, the next binary number indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next binary number indicates whether it is LM_L (0) or LM_A (1). In this case, when sps_cclm_enabled_flag is 0, the first binary number of the corresponding intra_chroma_pred_mode binarization table can be discarded before entropy encoding. Or in other words, the first binary number is presumed to be 0 and therefore not encoded. This single binarization table is used both when sps_cclm_enabled_flag is equal to 0 and when it is equal to 1. The first two binary numbers in Table 2 are context encoded using their own context model, and the remaining binary numbers are bypass encoded.
[0132] Additionally, to reduce luma-chroma latency in the dual tree, if a 64x64 luma coding tree node is not partitioned (and intra sub-partitions (ISP) are not used for the 64x64 CU) or is partitioned using QT, the chroma CUs in the 32x32 / 32x16 chroma coding tree node are allowed to use CCLM in the following manner: - If a 32x32 chroma node is not split or is split with QT split, all chroma CUs within the 32x32 node can use CCLM. - If a 32x32 chroma node is split using horizontal BT and the 32x16 child node is not split or uses vertical BT split, all chroma CUs within the 32x16 chroma node can use CCLM.
[0133] In all other luma and chroma coding split conditions, CCLM is not allowed for chroma CUs.
[0134] Multi-model LM (MMLM)
[0135] The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, adjacent reconstructed samples are classified into two classes using a threshold value that is the average of the luma-reconstructed samples. A linear model for each class is derived using the least-mean-square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Figures 6a and 6b show two luma-chroma models obtained for 17 luma (Y) thresholds in the sample domain and spatial domain, respectively. Each luma-chroma model has its own linear model parameters, α and β. As shown in Figure 6b, each luma-chroma model corresponds to a spatial segmentation of the content (i.e., to different objects or textures in the scene).
[0136] Convolutional cross-component model (CCCM)
[0137] An improved version of inter-component prediction, known as CCCM, uses a 2D filter kernel to derive a luma-chroma model. Using a set of reconstructed input data and chroma samples, filter coefficients are derived at the decoder side. To derive the filter coefficients, a co-located reference sample area (consisting of reconstructed luma and chroma samples) is defined for both luma and chroma, as shown in FIG. 10, where the commonly used 4:2:0 chroma downsampling is applied. The reference sample area for a particular block can be, for example, six lines above and to the left of the particular block, as shown in FIG. 10, although any number of reference lines (which can be realized by both the encoder and the decoder) can be used. In general, the reference samples can include any chroma and luma samples that have been reconstructed by both the encoder and the decoder. After the reference samples are determined, the filter coefficients can be derived using various types of linear regression tools, such as ordinary least squares estimation, orthogonal matching pursuit, optimized orthogonal matching pursuit, ridge regression, or least absolute shrinkage and selection operators.
[0138] 11 shows various example filter kernel dimensions, which can be, for example, 1x3 (1D vertical), 3x1 (1D horizontal), 3x3, 7x7, or any dimension, and can be shaped (by selecting only a subset of all possible kernel positions) as a cross or diamond (as shown in FIG. 11), or any particular shape. When referring to samples within a filter kernel, the letters N, E, S, W, C are used to denote north (top), east (right), south (bottom), west (left), and center, as shown in FIG. 11.
[0139] The overall method of reconstructing chroma samples using the convolution between the filter kernel obtained at the decoder side and a set of input data is referred to herein as the Convolutional Component-Component Model (CCCM). The following steps can be applied to perform the operation of CCCM: 1. Define a reference area at the same location on the luma and chroma components. 2. Downsample luma samples to match the chroma grid (optional). 3. Scan the luma and chroma samples of the reference area and collect available statistics (such as autocorrelation matrices and cross-correlation vectors) based on the filter shape. 4. Solve for the filter coefficients by minimizing the squared error (or any other metric) based on available statistics (such as the autocorrelation matrix and cross-correlation vectors). 5. Compute the predicted chroma blocks by convolving the downsampled luma samples with a filter kernel.
[0140] The luma samples that may be downsampled may be defined as a 2D array Y(x,y) indexed using a horizontal x-coordinate and a vertical y-coordinate. The co-located chroma samples are also defined as a 2D array C(x,y), and the filter kernel (i.e., coefficients) is defined as a 3x3 array F(i,j). At the sample level, the convolution between Y and F may be defined as follows:
[0141]
number
number
number
[0142] Similarly, a bias term can be added to the convolution as follows:
number
[0143] F and
number
[0144] Multiple reference line (MRL) intra prediction
[0145] Multiple Reference Line (MRL) intra prediction uses more reference lines for intra prediction. Figure 7 shows an example of four reference lines, where samples from segments A and F are filled with the nearest samples from segments B and E, respectively, rather than being obtained from reconstructed neighboring samples. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 3) are used.
[0146] The index of the selected reference line (mrl_idx) is signaled and used to generate the intra predictor. For reference line idx greater than 0, only the additional reference line modes are included in the MPM list, and only the mpm index is signaled, without the remaining modes. The reference line index is signaled before the intra prediction modes, and if a non-zero reference line index is signaled, planar modes are excluded from the intra prediction modes.
[0147] For the first line of a block in a CTU, MRL is disabled to prevent the use of extended reference samples outside the lines of the current CTU. Also, PDPC is disabled if additional lines are used. For MRL mode, the derivation of DC values for non-zero reference line indices in DC intra prediction modes is aligned with the derivation for reference line index 0. MRL requires the storage of three adjacent luma reference lines along with the CTU to generate the prediction. The CCLM tool also requires three adjacent luma reference lines for the downsampling filter. The definition of MLR to use the same three lines is aligned with CCLM to reduce decoder storage requirements.
[0148] Intra Subdivision (ISP)
[0149] Intra subpartitioning (ISP) divides a luma intra-prediction block vertically or horizontally into two or four subpartitions, depending on the block size. For example, the minimum block size for ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into four subpartitions. Note that Mx128 (M≦64) and 128xN (N≦64) ISP blocks can cause potential problems with a 64x64 VDPU. For example, an Mx128 CU in a single tree case includes an Mx128 luma TB and two corresponding M / 2x64 chroma TBs. When a CU uses ISP, the luma TB is divided into four Mx32 TBs (only horizontal division is possible), and each of these TBs is smaller than a 64x64 block. However, in current ISP designs, chroma blocks are not divided. Therefore, both chroma components have a size larger than a 32x32 block. Similarly, a similar situation can be created using an ISP with a 128xN CU. Therefore, these two cases are problematic for a 64x64 decoder pipeline. For this reason, the size of the CU that can use the ISP is limited to a maximum of 64x64. All subdivisions must contain at least 16 samples.
[0150] Matrix weighted intra prediction (MIP)
[0151] The matrix-weighted intra prediction (MIP) method is a new intra prediction technique added to VVC. To predict samples of a rectangular block of width W and height H, matrix-weighted intra prediction (MIP) takes as input the reconstructed adjacent boundary samples of one line of H on the left side of the block and the reconstructed adjacent boundary samples of one line of W above the block. If reconstructed samples are not available, reconstructed samples are generated in the same way as traditional intra prediction. The generation of the prediction signal is based on three steps: averaging, matrix-vector multiplication, and linear interpolation, as shown in Figure 8.
[0152] Inter Prediction in VVC
[0153] The merge list may include the following candidates: (1) Spatial MVP from spatially adjacent CUs (2) Time MVP from collocated CU (3) History-based MVP from FIFO tables (4) Pairwise average MVP (using candidates already in the list) (5) Zero's music video
[0154] Merged mode width motion vector difference (MMVD) is for signaling MVD and resolution index after signaling merge candidates.
[0155] In symmetric MVD, the motion information of list 1 is derived from the motion information of list 0 in the bi-predictive case.
[0156] In affine prediction, multiple motion vectors are indicated / signaled for different corners of a block and used to derive motion vectors for sub-blocks. In affine merge prediction, affine motion information for a block is generated based on regular or affine motion information of neighboring blocks.
[0157] In sub-block-based temporal motion vector prediction, the motion vectors of the sub-blocks of the current block are predicted from the appropriate sub-blocks in the reference frame indicated by the motion vectors of spatial neighboring blocks (if available).
[0158] In adaptive motion vector resolution (AMVR), the precision of the MVD is signaled on a per-CU basis.
[0159] In CU-level weighted bi-prediction, the index indicates the weight value of the weighted average of the two prediction blocks.
[0160] Bi-directional optical flow (BDOF) refines motion vectors in the bi-predictive case. BDOF uses the signaled motion vectors to generate two prediction blocks. Then, using the gradient values of the two prediction blocks, a motion refinement is calculated to minimize the error between the prediction blocks. The motion refinement and gradient values are used to refine the final prediction block.
[0161] CU-level Weighted Bi-Prediction (BCW) and Weighted Prediction (WP)
[0162] In HEVC, a bi-predictive signal is generated by averaging two prediction signals obtained from two different reference pictures and / or by using two different motion vectors. In VVC, the bi-predictive mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred =((8-w)*P0+w*P1+4>>3
[0163] Weighted average bi-prediction allows five weights, w∈{-2,3,4,5,10}. For each bi-predicted CU, the weight w is determined in one of two ways: (1) for non-merged CUs, the weight index is signaled after the motion vector differential; (2) for merged CUs, the weight index is estimated from neighboring blocks based on the merge candidate index. BCW is only applied to CUs that contain 256 or more luma samples (i.e., when the CU width multiplied by the CU height is 256 or more). For low-latency images, all five weights are used. For non-low-latency images, only three weights (w∈{3,4,5}) are used.
[0164] In the encoder, fast search algorithms are applied to find the weight indices without significantly increasing the encoder complexity. These algorithms can be summarized as follows: - When combined with AMVR, unequal weights are only conditionally checked for 1-pixel and 4-pixel motion vector accuracy if the current picture is a low-latency picture. - When combined with affine, affine ME is performed for unequal weights if, and only if, the affine mode is selected as the current best mode. - If the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked.
[0165] Depending on the POC distance between the current image and its reference image, the coding QP, and the temporal level, unequal weights are not searched for if certain conditions are met.
[0166] The BCW weight index is coded using one context-coded binary number followed by a bypass-coded binary number. The first context-coded binary number indicates whether equal weights are used, and if unequal weights are used, an additional binary number is signaled using bypass coding to indicate which unequal weights are used.
[0167] Weighted prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficiently encoding video content with fading. WP support was also added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. The weights and offsets of the corresponding reference pictures are then applied during motion compensation. WP and BCW are designed for various types of video content. To avoid interactions between WP and BCW, which complicates the design of VVC decoders, when a CU uses WP, the BCW weight index is not signaled and w is estimated to be 4 (i.e., equal weights are applied). For merged CUs, the weight index is estimated from neighboring blocks based on the merge candidate index. This applies to both the normal merge mode and the inherited affine merge mode. For the constructed affine merge mode, affine motion information is constructed based on the motion information of up to three blocks. The BCW index of a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV.
[0168] In VVC, CIIP and BCW cannot be applied to a CU together. If a CU is coded in CIIP mode, the BCW index of the current CU is set to 2 (e.g., equal weight).
[0169] Joint Inter and Intra Prediction (CIIP)
[0170] In VVC, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is 64 or more), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether combined inter and intra prediction (CIIP) mode is applied to the current CU. As the name suggests, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter is derived using the same inter prediction process as applied in the standard merge mode, resulting in an intra predicted signal P intra is derived according to the standard intra prediction process using planar mode. The intra prediction signal and the inter prediction signal are then combined using a weighted average, with the weight value calculated depending on the coding modes of the upper and left neighboring blocks (shown in Figure 12) as follows: - If the neighboring block above is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0. - If the left neighboring block is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0. - If (isIntraLeft+isIntraTop) is equal to 2, set wt to 3. - Else, if (isIntraLeft+isIntraTop) is equal to 1, then wt is set to 2. - Otherwise, set wt to 1.
[0171] The CIIP forecast is formed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2
[0172] Local Lighting Compensation (LIC)
[0173] LIC is an inter-prediction technique for modeling local illumination variations between a current block and its predicted block as a function of the local illumination variations between the current block template and the reference block template. The parameters of this function can be denoted by a scale α and an offset β that form a linear equation to compensate for illumination changes: α*p[x]+β, where p[x] is the reference sample pointed to by the MV at location x on the reference image. Because α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for α and β, except that the LIC flag is signaled in AMVP mode to indicate the use of LIC.
[0174] The local illumination compensation proposed in JVET-O0066 is used in ECM for uni-predictive inter CUs with the following modifications: Intra neighbor samples can be used to derive LIC parameters. LIC is disabled for blocks containing less than 32 luma samples. For both non-subblock and affine modes, the derivation of LIC parameters is performed based on the template block samples corresponding to the current CU instead of the partial template block samples corresponding to the first top-left 16x16 unit. · The reference block template samples are generated by using MC without rounding the block MV to integer pixel precision.
[0175] Inter-coding in modern video coding standards (e.g., H.266) uses single-tree motion-compensated prediction, where areas from one or more reference frames are copied to a target frame using subpixel interpolation. To account for variations in illumination and / or color differences between the reference and target frames, additional tools such as weighted prediction (WP), control unit level weighted bi-prediction (BCW), joint inter- and intra-prediction (CIIP), or local illumination compensation (LIC) are used. Single-tree coding signals one set of partition information per block, and this partition is shared between the luminance and chrominance components. Due to the paucity of chrominance-specific coding tools, inter-coding often implicitly prioritizes luminance blocks compared to chrominance. Even when chrominance-specific tools such as CCCM or CCLM are used, they are less effective because chroma blocks in the reference frame often have higher quality compared to intra-predicted inter-component blocks. To benefit from inter-component tools in inter prediction, there should be a way to model chrominance using primarily co-located motion compensated luminance and chrominance.
[0176] Currently, an improved inter-component residual model for prediction is introduced.
[0177] According to an embodiment, a method is provided for directly mapping luminance residual to chrominance prediction. In particular, this mapping may be applied in single-tree inter-coded blocks where the same partitioning is shared between luminance and chrominance blocks. An intuitive method for performing this mapping after decoding the luminance residual between single-tree blocks is disclosed. This method can also be extended to dual-tree blocks or any other partitioning schemes. Alternative configurations, such as between chrominance, are also possible.
[0178] Single-tree coding shares the same partition and motion between the luminance (Y) and chrominance (Cb, Cr) components. Typically, the reconstruction process for a single-tree inter-coded block is as follows: First, motion-compensated prediction is performed using a reference frame and sub-pixel motion vectors, then additional adjustments are made using tools such as WP, BCW, CIIP, or LIC, and then the residual is applied. These steps may be applied sequentially to both the luminance and chrominance components. Inter-component prediction has shown that chrominance can be locally modeled using a linear or nonlinear model that maps the luminance component to chrominance. Such a model may be formulated as follows: C=f(Y)+ε c where C is chrominance, f(·) is the function that maps luminance to chrominance, and ε c is the remaining chrominance residual. In existing inter-component tools such as CCCM, the function f(·) is obtained using adjacent reference lines and applied using the co-located luminance values as input. Typically, inter-component tools are used for dual-tree or single-tree intra-coding. In the case of single-tree intra-coding, the function f(·) is obtained after inter-prediction, including motion compensation and optionally additional tools such as WP / BCW / CIIP / LIC, such that the final chrominance prediction maps the luminance residual to the chrominance prediction, but only after the luminance residual ε y may be obtained using the intensities at the same location before applying C=f(Y+ε y )+ε c
[0179] The parameters of the model f(·) can be obtained using CCCM, CCLM, or any other component-to-component method. The model can be one-dimensional (1D) or two-dimensional (2D) and may include nonlinear terms.
[0180] According to the embodiment, the cross-component residual model (CCRM) has four main steps, which are as follows: 1) Perform motion compensated prediction (optionally including additional refinements such as WP / BCW / CIIP / LIC) on both the luminance component Y and the chrominance components Cb, Cr. 2) Using the co-located luminance and chrominance predictions, for the chrominance components Cb and Cr, the function f Cb , f Cr are obtained separately. 3) Luminance residual ε y Reconstruct. 4) Chrominance Prediction C Cb =f Cb (Y+ε y ) and C Cr =f Cr (Y+ε y ) to get the
[0181] The signaling of CCRM can be done using, for example, a flag at the TU / PU / CU level (a joint flag for Cb, Cr) or multiple flags (separate flags for Cb, Cr). Alternatively, after CCRM, a tool that provides further prediction improvement can be applied. In such a case, the modified steps are as follows: 1. For both the luminance component Y and the chrominance components Cb, Cr, additional refinements such as WP / BCW / CIIP / LIC are applied only to the luminance component Y to perform motion compensated prediction. 2. (Unchanged) Using the co-located luminance and chrominance predictions, we apply the function f Cb , f Cr are obtained separately. 3. (Unchanged) Luminance residual ε y Reconstruct. 4. (Unchanged) Chrominance Prediction C Cb =f Cb (Y+ε y ) and C Cr =f Cr (Y+ε y ) to get the 5. Apply additional improvements such as WP / BCW / CIIP / LIC to Cb and Cr predictions.
[0182] After the last step (either step 4 or 5, depending on the two cases above), the decoder calculates the chrominance residual ε Cb , ε Cr can be reconstructed and added to further improve the reconstruction. Although the above list predicts chrominance from luminance, as will be shown later, CCRM can also be configured to perform chrominance-to-chrominance or chrominance-to-luminance prediction.
[0183] The derived model f(·) should remain stable even after adding the luminance residual. Step 4a, an alternative to step 4 above, can be used to mitigate potential problems caused by large magnitudes or large variations in the luminance residual. 4a. Chrominance Prediction
[0184]
number
number
number
[0185] In some cases, it may be beneficial to perform step 4 as a weighted sum of the original chroma prediction and the CCRM prediction, for example as follows: 4b. Chrominance Prediction
[0186]
number
number
number
number
[0187] As another example, step 4 may be performed by applying a mapping function f(·) to only the luminance residuals, for example: 4c. Chrominance Prediction
[0188]
number
number
number
number
[0189] Any combination of the above steps 4, 4a, 4b, and 4c is possible, for example, where w0 and w1 are weights, a weighted sum w0*4a+w1*4b may be used. The choice between 4, 4a, 4b, and 4c can occur at the sub-block level, or even at the sample level. A further extension of the above approach is to include chrominance residuals in the model to obtain inter-chrominance predictions as follows:
[0190]
number
number
number
[0191]
number
[0192] Some example embodiments are described below.
[0193] In an embodiment, the co-located luma component may be downsampled to match the sub-sampled chroma grid.
[0194] In an embodiment, the co-located luminance components (including the residual) may be filtered using any 1D or 2D filter.
[0195] In an embodiment, the mapping function f(·) may be obtained using any method used to derive a component-component model, such as CCCM or CCLM.
[0196] In an embodiment, the mapping function f(·) may include any number of filter taps and may include non-linear terms, such as quadratic terms.
[0197] In an embodiment, the mapping function f(·) can be 1D or 2D.
[0198] In an embodiment, the mapping function f(·) can map from luminance to chrominance or from chrominance to chrominance. Any weighted sum of the two versions can be used to obtain the final prediction.
[0199] In an embodiment, the use of CCRM may be signaled from the encoder to the decoder at the level of a TU, PU, CU, CTU, slice, frame, or sequence.
[0200] In embodiments, the use of CCRM can be signaled for different components such as Cb, Cr, and Y separately or together.
[0201] In an embodiment, the luminance residual may be clipped or spatially filtered before applying the inter-component model.
[0202] In an embodiment, the type and amount of filtering applied to the luminance residual in step 4a may be determined separately for Cb and Cr.
[0203] In an embodiment, the type and amount of filtering applied to the luminance residual in step 4a may be determined based on the output of f(·).
[0204] In an embodiment, the type and amount of filtering applied to the luma residual in step 4a may be applied differently at the sub-block level or at the sample level.
[0205] In an embodiment, the mixing weights α and β in step 4b may be defined for Cb and Cr independently or jointly.
[0206] In an embodiment, the mixing weights α and β in step 4b may be defined at the sub-block level or even at the sample level.
[0207] In an embodiment, the mixing weights α and β of step 4b may be signaled from the encoder to the decoder at the level of a TU, CU, CTU, slice, frame, or sequence.
[0208] In an embodiment, the block prediction may be obtained as a weighted sum of different mapping functions, such as the weighted sums described in steps 4, 4a, and 4b.
[0209] In an embodiment, block prediction may use different mapping functions at the sub-block level.
[0210] In an embodiment, block prediction may use different mapping functions based on thresholds derived for luminance, luminance residual, or chrominance. For example, a threshold may be set (e.g., based on the average luminance over the block), and different mapping functions are derived and used below or above that threshold.
[0211] In embodiments, a mapping function f(·) may map from chrominance to luminance. In chrominance-to-luminance prediction, the input of f(·) may be upsampled to match the luminance grid. In chrominance-to-luminance prediction, the output of f(·) may be upsampled to match the luminance grid.
[0212] A method according to one embodiment is shown in FIG. 9 and includes receiving an image block unit of a frame, the image block unit including samples in a first chrominance channel, a second chrominance channel, and one luminance channel (900); defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel (902); and using the reference frame and motion vectors to define a co-located reference area across the luminance component, the first chrominance component, and the second chrominance channel. The method includes performing motion compensated prediction of two chrominance components (904), using the co-located luminance and chrominance predictions to obtain a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component (906), reconstructing a luminance residual (908), and obtaining a first chrominance prediction by using the first function and the luminance residual, and obtaining a second chrominance prediction by using the second function and the luminance residual (910).
[0213] An apparatus according to one aspect includes means for receiving an image block unit of a frame, the image block unit including samples in color channels including at least one chrominance channel and one luminance channel; means for determining an intra-prediction direction based on the samples of the luminance channel of the image block unit; means for selecting a low-frequency non-separable transform index for a low-frequency non-separable transform using the determined intra-prediction direction; and means for performing the low-frequency non-separable transform using the selected index.
[0214] According to an embodiment, the device comprises means for downsampling the samples of said luminance channel to correspond to the sample size of the chrominance samples before determining the filter coefficients.
[0215] In a further aspect, an apparatus is provided having at least one processor and at least one memory, wherein the at least one memory has stored therein code that, when executed by the at least one processor, causes the apparatus to at least perform the following steps: receive an image block unit of a frame, the image block unit including samples in color channels including at least one chrominance channel and one luminance channel; reconstruct samples of the luminance channel of the image block unit; determine an intra-prediction direction based on the samples of the luminance channel of the image block unit; select a low-frequency non-separable transform index for a low-frequency non-separable transform using the determined intra-prediction direction; and perform the low-frequency non-separable transform using the selected index.
[0216] Such an apparatus may, for example, comprise the functional units disclosed in any of Figures 1, 2, 4a and 4b for implementing the embodiments.
[0217] Such an apparatus further comprises code stored in said at least one memory, which code, when executed by said at least one processor, causes the apparatus to perform one or more of the embodiments disclosed herein.
[0218] FIG. 13 is a graphical representation of an exemplary multimedia communication system in which various embodiments may be implemented. A data source 1510 provides a source signal in analog, uncompressed, or compressed digital format, or any combination of these formats. An encoder 1520 may include or be connected to preprocessing, such as data format conversion and / or filtering of the source signal. The encoder 1520 encodes the source signal into a coded media bitstream. It should be noted that the bitstream to be decoded may be received directly or indirectly from a remote device within virtually any type of network. Additionally, the bitstream may be received from local hardware or software. An encoder 1520 may be capable of encoding two or more media types, such as audio and video, or two or more encoders 1520 may be required to encode different media types of the source signal. An encoder 1520 may also take synthetically generated input, such as graphics and text, or may be capable of generating coded bitstreams of synthetic media. In the following, for simplicity of explanation, only the processing of one coded media bitstream of one media type is considered. It should be noted, however, that a real-time broadcast service typically includes multiple streams (typically at least one audio, video, and text subtitle stream). It should also be noted that while a system may include many encoders, only one encoder 1520 is depicted in this diagram for simplicity of explanation without loss of generality. It should be further understood that while the text and examples contained herein may specifically describe an encoding process, those skilled in the art will understand that the same concepts and principles apply to the corresponding decoding process, and vice versa.
[0219] The encoded media bitstreams may be transferred to storage 1530. Storage 1530 may comprise any type of mass memory for storing the encoded media bitstreams. The format of the encoded media bitstreams in storage 1530 may be a basic self-contained bitstream format, or one or more encoded media bitstreams may be encapsulated in a container file, or the encoded media bitstreams may be encapsulated in a segment format suitable for DASH (or a similar streaming system) and stored as a sequence of segments. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown) may be used to store the one or more media bitstreams in the file and to create file format metadata that may also be stored in the file. The encoder 1520 or storage 1530 may comprise the file generator, or the file generator may be operably attached to either the encoder 1520 or storage 1530. Some systems operate “live,” i.e., omitting storage and transferring the encoded media bitstreams directly from the encoder 1520 to the sender 1540. The encoded media bitstreams may then be forwarded to a sender 1540, also referred to as a server, as needed. The format used in the transmission may be a basic self-contained bitstream format, a packet stream format, a segment format suitable for DASH (or a similar streaming system), or one or more encoded media bitstreams may be encapsulated in a container file. The encoder 1520, storage 1530, and server 1540 may reside on the same physical device or may be included in separate devices.The encoder 1520 and server 1540 may operate using live real-time content, in which case the encoded media bitstream is typically not stored persistently but is buffered for a short period of time within the content encoder 1520 and / or server 1540 to smooth out fluctuations in processing delays, transmission delays, and encoded media bitrates.
[0220] The server 1540 transmits the encoded media bitstream using a communication protocol stack. The stack may include, but is not limited to, one or more of Real-Time Transport Protocol (RTP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). If the communication protocol stack is packet-oriented, the server 1540 encapsulates the encoded media bitstream into packets. For example, if RTP is used, the server 1540 encapsulates the encoded media bitstream into RTP packets according to the RTP payload format. Typically, each media type has its own RTP payload format. It should also be noted that although the system may include more than one server 1540, for simplicity, the following description considers only one server 1540.
[0221] If the media content is encapsulated in a container file for storage 1530 or for inputting data to the transmitter 1540, the transmitter 1540 may comprise or be operably attached to a “transmission file parser” (not shown). In particular, if the container file is not transmitted as such but at least one of the contained encoded media bitstreams is encapsulated for transport via a communication protocol, the transmission file parser finds the appropriate portion of the encoded media bitstream to be conveyed via the communication protocol. The transmission file parser may also help create the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may include encapsulation instructions, such as hint tracks in ISOBMFF, for encapsulation of at least one of the contained media bitstreams over the communication protocol.
[0222] The server 1540 may or may not be connected to the gateway 1550 via a communication network, which may be, for example, a CDN, the Internet, and / or a combination of one or more access networks. The gateway may additionally or alternatively be referred to as a middlebox. In the case of DASH, the gateway may be an edge server (of a CDN) or a web proxy. Note that while a system may generally comprise any number of gateways or the like, for simplicity, the following description considers only one gateway 1550. The gateway 1550 may perform various types of functions, such as converting packet streams according to one communication protocol stack to another, merging and forking data streams, and manipulating data streams according to downlink capabilities and / or receiver capabilities, such as controlling the bit rate of the forwarded stream according to prevailing downlink network conditions. The gateway 1550 may be a server entity in various embodiments.
[0223] The system typically includes one or more receivers 1560 capable of receiving, demodulating, and decapsulating the transmitted signal into a coded media bitstream. The coded media bitstream may be transferred to a recording storage 1570. The recording storage 1570 may comprise any type of mass memory for storing the coded media bitstream. Alternatively, or additionally, the recording storage 1570 may comprise computational memory, such as random access memory. The coded media bitstreams in the recording storage 1570 may be in a basic self-contained bitstream format, or one or more coded media bitstreams may be encapsulated in a container file. When there are multiple coded media bitstreams, such as audio and video streams, that are associated with each other, a container file is typically used, and the receiver 1560 includes or is attached to a container file generator that generates a container file from the input stream. Some systems operate “live,” i.e., omitting the recording storage 1570 and transferring the coded media bitstream directly from the receiver 1560 to the decoder 1580. In some systems, only the most recent portion of the recorded stream, for example the most recent 10 minute excerpt of the recorded stream, is maintained in recording storage 1570, while previously recorded data is discarded from recording storage 1570.
[0224] The encoded media bitstreams may be transferred from the recording storage 1570 to the decoder 1580. If there are many encoded media bitstreams, such as audio and video streams, associated with each other and encapsulated in a container file, or if a single media bitstream is encapsulated in a container file, for example for ease of access, a file parser (not shown) is used to decapsulate each encoded media bitstream from the container file. The recording storage 1570 or the decoder 1580 may comprise the file parser, or the file parser may be attached to either the recording storage 1570 or the decoder 1580. It should also be noted that while a system may include many decoders, only one decoder 1570 is described herein for simplicity of explanation without loss of generality.
[0225] The encoded media bitstream is further processed by a decoder 1570, the output of which is one or more uncompressed media streams. Finally, a renderer 1590 may play the uncompressed media streams, for example using loudspeakers or a display. The receiver 1560, recording storage 1570, decoder 1570, and renderer 1590 may reside in the same physical device or may be included in separate devices.
[0226] The sender 1540 and / or the gateway 1550 may be configured to perform switching between different representations, view switching, bitrate adaptation, and / or fast start-up, e.g., for switching between different viewports of 360-degree video content, and / or the sender 1540 and / or the gateway 1550 may be configured to select a transmitted representation. Switching between different representations may be performed for multiple reasons, such as in response to a request from the receiver 1560 or due to general conditions such as the throughput of the network over which the bitstream is transmitted. In other words, the receiver 1560 may initiate the switch between representations. A request from the receiver can be, for example, a request for a segment or subsegment from a different representation than previously, a request for a change in the transmitted scalability layer and / or sublayer, or a request for a change in a rendering device with different capabilities compared to previous capabilities. A request for a segment may be an HTTP GET request. A request for a subsegment may be an HTTP GET request with a byte range. Additionally or alternatively, bitrate adjustment or bitrate adaptation may be used, for example, in streaming services, to provide so-called fast start-up, where the bitrate of the transmitted stream is lower than the channel bitrate after the start of streaming or random access, in order to start playback immediately and to achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions. Bitrate adaptation may involve multiple representation or layer up-switching operations and representation or layer down-switching operations, performed in various orders.
[0227] Decoder 1580 may be configured to perform switching between different representations, view switching, bitrate adaptation, and / or fast start-up, e.g., for switching between different viewports of 360-degree video content, and / or decoder 1580 may be configured to select a transmitted representation. Switching between different representations may be performed for multiple reasons, such as to achieve faster decoding operations or to adapt the transmitted bitstream, e.g., in terms of bitrate, to prevailing conditions such as the throughput of the network over which the bitstream is conveyed. Faster decoding operations may be required, for example, when a device including decoder 1580 is multitasking and uses computing resources for purposes other than decoding the video bitstream. In another example, faster decoding operations may be required when content is played back at a pace faster than normal playback speed (e.g., two or three times faster than conventional real-time playback speed).
[0228] In the above, some embodiments have been described with reference to and / or using HEVC and / or VVC terminology. It should be understood that the embodiments may similarly be implemented using any video encoder and / or video decoder.
[0229]
[0013] Where example embodiments are described above with reference to an encoder, it should be understood that the resulting bitstream and decoder may include corresponding elements. Similarly, where example embodiments are described with reference to a decoder, it should be understood that the encoder may include structure and / or computer program for generating a bitstream that is decoded by the decoder. For example, some embodiments have been described with reference to generating predictive blocks as part of encoding. Embodiments may similarly be realized by generating predictive blocks as part of decoding, with the difference being that coding parameters such as horizontal and vertical offsets are decoded from the bitstream rather than being determined by the encoder.
[0230] The above-described embodiments of the present invention describe the codec in terms of separate encoder and decoder devices to aid in understanding the processes involved. However, it will be understood that the devices, structures, and operations may be implemented as a single encoder / decoder device / structure / operation. Furthermore, it is possible that the coder and decoder may share some or all common elements.
[0231] While the above examples describe embodiments of the invention operating within a codec within an electronic device, it will be understood that the invention as defined in the claims may be implemented as part of any video codec. Thus, for example, embodiments of the invention may be implemented in a video codec that is capable of performing video encoding over fixed or wired communication paths.
[0232] Thus, the user equipment may be equipped with a video codec, such as the video codec described in the embodiments of the present invention above. It should be understood that the term user equipment is intended to cover any suitable type of wireless user equipment, such as a mobile phone, a portable data processing device, or a portable web browser.
[0233] Additionally, elements of a public land mobile network (PLMN) may also be equipped with video codecs such as those mentioned above.
[0234] In general, various embodiments of the present invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. While various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing device, or some combination thereof.
[0235] Embodiments of the present invention may be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that any block of logic flow as in the figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software may be stored on physical media such as memory chips, or memory blocks embodied within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc.
[0236] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), and processors based on multi-core processor architectures.
[0237] Embodiments of the present invention may be practiced in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs ready to be etched into semiconductor substrates.
[0238] Programs such as those offered by Synopsys, Inc. of Mountain View, California, and Cadence Design Systems, Inc. of San Jose, California, use established design rules and pre-stored libraries of design modules to automatically route conductors and place components on semiconductor chips. After the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) may be sent to a semiconductor manufacturing facility or "fab" for fabrication.
[0239] The foregoing description has provided a complete and informative description of example embodiments of the present invention, by way of illustrative and non-limiting examples. However, various modifications and adaptations may become apparent to those skilled in the art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention remain within the scope of this invention.
Claims
1. means for receiving an image block unit of a frame, the image block unit including samples in a first chrominance channel, a second chrominance channel, and one luminance channel; means for defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel; means for performing motion compensated prediction of the luminance component, the first chrominance component, and the second chrominance component using a reference frame and a motion vector; means for using the co-located luminance prediction and the chrominance prediction to obtain a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component; means for reconstructing the luminance residual; means for obtaining a first chrominance prediction by using the first function and the luminance residual, and for obtaining a second chrominance prediction by using the second function and the luminance residual; An apparatus comprising:
2. 2. The apparatus of claim 1, further comprising means for downsampling the co-located luminance component to match the subsampled chrominance grid.
3. 3. The apparatus of claim 1, further comprising means for filtering the co-located luminance components using a one-dimensional or two-dimensional filter, and wherein the texture analysis comprises a gradient calculation of the samples.
4. 4. The apparatus of claim 1, comprising means for using either the first mapping function or the second mapping function to map the luminance component to the first chrominance component or to the second chrominance component, or to map the first chrominance component to the second chrominance component.
5. 5. The apparatus of claim 1, further comprising means for signaling the use of an inter-component residual model from an encoder to a decoder at the level of a transform unit, a prediction unit, a coding unit, a coding tree unit, a slice, a frame, or a sequence.
6. An apparatus according to any preceding claim, comprising means for clipping or spatially filtering the luminance residual before applying the inter-component model.
7. 7. The apparatus of claim 1, further comprising means for determining the type and amount of filtering to be applied to the luminance residual separately for the first chrominance component and the second chrominance component.
8. 8. The apparatus of claim 1, comprising means for determining the type and amount of filtering to apply to the luminance residual based on an output of the first mapping function or an output of the second mapping function.
9. The device according to any one of claims 1 to 8, comprising means for applying the type and amount of filtering to the luminance residual differently at a sub-block level or at a sample level.
10. 10. The apparatus of claim 1, comprising means for using different mapping functions in the block prediction based on one or more thresholds derived for luminance, luminance residual, or chrominance.
11. Apparatus according to any one of claims 1 to 10, comprising means for mapping chrominance components onto luminance components.
12. means for upsampling the input of said mapping function to match a luminance grid; means for upsampling the output of said mapping function to match a luminance grid; The device according to any one of claims 1 to 11, comprising at least one of:
13. Apparatus according to any preceding claim, comprising means for mapping the first chrominance component onto the second chrominance component.
14. 1. An apparatus for encoding comprising at least one processor and a memory containing computer program code, the memory and the computer program code operating in conjunction with the at least one processor causing the apparatus to: receiving an image block unit of a frame, the image block unit including samples in a first chrominance channel, a second chrominance channel, and one luminance channel; defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel; performing motion compensated prediction of the luminance component, the first chrominance component, and the second chrominance component using a reference frame and a motion vector; using the co-located luminance prediction and the chrominance prediction to obtain a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component; and reconstructing a luminance residual; obtaining a first chrominance prediction by using the first function and the luminance residual, and obtaining a second chrominance prediction by using the second function and the luminance residual; An apparatus configured to cause at least
15. receiving an image block unit of a frame, the image block unit including samples in a first chrominance channel, a second chrominance channel, and one luminance channel; defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel; performing motion compensated prediction of the luminance component, the first chrominance component, and the second chrominance component using a reference frame and a motion vector; using the co-located luminance prediction and the chrominance prediction to obtain a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component; and reconstructing a luminance residual; obtaining a first chrominance prediction by using the first function and the luminance residual, and obtaining a second chrominance prediction by using the second function and the luminance residual; A method comprising:
Citation Information
Patent Citations
Encoding device, decoding device, and program
JP2011193363A
Inter-component Prediction in Video Coding
JP2017523672A
Cross-component prediction in video coding
US20150373349A1
Image processing device and image processing method
WO2021054437A1
Cited By
Encoding / Decoding Method, Codestream, Encoder, Decoder, and Storage Medium
JP2025533266A
Simplifying the derivation of filter coefficients for inter-component prediction
JP2025534504A
Systems and methods for applying non-separable transforms on inter prediction residuals
US12621489B2