UM APARELHO, UM MÉTODO E UM PROGRAMA DE COMPUTADOR PARA CODIFICAÇÃO E DECODIFICAÇÃO DE VÍDEO
Patent Information
- Application Number
- BR112025019737
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-16
- Filing Date
- 2024-02-06
- Publication Date
- 2026-08-04
Smart Images

Figure 00000084_0000 
Figure 00000084_0001 
Figure 00000085_0000
Abstract
Description
1 / 79 A device, a method, and a computer program for encoding and decoding video. TECHNICAL FIELD
[0001] The present invention relates to an apparatus, a method and a computer program for encoding and decoding video. BACKGROUND
[0002] In video encoding, video and image samples are typically encoded using color representations such as YUV or YCbCr, which consist of one luminance (luma) channel and two chrominance (chroma) channels. In these cases, the luminance channel, primarily representing the scene's illumination, is typically encoded at a certain resolution, while the chrominance channels, typically representing differences between certain color components, are often encoded at a second, lower resolution than the luminance signal. The intention of this type of differential representation is to decorrelate the color components and be able to compress the data more efficiently.
[0003] The Linear Between Components Model (CCLM) and the Convolutional Between Components Model (CCCM) are used to predict samples in chroma channels (e.g., Cb and Cr) through cross-channel correlation (e.g., using luminance samples). Model parameters are derived based on samples reconstructed in the vicinity of the chroma block, neighboring samples colocalized in the luma block, as well as samples reconstructed within the colocalized luma block.
[0004] The assumption is that the reconstructed neighboring samples have a good correlation with the samples within the prediction block. However, if the neighboring samples have no correlation with the samples within the current block, or if the correlation is very small, then the calculated model may not be able to predict the samples within the block efficiently. SUMMARY Petition 870250083300, dated 09 / 16 / 2025, page 11 / 298 2 / 79
[0005] Now, to at least mitigate the above problems, an improved method for enhancing the efficiency of cross-component prediction tools is introduced in this document.
[0006] The scope of protection sought for various embodiments of the invention is established by the independent claims. The embodiments and features, if any, described in this descriptive report that are not covered by the scope of the independent claims should be interpreted as useful examples for understanding various embodiments of the invention.
[0007] A method according to a first aspect comprises receiving an image block unit from a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; reconstructing samples of said luminance channel from the image block unit; determining a reference area for predicting target samples of at least one color channel from the image block unit, wherein said reference area comprises one or more reference samples in an actual block or colocalized in an actual color channel or in an actual frame or in a reference color channel or in a reference frame encoded using an intrablock copy (IBC) method; determining a block vector indicating a spatial distance from a target sample area to the reference area;and predict said target samples of at least one color channel of the image block unit using a cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
[0008] An apparatus according to a second aspect comprises means for receiving an image block unit from a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; means for reconstructing samples of said luminance channel from the image block unit; means for determining a reference area. Petition 870250083300, dated 09 / 16 / 2025, page 12 / 298 3 / 79 for predicting target samples of at least one color channel of the image block unit, wherein said reference area comprises one or more reference samples in a current block or colocalized in a current color channel or in a current frame or in a reference color channel or in a reference frame encoded using an intrablock copy (IBC) method; means for determining a block vector indicating a spatial distance from a target sample area to the reference area; and means for predicting said target samples of at least one color channel of the image block unit using a cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
[0009] Depending on one embodiment, the cross-component prediction method is either a cross-component linear model (CCLM) or a cross-component convolutional model (CCCM).
[0010] According to one embodiment, the apparatus comprises means for deriving the parameters for the cross-component model using samples from a reference block in the current channel and a reference block in the reference channel.
[0011] According to one embodiment, the apparatus comprises means for inferring the block vector of the colocalized block in the reference channel; and means for scaling the block vector to correspond to a sampling density of the current channel.
[0012] According to one embodiment, the block vector inferred from the reference channel is at least partially misaligned with colocalized block coordinates.
[0013] According to one embodiment, the apparatus comprises means for identifying the nearest match to the reference area in the current channel using a template match-based search in the current channel.
[0014] According to one embodiment, the apparatus comprises means for calculating the cross-component model parameters and applying the model. Petition 870250083300, dated 09 / 16 / 2025, page 13 / 298 4 / 79 calculated for prediction based on one or more of the following: - one or more samples of the reference block in the current channel; - one or more samples of the reference block in the reference channel; - one or more samples of the block colocalized in the reference channel; - one or more neighborhood samples of the reference block in the current channel; - one or more neighborhood samples of the reference block in the reference channel; - one or more neighborhood samples of the current block in the current channel; - one or more neighborhood samples of the colocalized block in the reference channel.
[0015] According to one embodiment, the apparatus comprises means for calculating cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - Horizontal and / or vertical coordinates of said samples; - Directional gradient values of said samples.
[0016] According to one embodiment, the apparatus comprises means for using the block vector of the colocalized block inferred from the reference channel as an initial block vector for the current block.
[0017] According to one embodiment, the apparatus comprises means for generating a list of candidates for a block vector of encoded blocks using the intrablock copy (IBC) method in the reference channel in the colocalized block and / or in the neighborhood of the current block in the current channel.
[0018] According to one embodiment, the apparatus comprises means for obtaining the final cross-component prediction parameters of the current block using a two-stage derivation, wherein a first-stage derivation is performed using the reference block in the reference channel pointed to by the block vector and a second-stage derivation is performed in the vicinity of the current block.
[0019] As a third aspect, an apparatus is provided comprising: Petition 870250083300, dated 09 / 16 / 2025, page 14 / 298 5 / 79 at least one processor and at least one memory, said at least one memory stored with code therein which, when executed by said at least one processor, causes the apparatus to perform at least: receive an image block unit from a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; reconstruct samples of said luminance channel from the image block unit; determine a reference area for predicting target samples of at least one color channel from the image block unit, wherein said reference area comprises one or more reference samples in a current block or colocalized in a current color channel or in a current frame or in a reference color channel or in a reference frame encoded using an intrablock copy (IBC) method;Determine a block vector indicating a spatial distance from a target sample area to the reference area; and predict said target samples from at least one color channel of the image block unit using a cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
[0020] The devices and computer-readable storage media stored with code on them, as described above, are thus arranged to perform the above methods and one or more of the modalities related thereto. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To better understand the present invention, reference will now be made by way of example to the accompanying drawings, in which:
[0022] Figure 1 schematically shows an electronic device employing embodiments of the invention;
[0023] Figure 2 schematically shows a user-friendly device suitable for employing embodiments of the invention;
[0024] Figure 3 also schematically shows electronic devices that Petition 870250083300, dated 09 / 16 / 2025, page 15 / 298 6 / 79 employ embodiments of the invention connected with the use of wired or wireless network connections;
[0025] Figures 4a and 4b schematically show an encoder and a decoder suitable for implementing embodiments of the invention;
[0026] Figure 5 illustrates the locations of the samples used for deriving parameters for a Cross-Component Linear Model (CCLM);
[0027] Figures 6a and 6b show examples of classifying luma samples into two classes in the sample domain and the spatial domain, respectively;
[0028] Figure 7 shows an example of colocalized reference sample areas consisting of reconstructed luma and chroma samples defined for luma and chroma for a Cross-Component Convolutional Model (CCCM);
[0029] Figure 8 shows several examples of the filter kernel dimensions in CCCM;
[0030] Figure 9 illustrates an example of four reference lines adjacent to a prediction block;
[0031] Figure 10 illustrates an intraweighted matrix prediction process;
[0032] Figure 11 shows an example of a low-frequency non-separable transform (LFNST) process;
[0033] Figure 12 shows an example of intramold matching prediction (Intra TMP);
[0034] Figures 13a and 13b show examples of defining reference areas for intrablock copy mode (IBC);
[0035] Figures 14a and 14b illustrate the horizontal inversion and vertical inversion methods used in Reconstruction-Reordered IBC;
[0036] Figure 15 shows a flowchart of a method for predicting samples in at least one color channel according to an embodiment of the invention;
[0037] Figure 16 shows the basic principle of the current block and the block of Petition 870250083300, dated 09 / 16 / 2025, page 16 / 298 7 / 79 reference using a block vector in intrablock copy (IBC) mode;
[0038] Figures 17a - 17f show examples of using an IBC-encoded block vector to indicate the reference area for parameter derivation for cross-component prediction according to various embodiments of the invention; and
[0039] Figure 18 shows a schematic diagram of an example of a multimedia communication system within which various modalities can be implemented. DETAILED DESCRIPTION OF SOME EXEMPLARY MODALITIES
[0040] The following describes in further detail suitable apparatus and possible mechanisms for improving the efficiency of cross-component prediction tools. In this regard, reference is first made to Figures 1 and 2, wherein Figure 1 shows a block diagram of a video coding system according to an exemplary embodiment as a schematic block diagram of an exemplary electronic apparatus or device 50, which may incorporate a codec according to an embodiment of the invention. Figure 2 shows an arrangement of an apparatus according to an exemplary embodiment. The elements of Figs. 1 and 2 will be explained below.
[0041] The electronic device 50 may, for example, be a mobile terminal or user equipment of a wireless communication system. However, it would be appreciated that embodiments of the invention may be implemented within any electronic device or apparatus that may require encoding and decoding or encoding or decoding of video images.
[0042] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 may further comprise a display 32 in the form of a liquid crystal display. In other embodiments of the invention, the display may be any display technology suitable for displaying an image or video. The apparatus 50 may further comprise a Petition 870250083300, dated 09 / 16 / 2025, page 17 / 298 8 / 79 numeric keypad 34. In other embodiments of the invention, any suitable data mechanism or user interface may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0043] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analog signal input. The apparatus 50 may further comprise an audio output device which, in embodiments of the invention, may be any of: a headphone 38, a loudspeaker or an analog audio output connection or digital audio. The apparatus 50 may also comprise a battery (or, in other embodiments of the invention, the device may be powered by any suitable mobile energy device such as a solar cell, fuel cell or mechanical generator). The apparatus may further comprise a camera capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short-range line-of-sight communication with other devices.In other embodiments, the 50 device may additionally include any suitable short-range communication solution such as, for example, a wireless connection via Bluetooth or a wireless connection via USB / firewire.
[0044] The apparatus 50 may comprise a controller 56, processor or set of processor circuits for controlling the apparatus 50. The controller 56 may be connected to memory 58 which, in embodiments of the invention, may store both image data and audio data and / or may also store instructions for implementation in the controller 56. The controller 56 may additionally be connected to a set of codec circuits 54 suitable for performing the encoding and decoding of audio and / or video data or assisting in the encoding and decoding performed by the controller.
[0045] The device 50 may additionally comprise a card reader 48 and a smart card 46, for example, a UICC and UICC reader for Petition 870250083300, dated 09 / 16 / 2025, page 18 / 298 9 / 79 provides user information and is suitable for providing authentication information for user authentication and authorization on a network.
[0046] The device 50 may comprise a set of radio interface circuits 52 connected to the controller and suitable for generating wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system or a wireless local area network. The device 50 may further comprise an antenna 44 connected to the set of radio interface circuits 52 for transmitting radio frequency signals generated in the set of radio interface circuits 52 to other device(s) and for receiving radio frequency signals from other device(s).
[0047] The device 50 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing. The device may receive video image data for processing from another device prior to transmission and / or storage. The device 50 may also receive the image wirelessly or via a wired connection for encoding / decoding. The structural elements of device 50 described above represent examples of means for performing a corresponding function.
[0048] With regard to Figure 3, an example of a system within which embodiments of the present invention can be used is shown. System 10 comprises multiple communication devices that can communicate via one or more networks. System 10 may comprise any combination of wired or wireless networks including, but not limited to, a wireless cellular telephone network (such as a GSM, UMTS, CDMA, etc. network), a wireless local area network (WLAN) as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network and the Internet.
[0049] System 10 may include both communication and / or devices Petition 870250083300, dated 09 / 16 / 2025, page 19 / 298 10 / 79 device 50 suitable wired as well as wireless to implement embodiments of the invention.
[0050] For example, the system shown in Figure 3 shows a mobile phone network 11 and a representation of the internet 28. Internet connectivity 28 may include, but is not limited to, long-range wireless connections, short-range wireless connections and various wired connections including, but not limited to, telephone lines, power lines, power lines and similar communication pathways.
[0051] The exemplary communication devices shown in the system 10 may include, but are not limited to, an electronic device or appliance 50, a combination of a personal digital assistant (PDA) and a mobile phone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22. The device 50 may be stationary or mobile when carried by an individual who is in motion. The device 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a vessel, an aircraft, a bicycle, a motorcycle or any similar suitable mode of transport.
[0052] The modalities can also be implemented in a set-top box; that is, a digital TV receiver, which may or may not have a display or wireless capabilities, in tablets or personal computers (laptops) (PCs), which have hardware or software or a combination of encoder / decoder implementations, in various operating systems, and in sets of chips, processors, DSPs and / or embedded systems that offer hardware / software based encoding.
[0053] Some device may send and receive calls and messages and communicate with service providers via a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that enables communication between the mobile phone network 11 and the internet 28. The system may include communication devices Petition 870250083300, dated 09 / 16 / 2025, page 20 / 298 11 / 79 additional and communication devices of various types.
[0054] Communication devices can communicate using various transmission technologies including, but not limited to, Code Division Multiple Access (CDMA), Global Systems for Mobile Communications (GSM), Universal Mobile Telecommunication System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol-Internet Protocol (TCP-IP), Short Message Service (SMS), Multimedia Message Service (MMS), email, Instant Message Service (IMS), Bluetooth, IEEE 802.11, and any similar wireless communication technology. A communication device involved in implementing various embodiments of the present invention can communicate using various means including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0055] In telecommunications and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium, such as a wire, while a logical channel can refer to a logical connection through a multiplexed medium capable of transmitting multiple logical channels. A channel can be used to transmit an information signal, for example, a bit stream, from one or more senders (or transmitters) to one or more receivers.
[0056] An MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format for transporting audio, video, and other media as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify an elementary stream (also known as a packed elementary stream) within the TS. Therefore, a logical channel within an MPEG-2 TS can be considered to correspond to a specific PID value.
[0057] Available media file format standards include ISO-based media file format (ISO / IEC 14496-12, which may be Petition 870250083300, dated 09 / 16 / 2025, page 21 / 298 12 / 79 (abbreviated as ISOBMFF) and a file format for video structured in NAL units (ISO / IEC 14496-15), which is derived from ISOBMFF.
[0058] A video codec consists of an encoder that transforms the input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. A video encoder and / or a video decoder can also be separate from each other, i.e., they do not need to form a codec. Typically, the encoder discards some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bitrate).
[0059] Typical hybrid video encoders, for example, many implementations of ITU-T H.263 and H.264 encoders, encode video information in two stages. First, the pixel values in a given image area (or “block”) are predicted, for example, by means of motion compensation (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by means of spatial compensation (using the pixel values around the block to be encoded in a specified manner). Second, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded.This is typically done by transforming the difference into pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (image quality) and the size of the resulting encoded video representation (file size or transmission bitrate).
[0060] In time forecasting, the prediction sources are previously decoded images (also known as reference images). Petition 870250083300, dated 09 / 16 / 2025, page 22 / 298 13 / 79 In intrablock copy (IBC; also known as intrablock copy prediction), prediction is applied similarly to temporal prediction, but the reference image is the current image, and only previously decoded samples can be referenced in the prediction process. Interlayer or interview prediction can be applied similarly to temporal prediction, but the reference image is a decoded image from another scalable layer or another view, respectively. In some cases, inter-prediction may refer only to temporal prediction, although in other cases, inter-prediction may collectively refer to temporal prediction and any intrablock copy, interlayer prediction, and interview prediction, provided they are performed with a process equal to or similar to temporal prediction. Inter-prediction or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0061] Motion compensation can be performed with full-sample or subsample precision. In the case of full-sample precision motion compensation, motion can be represented as a motion vector with integer values for horizontal and vertical displacements, and the motion compensation process effectively copies samples from the reference image using these displacements. In the case of subsample precision motion compensation, motion vectors are represented by fractional or decimal values for the horizontal and vertical components of the motion vector. If a motion vector refers to a non-integer number position in the reference image, a subsample interpolation process is typically required to calculate predicted sample values based on the reference samples and the selected subsample position.The subsample interpolation process typically consists of horizontal filtering that compensates for horizontal deviations from the full sample positions, followed by vertical filtering that compensates for vertical deviations from the full sample positions. Petition 870250083300, dated 09 / 16 / 2025, page 23 / 298 14 / 79 full sample positions. However, vertical processing can also be done before horizontal processing in some environments.
[0062] Inter-prediction, which may also be called temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter-prediction, the prediction sources are previously decoded images. Intra-prediction utilizes the fact that adjacent pixels within the same image are likely to be correlated. Intra-prediction can be performed in the transform or spatial domain, i.e., sample values or transform coefficients can be predicted. Intra-prediction is typically exploited in intra-coding, where no inter-prediction is applied.
[0063] One result of the encoding procedure is a set of encoding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-encoded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector can be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor can be encoded. Encoding parameter prediction and intra-prediction can be collectively referred to as image prediction.
[0064] Figures 4a and 4b show an encoder and a decoder suitable for employing embodiments of the invention. A video codec consists of an encoder that transforms an input video into a compressed representation suitable for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically, the encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bitrate). An example of an encoding process is illustrated in Figure 4a. Figure 4a illustrates an image to be encoded (ln); a predicted representation of an image block. Petition 870250083300, dated 09 / 16 / 2025, page 24 / 298 15 / 79 (P'n); a prediction error signal (Dn); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (l'n); a final reconstructed image (R'n); a transform (T) and an inverse transform (T-1); a quantization (Q) and inverse quantization (Q-1); entropy coding (E); a frame of reference memory (RFM); inter-prediction (Pinter); intra-prediction (Pintra); mode selection (MS) and filtering (F).
[0065] An example of a decoding process is illustrated in Figure 4b. Figure 4b illustrates a predicted representation of an image block (P'n); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (l'n); a final reconstructed image (R'n); an inverse transform (T-1); an inverse quantization (Q-1); an entropy decoding (E-1); a frame of reference memory (RFM); a prediction (inter or intra) (P); and filtering (F).
[0066] Many hybrid video encoders encode video information in two phases. First, the pixel values in a given image area (or “block”) are predicted, for example, by means of motion compensation (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by spatial means (using the pixel values around the block to be encoded in a specified manner).Secondly, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded. This is typically done by transforming the difference into pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding of the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (image quality) and the size of the resulting encoded video representation (file size or transmission bitrate). Video codecs may also provide a transform-hopping mode, which encoders can choose to use. In transform-hopping mode, the error... Petition 870250083300, dated 09 / 16 / 2025, page 25 / 298 16 / 79 of prediction is encoded in a sample domain, for example, by deriving a sample difference value relative to certain adjacent samples and encoding the sample difference value with an entropy encoder.
[0067] Entropy encoding / decoding can be performed in many ways. For example, context-based encoding / decoding can be applied, where both the encoder and decoder modify the context state of an encoding parameter based on previously encoded / decoded encoding parameters. Context-based encoding can, for example, be context-adaptive binary arithmetic coding (CABAC) or context-based variable-length coding (CAVLC) or any similar entropy encoding. Entropy encoding / decoding can alternatively or additionally be performed using a variable-length encoding scheme, such as Huffman encoding / decoding or ExpGolomb encoding / decoding. Decoding encoding parameters from an entropy-encoded bitstream or codeword can be referred to as parsing.
[0068] The expression along the bitstream (e.g., indicating along the bitstream) can be defined to refer to out-of-band transmission, signaling, or storage in a way that out-of-band data is associated with the bitstream. The expression decoding along the bitstream or similar can refer to the decoding of said out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream. For example, an indication along the bitstream can refer to metadata in a container file that encapsulates the bitstream.
[0069] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Expert Group (VCEG) of the Telecommunications Standardization Sector of the International Telecommunication Union (ITU-T) and the Moving Picture Expert Group (MPEG) of Petition 870250083300, dated 09 / 16 / 2025, page 26 / 298 17 / 79 International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations and is referred to as ITUT Recommendation H.264 and International Standard ISO / IEC 14496-10, also known as Advanced Video Coding (AVC) MPEG-4 Part 10. There have been several versions of the H.264 / AVC standard, which integrate new extensions or features into the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0070] Version 1 of the High Efficiency Video Coding (H.265 / HEVC, also known as HEVC) standard was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by both parent standardization organizations and is referred to as ITU-T Recommendation H.265 and International Standard ISO / IEC 23008-2, also known as MPEG-H High Efficiency Video Coding Part 2 (HEVC). Later versions of H.265 / HEVC included scalable, multiview, fidelity range, three-dimensional, and screen content coding extensions, which can be abbreviated as SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.
[0071] Versatile Video Coding (VVC) (MPEG-I Part 3), also known as ITU-T H.266, is a video compression standard developed by the Joint Video Expert Team (JVET) of the Moving Picture Expert Group (MPEG), (formally ISO / IEC JTC1 SC29 WG11) and the Video Coding Expert Group (VCEG) of the International Telecommunication Union (ITU) to be the successor to HEVC / H.265.
[0072] Some key definitions, bitstream and encoding structures, and concepts of H.264 / AVC and HEVC are described in this section as an example of a video encoder, decoder, encoding method, decoding method, and a bitstream structure, in which the modalities can be implemented. Some of the key definitions, bitstream and structures Petition 870250083300, dated 09 / 16 / 2025, page 27 / 298 The 18 / 79 encoding and concepts of H.264 / AVC are the same in HEVC – therefore, they are described below together. The aspects of the invention are not limited to H.264 / AVC or HEVC, but instead, the description is given for a possible basis on which the invention may be partially or completely implemented.
[0073] Similar to many previous video coding standards, the syntax and semantics of bitstreams, as well as the decoding process for error-free bitstreams, are specified in H.264 / AVC and HEVC. The encoding process is not specified, but encoders must generate conforming bitstreams. Bitstream and decoder conformance can be verified with the Hypothetical Reference Decoder (HRD). The standards contain encoding tools that aid in copying transmission errors and losses, but the use of these tools in encoding is optional, and no decoding process has been specified for erroneous bitstreams.
[0074] The elementary unit for input to an H.264 / AVC and HVEC encoder and for output to an H.264 / AVC or HEVC decoder, respectively, is an image.An image given as input to an encoder can also be referred to as a source image, and an image decoded by a decoder can be referred to as a decoded image.
[0075] The source and decoded images are each comprised of one or more sample arrangements, as one of the following sets of sample arrangements: - Luma (Y) only (monochrome). - Luma and two chromas (YCbCr or YCgCo). Green, Blue, and Red (GBR, also known as RGB). - Arrangements that represent other unspecified monochromatic color samplings or tristimuli (e.g., YZX, also known as XYZ).
[0076] In H. 264 / AVC and HEVC, an image can be a frame or a Petition 870250083300, dated 09 / 16 / 2025, p. 28 / 298 19 / 79 field. A frame comprises an array of luma samples and possibly the corresponding chroma samples. A field is a set of alternating sample lines from a frame, and can be used as encoder input when the source signal is crosstalked. Chroma sample arrays may be absent (and therefore monochrome sampling may be in use) or chroma sample arrays may be undersampled when compared to luma sample arrays. Chroma formats can be summarized as follows: In monochromatic sampling, there is only one sample array, which can be nominally considered the luma array. - In 4:2:0 sampling, each of the two chroma arrays is half the height and half the width of the luma array. - In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array. - In 4:4:4 sampling, when no separate color plane is in use, each of the two chroma arrays has the same height and width as the luma array.
[0077] In H.264 / AVC and HEVC, it is possible to encode sample arrays as separate color planes in the bitstream and respectively decode separately encoded color planes from the bitstream. When separate color planes are in use, each of them is separately processed (by the encoder and / or the decoder) as a monochrome sampled image.
[0078] A partition can be defined as a division of a set into subsets such that each element of the set is in exactly one of the subsets.
[0079] When describing the HEVC encoding and / or decoding operation, the following terms may be used. A coding block may be defined as a block of NxN samples for some value of N such that the division of a coding tree block into coding blocks Petition 870250083300, dated 09 / 16 / 2025, page 29 / 298 20 / 79 is a partition. A coding tree block (CTB) can be defined as an NxN sample block for some value of N such that dividing a component into coding tree blocks is a partition. A coding tree unit (CTU) can be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples from an image that has three sample arrangements, or a coding tree block of samples from a monochrome image or an image that is encoded using three separate color planes and syntax structures used to encode the samples.A coding unit (CU) can be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples from an image that has three sample arrangements, or a coding block of samples from a monochrome image, or an image that is coded using three separate color planes and syntax structures used to encode the samples. A CU with the maximum allowed size can be called an LCU (last coding unit) or coding tree unit (CTU), and the video image is divided into non-overlapping LCUs.
[0080] A CU consists of one or more prediction units (PUs) that define the prediction process for the samples within the CU and one or more transform units (TUs) that define the prediction error coding process for the samples in said CU. Typically, a CU consists of a square block of samples with a size selectable from a predefined set of possible CU sizes. Each PU and TU can be further divided into smaller PUs and TUs in order to increase the granularity of the prediction and prediction error coding processes, respectively. Each PU has associated prediction information that defines what type of prediction should be applied to the pixels within that PU (e.g., motion vector information for PUs subjected to inter-prediction and intra-prediction directionality information for PUs subjected to intra-prediction). Petition 870250083300, dated 09 / 16 / 2025, page 30 / 298 21 / 79
[0081] Each TU can be associated with information describing the prediction error decoding process for the samples within that TU (including, for example, DCT coefficient information). It is typically signaled at the CU level whether prediction error coding is applied or not for each CU. If there is no prediction error residue associated with the CU, it can be considered that there are no TUs for that CU. The division of the image into CUs, and the division of CUs into PUs and TUs is typically signaled in the bitstream that allows the decoder to reproduce the intended structure of these units.
[0082] In HEVC, an image can be partitioned into slices, which are rectangular and contain an integer number of LCUs. In HEVC, the partitioning into slices forms a regular grid, where slice heights and widths differ from each other by at most one LCU. In HEVC, a slice is defined as an integer number of encoding tree units contained within an independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit. In HEVC, a slice segment is defined as an integer number of encoding tree units ordered consecutively in the slice scan and contained within a single NAL unit. Dividing each image into slice segments is a partition.In HEVC, an independent slice segment is defined as a slice segment for which the values of the slice segment header syntax elements are not inferred from the values for a preceding slice segment, and a dependent slice segment is defined as a slice segment for which the values of some slice segment header syntax elements are inferred from the values for the preceding independent slice segment in order of decoding. In HEVC, a slice header is defined as the slice segment header of the independent slice segment that is either the current slice segment or the slice segment. Petition 870250083300, dated 09 / 16 / 2025, page 31 / 298 22 / 79 independent slicing units that precede a current dependent slice segment, and a slice segment header is defined as being a portion of an encoded slice segment that contains the data elements belonging to the first or all of the encoding tree units represented in the slice segment. CUs are scanned in the scan order of LCU traces within clippings or within an image if clippings are not in use. Within an LCU, CUs have a specific scan order.
[0083] The decoder reconstructs the output video by applying a prediction method similar to the encoder to form a predicted representation of the pixel blocks (using the spatial or motion information created by the encoder and stored in the compressed representation) and prediction error decoding (the inverse operation of prediction error encoding, which recovers the quantized prediction error signal in the spatial pixel domain). After applying the prediction method and prediction error decoding, the decoder sums the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) may also apply additional filtering methods to improve the quality of the output video before displaying it and / or storing it as a prediction reference for subsequent frames in the video sequence.
[0084] Filtering may, for example, include one or more of the following: unblocking, adaptive sample bias (ASB), and / or adaptive loop filtering (ALF). H.264 / AVC includes unblocking, while HEVC includes both unblocking and ASB.
[0085] In typical video codecs, motion information is indicated with motion vectors associated with each image block subjected to motion compensation, as a prediction unit. Each of these motion vectors represents the displacement of the image block in the image to be encoded (on the encoder side) or decoded (on the decoder side). Petition 870250083300, dated 09 / 16 / 2025, page 32 / 298 23 / 79 decoder side) and the prediction source block in one of the previously encoded or decoded images. In order to represent motion vectors efficiently, these are typically differentially encoded with respect to specific block-predicted motion vectors. In typical video codecs, predicted motion vectors are created in a predefined way, for example, by calculating the median of the encoded or decoded motion vectors of adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference images and flag the chosen candidate as the motion vector predictor.In addition to predicting motion vector values, it can be predicted that the reference image(s) is / are used for motion-compensated prediction, and this prediction information can be represented, for example, by a previously encoded / decoded image reference index. The reference index is typically predicted from adjacent blocks and / or co-located blocks in the temporal reference image. Furthermore, typical high-efficiency video codecs employ an additional motion information encoding / decoding mechanism, often called blending / blend mode, in which all motion field information, including motion vector and corresponding reference image index for each available reference image list, is predicted and used without any modification / correction.Similarly, the prediction of motion field information is performed using motion field information from adjacent blocks and / or colocalized blocks in temporal reference images, and the motion field information used is selected from a list of motion field candidates populated with available motion field information from adjacent / colocated blocks.
[0086] In typical video codecs, the prediction residue after motion compensation is first transformed with a kernel of Petition 870250083300, dated 09 / 16 / 2025, page 33 / 298 24 / 79 is transformed (as DCT) and then encoded. The reason for this is that there is often still some correlation between the residuals, and the transformation can, in many cases, help to reduce this correlation and provide more efficient encoding.
[0087] Video encoding standards and specifications may allow encoders to divide an encoded image into encoded slices or similar. In-image prediction is typically disabled through slice boundaries. Thus, slices can be considered a way of dividing an encoded image into independently decodable pieces. In H.264 / AVC and HEVC, in-image prediction can be disabled through slice boundaries. Thus, slices can be considered a way of dividing an encoded image into independently decodable pieces, and slices are therefore frequently considered elementary units for transmission. In many cases, encoders can indicate in the bitstream which types of in-image prediction are turned off through slice boundaries, and the decoder operation considers this information, for example, when determining which prediction sources are available.For example, samples from a neighboring CU may be considered unavailable for intra-surface prediction if the neighboring CU resides in a different slice.
[0088] An elementary unit for the output of an H.264 / AVC or HEVC encoder and the input of an H.264 / AVC or HEVC decoder, respectively, is a Network Abstraction Layer (NAL) unit. For transport over networks or packet-oriented storage in structured files, NAL units can be encapsulated in packets or similar structures. A byte stream format was specified in H.264 / AVC and HEVC for transmission or storage environments that do not provide frame structures. The byte stream format separates NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders implement a byte-oriented start code emulation prevention algorithm, which Petition 870250083300, dated 09 / 16 / 2025, p. 34 / 298 25 / 79 adds an emulation prevention byte to the NAL unit payload if an initial code has otherwise occurred. In order to allow direct gateway operation between packet-oriented and flow-oriented systems, initial code emulation prevention can always be performed regardless of whether the byte stream format is in use or not. A NAL unit can be defined as a syntax structure containing an indication of the data type to follow and bytes containing that data in the form of an RBSP interleaved as needed with emulation prevention bytes. A raw byte sequence (RBSP) payload can be defined as a syntax structure containing an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or in the form of a column of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0089] NAL units consist of a header and payload. In H.264 / AVC and HEVC, the NAL unit header indicates the NAL unit type.
[0090] In HEVC, a two-byte NAL unit header is used for all specified NAL unit types. The NAL unit header contains a reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plus1 indication for the temporal level (may be required to be greater than or equal to 1), and a six-bit nuh_layer_id syntax element. The temporal_id_plus1 syntax element can be considered a temporal identifier for the NAL unit, and a zero-based Temporalld variable can be derived as follows: Temporalld = temporal_id_plus1 - 1. The abbreviation TID can be used interchangeably with the Temporalld variable. Temporalld equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plus1 needs to be non-zero in order to avoid initial code emulation involving the two bytes of the NAL unit header.The bitstream created by excluding all VCL NAL units that have a Temporalld greater than or equal to a value. Petition 870250083300, dated 09 / 16 / 2025, p. 35 / 298 26 / 79 selected and including all other VCL NAL units remains compliant. Consequently, an image that has Temporalld equal to tid_value does not use any image that has a Temporalld greater than tid_value as an inter-prediction reference. A sublayer or temporal sublayer can be defined as being a temporal scalable layer (or temporal layer, TL) of a temporal scalable bitstream, consisting of VCL NAL units with a particular value of the Temporalld variable and the associated non-VCL NAL units. nuh_layer_id can be understood as a scalability layer identifier.
[0091] NAL units can be categorized into Video Coding Layer (VCL) non-VCL NAL units and NAL units. VCL NAL units are typically encoded slice NAL units. In HEVC, VCL NAL units contain syntax elements that represent one or more CUs.
[0092] A non-VCL NAL unit can be, for example, one of the following types: a sequence parameter set, an image parameter set, a supplementary accentuation information (SEI) NAL unit, an access unit delimiter, a sequence NAL unit end, a bitstream NAL unit end, or a filler data NAL unit. Parameter sets may be required for the reconstruction of decoded images, while many of the other non-VCL NAL units are not required for the reconstruction of decoded sample values.
[0093] Parameters that remain unchanged throughout an encoded video sequence can be included in a sequence parameter set. In addition to parameters that may be required by the decoding process, the sequence parameter set may optionally contain video usability information (VUI), which includes parameters that may be important for buffering, image output timing, rendering, and resource reservation. In Petition 870250083300, dated 09 / 16 / 2025, page 36 / 298 27 / 79 In HEVC, a sequence parameter set RBSP includes parameters that can be referenced by one or more image parameter set RBSPs or one or more SEI NAL units containing a buffering period SEI message. An image parameter set contains such parameters that are likely to remain unchanged across multiple encoded images. An image parameter set RBSP may include parameters that can be referenced by the encoded slice NAL units of one or more encoded images.
[0094] In HEVC, a video parameter set (VPS) can be defined as a syntax structure containing syntax elements that apply to zero or more entire encoded video sequences as determined by the content of a syntax element found in the SPS referenced by a syntax element found in the PPS referenced by a syntax element found in each slice segment header.
[0095] An RBSP video parameter set may include parameters that can be referenced by one or more sequence parameter set RBSPs.
[0096] The relationship and hierarchy between video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) can be described as follows. VPS resides one level above SPS in the parameter set hierarchy and in the context of scalability and / or 3D video. VPS may include parameters that are common to all slices across all layers (scalability or preview) in the entire encoded video sequence. SPS includes parameters that are common to all slices in a particular layer (scalability or preview) in the entire encoded video sequence and may be shared by multiple layers (scalability or preview). PPS includes parameters that are common to all slices in a particular layer representation (the representation of a scalability or preview layer in an access unit) and are likely to be shared by all slices. Petition 870250083300, dated 09 / 16 / 2025, page 37 / 298 28 / 79 in multiple layer representations.
[0097] VPS can provide information about the dependency relationships of the layers in a bitstream, as well as much other information that is applicable to all slices across all layers (scalability or visualization) in the entire encoded video sequence. VPS can be considered as comprising two parts, the base VPS and a VPS extension, where the VPS extension may optionally be present.
[0098] Out-of-band transmission, signaling, or storage may additionally or alternatively be used for purposes other than transmission error tolerance, such as ease of access or session negotiation. For example, a sample input of a track in a file conforming to the ISO Base Media File Format may comprise parameter sets, while the data encoded in the bitstream is stored anywhere in the file or in another file. The expression along the bitstream (e.g., indicating along the bitstream) or along an encoded unit of a bitstream (e.g., indicating along an encoded slice) may be used in the claims and embodiments described to refer to out-of-band transmission, signaling, or storage in a manner in which the out-of-band data is associated with the bitstream or the encoded unit, respectively.The expression "decoding along the bitstream" or "decoding along a coded unit of a bitstream" or similar can refer to the decoding of said out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that are associated with the bitstream or coded unit, respectively.
[0099] A SEI NAL unit may contain one or more SEI messages, which are not required for decoding emitted images, but may assist in related processes such as image output timing, rendering, error detection, error hiding, and resource reservation.
[0100] An encoded image is an encoded representation of a Petition 870250083300, dated 09 / 16 / 2025, page 38 / 298 29 / 79 image.
[0101] In HEVC, an encoded image can be defined as an encoded representation of an image containing all the image's encoding tree units. In HEVC, an access unit (AU) can be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and contain at most one image with any specific nuh_layer_id value. In addition to containing the VCL NAL units of the encoded image, an access unit can also contain non-VCL NAL units. The specified classification rule can, for example, associate images with the same output time or image output count value in the same access unit.
[0102] A bitstream can be defined as a sequence of bits, in the form of a NAL unit stream or a byte stream, that forms the representation of encoded images and associated data that make up one or more encoded video sequences. A first bitstream can be followed by a second bitstream on the same logical channel, such as in the same file or on the same connection of a communication protocol. An elementary stream (in the context of video encoding) can be defined as a sequence of one or more bitstreams. The end of the first bitstream can be indicated by a specific NAL unit, which can be referred to as the bitstream NAL unit end (EOB) and which is the last NAL unit of the bitstream. In HEVC and its current draft extensions, the EOB NAL unit needs to have nuh_layer_id equal to 0.
[0103] In H.264 / AVC, an encoded video sequence is defined as a sequence of consecutive access units in decoding order from one IDR access unit inclusive to the next IDR access unit exclusively, or to the end of the bitstream, whichever comes first.
[0104] In HEVC, a video-encoded sequence (CVS) can be defined, Petition 870250083300, dated 09 / 16 / 2025, page 39 / 298 30 / 79, for example, as a sequence of access units consisting, in decoding order, of an IRAP access unit with NoRaslOutputFlag equal to 1, followed by zero or more access units that are not IRAP access units with NoRaslOutputFlag equal to 1, including all subsequent access units up to, but not including, any subsequent access unit that is an IRAP access unit with NoRaslOutputFlag equal to 1. An IRAP access unit can be defined as an access unit in which the base layer image is an IRAP image. The value of NoRaslOutputFlag is equal to 1 for each IDR image, each BLA image, and each IRAP image that is the first image in which the particular layer in the bitstream in decoding order is the first IRAP image that follows a sequence NAL unit end that has the same nuh_layer_id value in decoding order.There may be ways to provide the HandleCraAsBlaFlag value to the decoder from an external entity, such as a player or receiver, that can control the decoder. HandleCraAsBlaFlag can be set to 1, for example, by a player that searches for a new position in a bitstream or tunes into a broadcast and starts decoding, and then starts decoding from a CRA image. When HandleCraAsBlaFlag is equal to 1 for a CRA image, the CRA image is handled and decoded as if it were a BLA image.
[0105] In HEVC, an encoded video sequence may additionally or alternatively (to the above specification) be specified to terminate when a specific NAL unit, which may be referred to as a sequence end-of-NAL unit (EOS), appears in the bitstream and has nuh_layer_id equal to 0.
[0106] A group of images (GOP) and its characteristics can be defined as follows. A GOP can be decoded independently of whether any previous images have been decoded. An open GOP is such a group of images in which images preceding the initial intraimage are included. Petition 870250083300, dated 09 / 16 / 2025, p. 40 / 298 31 / 79 output order may not be correctly decoded when decoding starts from the initial intraimage of the open GOP. In other words, images from an open GOP may refer (in inter-prediction) to images belonging to a previous GOP. An HEVC decoder can recognize an intraimage starting from an open GOP, due to the fact that a specific NAL unit type, CRA NAL unit type, can be used for its encoded slices. A closed GOP is such a group of images in which all images can be correctly decoded when decoding starts from the initial intraimage of the closed GOP. In other words, no image in a closed GOP refers to any images in previous GOPs. In H.264 / AVC and HEVC, a closed GOP can start from an IDR image. In HEVC, a closed GOP can also start from a BLA_W_RADL or BLA_N_LP image.An open GOP encoding framework is potentially more efficient at compression compared to a closed GOP encoding framework, due to greater flexibility in selecting reference images.
[0107] A Decoded Image Buffer (DPB) can be used in the encoder and / or decoder. There are two reasons for buffering decoded images: for reference in inter-prediction and for reordering decoded images in output order. Because H.264 / AVC and HEVC provide great flexibility for both reference image tagging and output reordering, separate buffers for reference image buffering and output image buffering can waste memory resources. Therefore, the DPB can include a unified process for decoded image buffering for reference images and output reordering. A decoded image can be removed from the DPB when it is no longer used as a reference and is not needed for output.
[0108] In many H.264 / AVC and HEVC encoding modes, the reference image for inter-prediction is indicated with an index to a list of Petition 870250083300, dated 09 / 16 / 2025, page 41 / 298 32 / 79 reference image. The index can be encoded with variable-length encoding, which usually causes a smaller index to have a smaller value for the corresponding syntax element. In H.264 / AVC and HEVC, two reference image lists (reference image list 0 and reference image list 1) are generated for each bipredictive slice (B), and one reference image list (reference image list 0) is formed for each intercoded slice (P).
[0109] Many encoding standards, including H.264 / AVC and HEVC, may have a decoding process to derive a reference image index for a list of reference images, which can be used to indicate which of the multiple reference images is used for inter-prediction for a particular block. A reference image index may be encoded by an encoder in the bitstream in some inter-encoding modes or may be derived (by an encoder and a decoder), for example, using neighboring blocks in some other inter-encoding modes.
[0110] Motion parameter types or motion information may include, but are not limited to, one or more of the following types: - an indication of a prediction type (e.g., intra-prediction, uni-prediction, bi-prediction) and / or a number of reference images; - an indication of a prediction direction, such as inter-prediction (also known as temporal), interlayer prediction, interview prediction, visualization synthesis (VSP) prediction, and intercomponent prediction (which may be indicated by reference image and / or by prediction type and in which, in some modalities, interview prediction and visualization synthesis prediction may be considered together as a prediction direction) and / or - an indication of a reference image type, such as a short-term reference image and / or a long-term reference image and / or an interlayer reference image (which may be indicated, for example, by reference image) Petition 870250083300, dated 09 / 16 / 2025, p. 42 / 298 33 / 79 - a reference index for a list of reference images and / or any other identifier of a reference image (which may be indicated, for example, by reference image and the type of which may depend on the direction of prediction and / or the type of reference image and which may be accompanied by other relevant pieces of information, such as the list of reference images or similar to which the reference index applies); - a horizontal motion vector component (which can be indicated, for example, by prediction block or reference index or similar); - a vertical motion vector component (which can be indicated, for example, by prediction block or reference index or similar); - one or more parameters, such as image order count difference and / or a relative camera separation between the image containing or associated with the motion parameters and its reference image, which can be used for scaling the horizontal motion vector component and / or the vertical motion vector component in one or more motion vector prediction processes (wherein said one or more parameters may be indicated, for example, by each reference image or each reference index or the like); - coordinates of a block to which motion parameters and / or motion information apply, for example, coordinates of the upper left sample of the block in luma sample units; - extensions (for example, a width and a height) of a block to which motion parameters and / or motion information apply.
[0111] Compared to previous video encoding standards, the Versatile Video Codec (H.266 / VVC) introduces a plurality of new encoding tools, such as the following: • Intraprediction Petition 870250083300, dated 09 / 16 / 2025, page 43 / 298 34 / 79 - Intra 67 mode with wide-angle mode extension - Block size and mode-dependent 4-lead interpolation filter - Combination of intra-position-dependent prediction (PDPC) - Intra-component linear model (CCLM) prediction Multireference intra-line prediction - Intra-subpartitions - Intraweighted prediction with matrix multiplication • Interimage prediction - Block movement copy with spatial, temporal, historical, and pairwise averaged embedding candidates - Inter-index prediction of affine motion - Prediction of temporal motion vector based on sub-block Adaptive motion vector resolution - 8x8 block-based motion compression for temporal motion prediction - High-precision (1 / 16 pel) motion vector storage and motion compensation with an 8-lead interpolation filter for the luma component and a 4-lead interpolation filter for the chroma component. - Triangular partitions Combined intra- and inter-combined prediction - Incorporate with MVD (MMVD) - Symmetric MVD Coding - Bidirectional optical flow - Decoder side motion vector refinement - Biprediction with CU level weighting • Transform, quantization, and coding of coefficients Multiple primary transform selection with DCT2, DST7 and Petition 870250083300, dated 09 / 16 / 2025, page 44 / 298 35 / 79 DCT8 Secondary transform for low frequency zone - Transformation from sub-block to interpredicted residue - Dependent quantization with QP max increased from 51 to 63 - Transform coefficient encoding with signal data hiding - Encoding of transform jump residue • Entropy encoding - Arithmetic coding mechanism with adaptive dual-window probability update • Loop filter Loop resizing - Blocking filter with longer, stronger filter Adaptive sample bias Adaptive Loop Filtering • Screen Content Encoding: - Current image reference with region-specific restriction • 360-degree video encoding - Horizontal surrounding motion compensation • High-level syntax and parallel processing - Reference image management with direct flagging of the reference image list. - Clipping groups with rectangular clipping groups
[0112] Partitioning in VVC is performed similarly to HEVC, i.e., each image is divided into coding tree units (CTUs). An image can also be divided into slices, clippings, bricks, and subimages. A CTU can be divided into smaller CUs using a quaternary tree structure. Each CU can be divided using a quadruple tree and a nested multi-type tree including ternary and binary splitting. However, there are specific rules for Petition 870250083300, dated 09 / 16 / 2025, page 45 / 298 36 / 79 infer partitioning at image boundaries, and redundant splitting patterns are prohibited in nested multi-type partitioning.
[0113] In the new coding tools listed above, the cross-component linear model (CCLM) prediction mode is used in VVC to reduce cross-component redundancy. In the same mode, chroma samples are predicted based on luma samples reconstructed from the same CU using a linear model as follows: predc(i, j) = a · recL'(i,j) + β (Eq. 1a)
[0114] where predc(i,j) represents the predicted chroma samples in a CU and recL'(i, j) represents the subsampled reconstructed chroma samples from the same CU.
[0115] Alternatively, the equation below can be used for CCLM: predc(i, j) = a · » k + β (Eq. 1b) where » operation denotes a bit shift to the right by the value k.
[0116] The CCLM parameters (ae β) are derived with at most four neighboring chroma samples and their corresponding subsampled luma samples. The actual chroma block dimensions are assumed to be WxH, then W' and H' are defined as - W' = W, H' = H when LM mode is applied; - W' = W + H when the LM-A mode is applied; - H' = H + W when the LM-L mode is applied;
[0117] In this document, the LM-A mode refers to linear model_above, where only the above model (i.e., sample values from neighboring positions above the CU) is used to calculate the linear model coefficients. To obtain more samples, the above model is extended to (W+H). The LM-L mode, in turn, refers to linear model_left, where only the left model (i.e., sample values from neighboring positions to the left of the CU) is used to calculate the linear model coefficients. To obtain more samples, the left model is extended to (H+W). For a non-square block, the above model Petition 870250083300, dated 09 / 16 / 2025, page 46 / 298 37 / 79 is extended to W+W, the mold on the left is extended to H+H.
[0118] The neighboring positions above are denoted as S[0,-1]...S[W' -1, -1] and the neighboring positions to the left are denoted as S[-1, 0]...S[-1, H' -1]. Then, the four samples are selected as - S[W' / 4, -1 ], S[3 * W7 4, -1 ], S[ -1, H7 4 ], S[ -1, 3 * H7 4 ] when LM mode is applied and both the above and left samples are available; - S[W' / 8, -1], S[3*W7 8, -1], S[5'W78, -1], S[7*W78, -1 ] when LM-A mode is applied or only the above neighboring samples are available; - S[-1, H78], S[-1, 3*H78], S[-1, 5*H78], S[-1, 7 * H78] when the LM-L mode is applied or only left neighbor samples are available;
[0119] The four neighboring luma samples at the selected positions are subsampled and compared four times to find two smaller values: xOA and x1A, and two larger values: xOB and x1B. Their corresponding chroma sample values are denoted as yOA, y1A, yOB, and y1B. Then, xA, xB, yA, and yB are derived as: Xa=(^A + x^A +1)»1; Xb=(x°B + x^B +1)»1; Ya=(y°A + / a +1)»1; Yb=(y°B + / b +1)»1 (Eq. 2)
[0120] Finally, the linear model parameters a and β are obtained according to the following equations: $ = Yb-a-Xb(Eq.4)
[0121] Figure 5 shows an example of the location of the samples to the left and above and of the sample of the current block involved in CCLM mode.
[0122] The division operation to calculate the parameter α is implemented with a lookup table. To reduce the memory required to store the table, the diff value (difference between maximum and minimum values) and the parameter α Petition 870250083300, dated 09 / 16 / 2025, page 47 / 298 38 / 79 are expressed using exponential notation. For example, diff is approximated with a significant part of 4 bits and an exponent. Consequently, the table for 1 / diff is reduced by 16 elements to 16 significant part values as follows: DivTable [ ] = { 0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0} (Eq. 5)
[0123] This provides the benefit of both reducing the complexity of the calculation and the amount of memory required to store the necessary tables.
[0124] To match chroma sample locations for 4:2:0 video sequences, two types of subsampling filters are applied to luma samples to achieve a 2:1 subsampling ratio in both horizontal and vertical directions. The subsampling filter selection is specified by an SPS level flag. The two subsampling filters are as follows, corresponding to “type-O” and “type-2” content, respectively: RecL'(i,;) = recL(2i - 1, 2; - 1) + 2 recL(2i - 1, 2; - 1) + recL(2i + 1, 2; - 1) + Ί recL(2i - 1, 2;) + 2 recL(2i, 2;) + recL(2i + 1, 2;) + 4 6)rpr>(í n =[recL(2i, 2; - 1) + recL(2í - 1,2;) + 4 recL(2i, 2;)1 l +recL(2i + 1,2;) + recL(2i, 2; + 1) + 4 \q' '
[0125] Note that only one luma line (general line storage in intra prediction) is used to produce the subsampled luma samples when the upper reference line is on the CTU boundary.
[0126] This parameter computation is performed as part of the decoding process and is not merely as an encoder fetch operation. As a result, no syntax is used to drive the α and β values to the decoder.
[0127] For intra-mode chroma encoding, a total of 8 intra modes are allowed for intra-mode chroma encoding. These modes include five Petition 870250083300, dated 09 / 16 / 2025, page 48 / 298 39 / 79 traditional intra modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). Chroma mode signaling and the derivation process are shown in Table 1. Chroma mode encoding depends directly on the intra prediction mode of the corresponding luma block. Since the separate block partitioning structure for luma and chroma components is allowed in I slices, a chroma block can correspond to multiple luma blocks. Therefore, for Chroma DM mode, the intra prediction mode of the corresponding luma block that covers the central position of the current chroma block is directly inherited. Chroma prediction mode Corresponding intra luma prediction mode 0 50 18 1 X ( 0 <= X <= 66 ) 0 66 0 0 0 0 1 50 66 50 50 50 2 18 18 66 18 18 3 1 1 1 66 1 4 0 50 18 1 X 5 81 81 81 81 81 6 82 82 82 82 82 7 83 83 83 83 83 Table 1.
[0128] A simple binarization table is used regardless of the sps_cclm_enabled_flag value as shown in Table 2. Value of intra_chroma_pred_mode Binary column 4 00 0 0100 1 0101 Petition 870250083300, dated 09 / 16 / 2025, page 49 / 298 40 / 79 2 0110 3 0111 5 10 6 110 7 111 Table 2.
[0129] In Table 2, the first binary indicates whether it is regular (0) or LM modes (1). If it is LM mode, then the next binary indicates whether it is LM_CHROMA (0) or not. If it is not LM_CHROMA, the next binary 1 indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first binary in the binarization table for the corresponding intra_chroma_pred_mode can be discarded before entropy encoding. Or, in other words, the first binary is inferred to be 0 and therefore not encoded. This simple binarization table is used for both cases of sps_cclm_enabled_flag equal to 0 and 1. The first two binaries in Table 2 are context-encoded with their own context model, and the remaining binaries are bypass-encoded.
[0130] In addition, in order to reduce luma-chroma latency in dual-tree, when the 64x64 luma encoding tree node is partitioned with Undivided (and intra-subpartitions (ISPs) are not used for the 64x64 CU) or QT, the chroma CUs in the 32x32 / 32x16 chroma encoding tree node are allowed to use CCLM as follows: If the 32x32 chroma node is not split or is split by QT partitioning, all chroma CUs in the 32x32 node can use CCLM. - If the 32x32 chroma node is partitioned with Horizontal BT, and the 32x16 child node does not partition or uses Vertical BT partitioning, all chroma CUs in the 32x16 chroma node can use CCLM.
[0131] In all other split luma and chroma encoding tree conditions, CCLM is not allowed for chroma CU. Petition 870250083300, dated 09 / 16 / 2025, page 50 / 298 41 / 79
[0132] Multiple Model Logistics (MMLM)
[0133] The CCLM included in VVC is extended by adding three Multiple LM Model (MMLM) modes. In each MMLM mode, reconstructed neighboring samples are classified into two classes using a threshold that is the average of the reconstructed neighboring luma samples. The linear model of each class is derived using the Least Squares Method (LMS). For the CCLM mode, the LMS method is also used to derive the linear model. Figures 6a and 6b illustrate two luma-to-chroma models obtained for luma threshold (Y) of 17 in the sample domain and spatial domain, respectively. Each luma-to-chroma model has its own linear model parameters α and β. As can be seen in Figure 6b, each luma-to-chroma model corresponds to a spatial segmentation of the content (i.e., they correspond to different objects or textures in the scene).
[0134] Cross-Component Convolutional Model (CCCM)
[0135] An improved version of cross-component prediction, known as CCCM, uses a 2D filter kernel to derive the luma-to-chroma model. Filter coefficients are derived on the decoder side using a reconstructed set of chroma samples and input data. For filter coefficient derivation, colocalized reference sample areas (consisting of reconstructed luma and chroma samples) are defined for both luma and chroma as shown in Figure 7, where the typically used 4:2:0 chroma subsampling has been applied. The reference sample area for a given block might be, for example, six lines above and to the left as shown in Figure 7, but any number of reference lines (which can be realized by either the encoder or the decoder) can be used.Generally, reference samples can contain any chroma and luma samples that have been reconstructed by both the encoder and the decoder. Once the reference samples are determined, the filter coefficients can be derived, for example, using different types of linear regression tools, such as estimation of... Petition 870250083300, dated 09 / 16 / 2025, page 51 / 298 42 / 79 least common squares, orthogonal match search, optimized orthogonal match search, ridge regression or selection operator, and absolute minimum reduction.
[0136] The dimensions of the filter kernel can be, for example, 1x3 (1D vertical), 3x1 (1D horizontal), 3x3, 7x7 or any dimensions and can be shaped (by selecting only a subset of all possible kernel locations) as a cross or a diamond (as shown in Figure 8) or as any determined shape. When referring to samples within the filter kernel, the following notation is used: north (above), east (right), south (below), west (left) and center, as illustrated in Figure 8 using the letters N, E, S, W, C.
[0137] The general method of chroma sample reconstruction using convolution between a filter kernel obtained from the decoder side and an input dataset is referred to as the cross-component convolutional model (CCCM) here. The following steps can be applied to perform a CCCM operation: 1) Define colocalized reference areas on luma and chroma components. 2) Subsample the luma samples to match the chroma grid (optional). 3) Scan the luma and chroma samples from the reference area and collect available statistics (such as autocorrelation matrix and cross-correlation vector) based on the filter format. 4) Solve for the filter coefficients by minimizing squared error (or any other metric) based on available statistics (such as the autocorrelation matrix and the cross-correlation vector). 5) Calculate a chroma block predicted by convolution of luma samples with reduced sampling using the filter kernel.
[0138] Let's define luma samples (possibly subsampled) as a 2D Y(x,y) array indexed using horizontal x-coordinate and Petition 870250083300, dated 09 / 16 / 2025, page 52 / 298 43 / 79 vertical y-coordinate. We will also define the colocalized chroma samples as a 2D array C(x,y) and the filter kernel (i.e., coefficients) as a 3x3 array F(i,j). At the sample level, we define the convolution between Y and F as, 7=1 t=l C(x,y) =ΣΣY(x + i,y + f) F(i + l,j + 1). j=-ií=—ι
[0139] When other data terms are used, such as the non-linear square root term, the above convolution becomes, (7=1 t = l \ ΣΣY(x + i,y + f) F(i + l,j + 1) j + F(0) · y]Y(x, y), j=-lí=-l / where F are filter coefficients that reside outside the 2D filter kernel, but which were obtained as part of the system of linear equations that were used to solve for the 2D filter coefficients in Step 4 above. Similarly, we can add the bias term to the convolution with, (7 = 1 t = l \ ΣΣY(x + i,y+j)· F(i + l,j + 1) j + F(0) + F(l) · ^Y(x,y). j=-lt= —1 /
[0140] Multireference intraline prediction (MRL)
[0141] Multireference intra-line prediction (MRL) uses more reference lines for intra-prediction. In Figure 9, an example of 4 reference lines is depicted, where the samples of segments A and F are not sourced from reconstructed neighboring samples, but populated with the nearest samples from Segments B and E, respectively. HEVC intra-image prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used.
[0142] The selected reference line index (mrljdx) is signaled and used to generate the intra predictor. For reference line idx, which is greater than 0, only include additional reference line modes in the MPM list and only Petition 870250083300, dated 09 / 16 / 2025, page 53 / 298 44 / 79 signaling mpm index without remnant mode. The reference line index is signaled before intra prediction modes, and the Flat mode is excluded from intra prediction modes if a non-zero reference line index is signaled.
[0143] MRL is disabled for the first block row within a CTU to prevent the use of extended reference samples outside the current CTU row. Additionally, PDPC is disabled when an additional row is used. For MRL mode, the DC value derivation in intra DC prediction mode for non-zero reference row indices is aligned with that of reference row index 0. MRL requires the storage of 3 neighboring luma reference rows with a CTU to generate predictions. The CCLM tool also requires 3 neighboring luma reference rows for its subsampling filters. The MRL definition to use the same 3 rows is aligned with CCLM to reduce storage requirements for decoders.
[0144] Intra subpartitions (ISP)
[0145] Intra subpartitions (ISPs) divide intrapredicted luma blocks vertically or horizontally into 2 or 4 subpartitions depending on the block size. For example, the minimum block size for ISPs is 4x8 (or 8x4). If the block size is greater than 4x8 (or 8x4), then the corresponding block is divided into 4 subpartitions. It has been observed that ISP blocks M*128 (with M<64) and 128χN (with N<64) could generate a potential problem with the 64x64 VDPU. For example, a CU Mχ128 in the case of a simple tree has one Mx128 luma TB and two corresponding M / 2x64 chroma TBs. If the CU uses ISP, then the luma TB will be split into four 32-bit TBs (only horizontal splitting is possible), each smaller than a 64x64 block. However, in the current ISP design, chroma blocks are not split. Therefore, both chroma components will be larger than a 32x32 block. Analogously, a similar situation could be created with a 128xN CU that uses ISP.Therefore, these two cases are a problem for the 64x64 decoder line. For this reason, the CU sizes that can use ISPs. Petition 870250083300, dated 09 / 16 / 2025, page 54 / 298 45 / 79 are restricted to a maximum of 64x64. All subpartitions satisfy the condition of having at least 16 samples.
[0146] Matrix Intraweighted Prediction (MIP)
[0147] Matrix-weighted intraprediction (MIP) is a newly added intraprediction technique in VVC. To predict samples from a rectangular block of width W and height H, matrix-weighted intraprediction (MIP) takes a line of reconstructed neighboring boundary samples H to the left of the block and a line of reconstructed neighboring boundary samples W above the block as input. If reconstructed samples are unavailable, they are generated as in conventional intraprediction. The generation of the prediction signal is based on the following three steps, which are averaging, matrix vector multiplication, and linear interpolation as shown in Figure 10.
[0148] Decoder-side intra-mode derivation (DIMD)
[0149] When DIMD is applied, two intra modes are derived from the reconstructed neighboring samples, and these two predictors are combined with the planar mode predictor with weights derived from the gradients, as described in JVET-O0449. The split operations in weight derivation are performed using the same lookup table-based integralization (LUT) scheme used by CCLM. For example, the split operation in orientation calculation. Orient = Gy / Gxé computed by the following scheme based on LUT: x = Piso( Log2( Gx )) normDiff = (( Gx« 4 ) » x ) & 15 x +=( 3 + ( normDiff != 0 ) ? 1 : 0 ) Orient = (Gy* ( DivSigTable[ normDiff] | 8 ) + ( 1«( x-1 )))» x where DivSigTable
[16] = { 0, 7, 6, 5 ,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1,0}. Petition 870250083300, dated 09 / 16 / 2025, page 55 / 298 46 / 79
[0150] Intra-derived modes are included in the primary list of most probable intra modes (MPM), so that the DIMD process is performed before the construction of the MPM list. The primary intra-derived mode of a DIMD block is stored with a block and is used for the construction of the MPM list of neighboring blocks.
[0151] Fusion for intra-mold-based mode derivation (TIMD)
[0152] For each intra prediction mode in MPMs, the SATD between the prediction and mold reconstruction samples is calculated. The first two intra prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are merged with the weights after applying the PDPC process, and this weighted intra prediction is used to encode the current CU. The combination of position-dependent intra prediction (PDPC) is included in the derivation of the TIMD modes.
[0153] The costs of the two selected modes are compared with a limit; in the test, the cost factor 2 is applied as follows: costMode2 < 2*costMode1.
[0154] If this condition is true, the merge is applied, otherwise only the model is used.
[0155] The mode weights are calculated from their SATD costs as follows: weightl = costMode2 / (costMode1+ costMode2) weight2 = 1 - weightl
[0156] Division operations are conducted using the same lookup table-based integerization (LUT) scheme used by CCLM.
[0157] Low-frequency non-separable transform (LFNST)
[0158] In VVC, the LFNST is applied between the forward primary transform and quantization (on the encoder side) and between dequantization and the inverse primary transform (on the decoder side), as shown in Figure 11. In the LFNST, the 4x4 non-separable transform or the 8x8 non-separable transform is applied according to the block size. For example, the 4x4 LFNST is Petition 870250083300, dated 09 / 16 / 2025, page 56 / 298 47 / 79 is applied to small blocks (i.e., min (width, height) < 8) and LFNST 8x8 is applied to larger blocks (i.e., min (width, height) > 4).
[0159] The application of a non-separable transform, which is being used in LFNST, is described below using the input as an example. To apply LFNST 4x4, the 4x4 input block X ^00 -^01 ^02^03 y _ Ύ10 ^12^13 ^20 ^21 ^22^23 -X30 ^31 ^32^33- is first represented as a vector Ύ: X — [^00 ^01 Y02 Y03 *10 *11 *12 *13 *20 *21 *22 *23 *30 *31 *32 *3s]T
[0160] The non-separable transform is calculated as F = TX, where F indicates the transform coefficient vector, and T is a transform matrix. 16x16. The 16x1 coefficient vector F is subsequently rearranged as a 4x4 block using the scanning order of that block (horizontal, vertical, or diagonal). Coefficients with lower indices will be placed with the lower scanning index in the 4x4 coefficient block.
[0161] Reduced non-separable transform
[0162] LFNST (low-frequency non-separable transform) is based on the direct matrix multiplication approach to apply non-separable transform, so that it is implemented in a single pass without multiple iterations. However, the dimension of the non-separable transform matrix needs to be reduced to minimize computational complexity and memory space for storing the transformation coefficients. Therefore, the reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (N is usually equal to 64 for 8x8 NSST) into an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Therefore, instead of the NxN matrix, the RST matrix becomes an RxN matrix as Petition 870250083300, dated 09 / 16 / 2025, page 57 / 298 48 / 79 ^11 ^21 ^Rl next: ^12 ^13 _ tlN ^22 ^23 ^2NfR2fR3 tRN. where the R rows of the transform are R bases of N-dimensional space.
[0163] The inverse transform matrix for RT is the transpose of its direct transform. For 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix, which is a conventional 8x8 non-separable transform matrix size, is reduced to a 16x48 direct matrix. Therefore, the 48*16 inverse RST matrix is used on the decoder side to generate principal (primary) transform coefficients in upper left 8*8 regions. When 16x48 matrices are applied instead of 16x64 with the same transform set configuration, each of these receives 48 input data from three 4x4 blocks in an upper left 8x8 block, excluding the lower right 4x4 block. Thanks to the reduced size, the memory usage for storing all LFNST matrices is reduced from 10 KB to 8 KB with a reasonable performance drop.To reduce complexity, LFNST is restricted to be applicable only if all coefficients outside the first coefficient subgroup are not significant. Therefore, all primary transform coefficients must be zero when LFNST is applied. This allows for conditioning the LFNST index signaling at the last significant position and thus avoids extra coefficient scanning in the current LFNST design, which is necessary to check significant coefficients only at specific positions. The worst-case treatment of LFNST (in terms of multiplications per pixel) restricts non-separable transforms for 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In these cases, the last significant scan position must be less than 8 when LFNST is applied for other sizes smaller than 16. For blocks with 4xN and Nx4 shapes and N > 8, the proposed restriction... Petition 870250083300, dated 09 / 16 / 2025, page 58 / 298 49 / 79 implies that the LFNST is now applied only once, and only in the upper left 4x4 region. Since all primary coefficients are zero when LFNST is applied, the number of operations required for primary transforms is reduced in these cases. From the encoder's perspective, coefficient quantization is notably simplified when LFNST transforms are tested. An optimized distortion rate quantization must be performed at most for the first 16 coefficients (in scan order); the remaining coefficients are forced to zero.
[0164] LFNST transform selection
[0165] There are a total of 4 transform sets and 2 non-separable transform matrices (kernels) per transform set are used in LFNST. The mapping of the intra prediction mode to the transform set is predefined as shown in the table below. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModelntra <= 83), transform set 0 will be selected for the current chroma block. For each transform set, the selected candidate non-separable secondary transform is further specified by the explicitly signaled LFNST index. The index is signaled in a bitstream once per Intra CU after the transform coefficients. IntraPredMode Tr. set index IntraPredMode < 0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= lntraPredMode <= 80 1 Petition 870250083300, dated 09 / 16 / 2025, page 59 / 298 50 / 79 81 <= lntraPredMode<= 83 0 Transform selection table.
[0166] LFNST Index Signaling and Interaction with Other Tools
[0167] Because LFNST is restricted to being applicable only if all coefficients outside the first coefficient subgroup are not significant, LFNST index coding depends on the position of the last significant coefficient. Furthermore, the LFNST index is context-coded, but does not depend on the intra prediction mode, and only the first bin is context-coded. Additionally, LFNST is applied to intra CU in intra and inter slices, and to Luma and Chroma. If a double tree is enabled, the LFNST indices for Luma and Chroma will be signaled separately. For inter slices (dual tree is disabled), a single LFNST index is signaled and used for Luma and Chroma.
[0168] Considering that a large CU larger than 64x64 is implicitly split (TU mosaic) due to the existing maximum transform size restriction (64x64), an LFNST index lookup could increase the data buffer by four times for a certain number of stages in the decoding pipeline. Therefore, the maximum allowed size for LFNST is restricted to 64x64. Note that LFNST is enabled only with DCT2. The LFNST index signaling is placed before the MTS index signaling.
[0169] The use of scaling matrices for perceptual quantization does not make it clear that the scaling matrices that are specified for the primary matrices can be useful for LFNST coefficients. Therefore, the use of scaling matrices for LFNST coefficients is not allowed. For the single-tree partition mode, the chroma LFNST is not applied.
[0170] Enhanced Multiple Transform Selection (MTS) for intra-encoding
[0171] In the current VVC project, for MTS, only DST7 and DCT8 transform kernels are used, which are used for intra and inter encoding.
[0172] Additional primary transforms including DCT5, DST4, DST1, and Petition 870250083300, dated 09 / 16 / 2025, page 60 / 298 51 / 79 identity transform (IDT) methods are employed. Furthermore, the set of MTS is made dependent on TU size and intra-mode information. 16 different TU sizes are considered, and for each TU size, 5 different classes are considered dependent on intra-mode information. For each class, 1, 4, or 6 different transform pairs are considered. The number of intra-mode MTS candidates is adaptively selected (between 1, 4, and 6 MTS candidates) depending on the sum of the absolute values of the transform coefficients. The sum is compared to two fixed thresholds to determine the total number of allowed MTS candidates. candidate: sum <= thO candidates: thO < sum <= th1 candidates: sum > th1
[0173] It is observed that, although a total of 80 different classes are considered, some of these different classes frequently share exactly the same transform set. So there are 58 (less than 80) unique entries in the resulting LUT.
[0174] For angular modes, a joint symmetry with respect to TU shape and intra-prediction is considered. Thus, a mode i (i > 34) with TU shape AχB will be mapped to the same class corresponding to the mode j = (68 - i) with TU shape BχA. However, for each transform pair, the order of the horizontal and vertical transform kernels is swapped. For example, a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same class. However, the vertical and horizontal transform kernels are swapped. For wide-angle modes, the nearest conventional angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80.
[0175] Inter Multiple Transform Selection (MTS) Optimization
[0176] For the MTS of intercoded CUs, four candidates: {(DST7, Petition 870250083300, dated 09 / 16 / 2025, p. 61 / 298 52 / 79 DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)} are used for each CU. For higher resolution sequences (width > 1080), the maximum CU size for Inter-MTS use is set to 32 (i.e., Inter-MTS is used for CUs with width <= 32 and height <= 32), and for the remaining sequences (lower resolution) it is set to 16. For 4-pt, 8-pt, and 16-pt transforms, the current AMT transform cores, i.e., DST-7 and DCT-8, are replaced by separable KLTs, as proposed in JVET-J0021.
[0177] Intramold correspondence
[0178] Intra-mold matching prediction (Intra TMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped mold matches the current mold. For a predefined search range, the encoder looks for the mold most similar to the current mold in a reconstructed portion of the current frame and uses the matching block as a prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side.
[0179] The prediction signal is generated by matching the causal neighbor in an L-shape of the current block with another block in a predefined search area shown in Figure 12, consisting of: - R1:CTU current - R2: Upper left CTU - R3: CTU above - R4: Left CTU
[0180] Sum of absolute differences (SAD) is used as a cost function.
[0181] Within each region, the decoder searches for the template that has the lowest SAD relative to the current one and uses its corresponding block as a prediction block.
[0182] The dimensions of all regions (SearchRange_w, SearchRange_h) are defined proportionally to the block dimension (BlkW, BlkH) to have a Petition 870250083300, dated 09 / 16 / 2025, p. 62 / 298 53 / 79 fixed number of SAD comparisons per pixel. That is: Search Range_w = a * BlkW - SearchRange_h = a * BlkH where “a” is a constant that controls the gain / complexity ratio. In practice, “a” is equal to 5.
[0183] The Intramold Match tool is enabled for CUs with a width and height of 64 or less. This maximum CU size for Intramold Match is configurable.
[0184] The Intramold Match prediction mode is signaled at the CU level via a dedicated flag when DIMD is not used for the current CU.
[0185] Intrablock Copy (IBC) with Mold Match
[0186] Mold matching is used in IBC for both IBC embedding mode and IBC AMVP mode.
[0187] The IBC-TM embedding list is modified compared to that used by common IBC embedding mode, so that candidates are selected according to a removal method with a movement distance between candidates as in common TM embedding mode. The final zero movement compliance is replaced by left (-W, 0), up (0, -H) and top left (-W, -H) movement vectors, where W is the width and H the height of the current CU.
[0188] In IBC-TM embedding mode, selected candidates are refined using the Mold Matching method before the decoding or RDO process. IBC-TM embedding mode was put into competition with common IBC embedding mode and a TM embedding flag is signaled.
[0189] In IBC™ AMVP mode, up to 3 candidates are selected from the IBC™ incorporation list. Each of these 3 selected candidates is refined using the Mold Matching method and ranked according to its resulting Mold Matching cost. Only the top 2 Petition 870250083300, dated 09 / 16 / 2025, page 63 / 298 The first 54 / 79 are then considered in the motion estimation process as usual.
[0190] The Mold Match refinement for both IBC-TM and AMVP embedding modes is quite simple, since IBC motion vectors are restricted (i) to being integers and (ii) within a reference region as shown in Figure 13a. Therefore, in IBC-TM embedding mode, all refinements are performed in integer precision, and in IBC-TM AMVP mode, they are performed in integer or 4-pel precision depending on the AMVR value. This refinement only accesses samples without interpolation. In both cases, the refined motion vectors and the mold used in each refinement step must respect the reference region restriction.
[0191] IBC reference area
[0192] The reference area for IBC is extended to two rows of CTU above. Figure 13b illustrates the reference area for CTU (m,n) encoding. Specifically, for CTU (m,n) to be encoded, the reference area includes CTUs with index (m-2,n-2)...(W,n-2),(0,n-1)...(W,n-1),(0,n)...(m,n), where W denotes the maximum horizontal index within the current clipping, slice, or image. When CTU size is 256, the reference area is limited to one row of CTU above. This configuration ensures that, for CTU size being 128 or 256, IBC does not require extra memory on the current ETM platform. The sample block vector search range (or so-called local search) is limited to [-(C « 1), C » 2] horizontally and [-C, C » 2] vertically to match the extent of the reference area, where C denotes the size of the CTU.
[0193] Reconstruction-Reordered IBC (RR-IBC)
[0194] A Reconstruction-Reordered (RR-IBC) mode of IBC is allowed for IBC-encoded blocks. When RR-IBC is applied, the samples in a reconstruction block are inverted according to a type of inversion of the current block. On the encoder side, the original block is inverted before motion search and residual calculation, while the prediction block is derived without Petition 870250083300, dated 09 / 16 / 2025, page 64 / 298 55 / 79 inversion. On the decoder side, the reconstruction block is inverted back to restore the original block.
[0195] Two inversion methods, horizontal inversion and vertical inversion, are supported for RR-IBC encoded blocks. A syntax flag is first signaled for an IBC AMVP encoded block, indicating whether the reconstruction is inverted, and if inverted, another flag is additionally signaled specifying the inversion type. For IBC embedding, the inversion type is inherited from neighboring blocks, without syntax signaling. Considering horizontal or vertical symmetry, the current block and the reference block are normally aligned horizontally or vertically. Therefore, when a horizontal inversion is applied, the vertical component of the BV is not signaled and is inferred to be equal to 0. Similarly, the horizontal component of the BV is not signaled and is inferred to be equal to 0 when a vertical inversion is applied.
[0196] To better utilize the symmetry property, an inversion-aware BV fitting approach is applied to refine the candidate block vector. For example, as shown in Figure 14, (xnbr, ynbr) and (xcur, ycur) represent the sample center coordinates of the neighboring block and the current block, respectively, BVnbr and BVcur denote the BV of the neighboring block and the current block, respectively. Instead of directly inheriting the BV of a neighboring block, the horizontal component of BVcur is calculated by adding a motion offset to the horizontal component of BVnbr (denoted as BVnbrh) in the case where the neighboring block is encoded with a horizontal inversion, i.e., BVcurh = 2(xnbr - xcur) + BVnbrh. Similarly, the vertical component of BVcur is calculated by adding a motion offset to the vertical component of BVnbr (denoted as BVnbrv) in the case where the neighboring block is encoded with a vertical inversion, that is, BVcurv = 2(ynbr - ycur) + BVnbrv.
[0197] IBC embedding mode with block vector differences (IBCMBVD)
[0198] MMVD and GPM-MMVD were adopted for ECM as a Petition 870250083300, dated 09 / 16 / 2025, page 65 / 298 56 / 79 extension of common MMVD mode. It is natural to extend the MMVD mode to the IBC embedding mode.
[0199] In IBC-MBVD, the distance defined is {1-pel, 2-pel, 4-pel, 8-pel, 12-pel, 16-pel, 24-pel, 32-pel, 40-pel, 48-pel, 56-pel, 64-pel, 72-pel, 80-pel, 88-pel, 96-pel, 104-pel, 112-pel, 120-pel, 128-pel}, and the BVD directions are two horizontal directions and two vertical directions.
[0200] Base candidates are selected from the first five candidates in the reordered IBC incorporation list. And, based on the SAD cost between the mold (one row above and one column to the left for the current block) and its reference for each refining position, all possible MBVD refining positions (20x4) for each base candidate are recorded. Finally, the top 8 refining positions with the lowest mold SAD costs are kept as available positions, consequently for MBVD index coding. The MBVD index is binarized by the rice code with the parameter equal to 1.
[0201] An IBC-MBVD encoded block does not inherit the inversion type of a neighboring RR-IBC encoded block.
[0202] Cross-component intra-prediction tools, such as CCLM and CCCM discussed above, use samples of reconstructed neighboring samples to calculate the parameters of the prediction model. The assumption is that the reconstructed neighboring samples have good correlation with the samples within the prediction block. However, if the neighboring samples have no correlation with the samples within the current block, or if the correlation is very small, then the calculated model may not be able to predict the samples within the block efficiently.
[0203] Now an improved method for improving the efficiency of cross-component prediction tools is introduced.
[0204] A method according to one aspect is shown in Figure 15, wherein the method comprises receiving (1500) a frame image block unit, the image block unit comprising samples in channels Petition 870250083300, dated 09 / 16 / 2025, page 66 / 298 57 / 79 of color, wherein the color channels comprise at least one chrominance channel and one luminance channel; reconstruct (1502) samples of said luminance channel from the image block unit; determine (1504) a reference area for predicting target samples of at least one color channel from the image block unit, wherein said reference area comprises one or more reference samples in an actual block or colocalized in an actual color channel / frame or in a reference color channel / frame encoded using an intrablock copy (IBC) method; determine (1506) a block vector indicating a spatial distance from a target sample area to the reference area; and predict (1508) said target samples of at least one color channel from the image block unit using the cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
[0205] Thus, in the method, when a block or area is encoded in cross-component prediction mode, if the colocalized block or area in the reference channel is encoded using an IBC mode, then the associated IBC mode block vector is used to indicate a reference region for cross-component prediction. The reference region samples, in the current reference channel, can then be used to calculate the prediction model parameters.
[0206] Figure 16 illustrates the basic principle of the current block and reference block using a block vector in intrablock copy (IBC) mode. In IBC methods, block prediction is achieved by copying reconstructed samples from a different region in the image that has the best content matching the current block. This could be done by various means, for example, encoder-side search and block vector signaling or offset or indexing the best block vector from a list of candidate block vectors of the corresponding block in the bitstream with or without explicit vector difference signaling, or this could be done using template-based methods on the encoder-side. Petition 870250083300, dated 09 / 16 / 2025, page 67 / 298 58 / 79 decoder without block vector signaling, or it could be a combination of both.
[0207] In intrablock copy methods, prediction for the current block is obtained by copying the reconstructed samples from another block or area. The spatial distance from the current block to the reference block is indicated by a block vector (BV).
[0208] Depending on one embodiment, the cross-component prediction method is a cross-component linear model (CCLM) or a cross-component convolutional model (CCCM).
[0209] Several embodiments are described below, primarily in the context of improving the performance of the cross-component convolutional model (CCCM). It is important to understand that the CCCM method is given as an example in this invention, and the proposed methods and embodiments can be used in any other method with similar concepts, such as local illumination compensation (LIC) and other intra- or inter-channel prediction methods. It is further noted that cross-channel prediction can be performed from one chroma channel to another chroma channel (e.g., Cb to Cr or vice versa), or from a chroma channel to luma (Cb / Cr to Y) or vice versa.
[0210] Figure 17a shows an example of using a colocalized IBC-coded block BV to indicate the reference area for parameter derivation for cross-component prediction.
[0211] According to one embodiment, parameter derivation for the cross-component model is performed using samples from a reference block in the current channel and a reference block in the reference channel. After the cross-component prediction parameters have been derived (also known as trained), the colocalized block samples in the reference channel are used to predict the current block using the derived model. Figure 17b shows an example illustrating this embodiment.
[0212] According to one embodiment, the method comprises inferring the block vector of a block colocalized in the reference channel; and scaling the vector of Petition 870250083300, dated 09 / 16 / 2025, page 68 / 298 59 / 79 block to match the current channel sampling density.
[0213] Thus, the inferred colocalized block vector in the reference channel can be scaled to match the sampling density of the current channel. For example, in 4:2:0 or 4:2:2 YUV / YCbCr chroma formats, the inferred colocalized luma channel BV can be scaled down to obtain a corresponding reference block in the current channel. Figure 17c shows an example of scaled BV when the current and reference channel sampling formats are different.
[0214] According to one embodiment, the method comprises identifying the closest match to the reference area in the current channel using a template-match-based search in the current channel.
[0215] Therefore, a mold matching (TM) based search, for example, similar to TM-based search in IBC, can be used in the current channel to identify the closest match to a reference block or area in the current channel. Then the BV from the TM-based search is used to identify the matching region in the reference channel. Some or all of the reference block samples in the current channel and reference block samples in the reference channel are used to calculate the parameters of the cross-component prediction model.
[0216] According to one embodiment, the block vector inferred from the reference channel is at least partially misaligned with colocalized block coordinates.
[0217] Thus, the BV derived from the reference channel may not be exactly aligned with the coordinates of the colocalized block. The BV may be derived from one or more blocks in the colocalized region in the reference channel. The colocalized region can be defined as an area containing the colocalized block and an area extending in different directions from the colocalized block. The size of the extended area can be predefined or could be determined using different criteria, such as block size, channel type, etc.
[0218] Figure 17d shows an example where TM-based search is Petition 870250083300, dated 09 / 16 / 2025, page 69 / 298 60 / 79 is used in the current channel to identify the best matching reference block in the current channel. As shown in Figure 17d, the BV of the TM-based search process is then used to identify the reference block in the reference channel.
[0219] According to one embodiment, the method comprises calculating the cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - one or more samples of the reference block in the current channel; - one or more samples of the reference block in the reference channel; - one or more samples of the block colocalized in the reference channel; - one or more neighborhood samples of the reference block in the current channel; - one or more neighborhood samples of the reference block in the reference channel; - one or more neighborhood samples of the current block in the current channel; - one or more neighborhood samples of the colocalized block in the reference channel. [ 0220] According to one embodiment, the method comprises calculating the cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - Horizontal and / or vertical coordinates of said samples; - Directional gradient values of said samples.
[0221] Thus, with regard to the samples mentioned in the previous embodiment, they can be applied taking into account their horizontal and / or vertical coordinates and / or directional gradients, for example. Gradient values can be calculated in different directions. For example, horizontal gradient can be calculated by subtracting left sample from right sample or vice versa, vertical gradient can be calculated by subtracting above sample from below sample or vice versa.
[0222] According to one embodiment, the method comprises using the block vector of the colocalized block inferred from the reference channel as a vector of Petition 870250083300, dated 09 / 16 / 2025, page 70 / 298 61 / 79 initial block to the current block.
[0223] In this document, the BV of the IBC-encoded colocalized block of the reference channel can be used as the initial BV for the current block. Then a TM-based refinement is performed on the initial BV in the current channel to find the best match in the current channel. A predefined search range can be defined when conducting the refinement process.
[0224] According to one embodiment, the method comprises generating a list of candidates for a block vector of encoded blocks using the intrablock copy (IBC) method in the reference channel in the colocalized region and / or in the neighborhood of the current block in the current channel.
[0225] Thus, a list of BV candidates can be generated from the IBC-encoded blocks in the reference channel in the colocalized region and / or IBC-encoded blocks in the neighborhood of the current block in the current channel. The encoder can test all BV candidates in the list by generating different cross-component predictions, as described above, and the index of the best-performing candidate is signaled in the bitstream. Alternatively, the best-performing candidate is determined by a TM-based process on both the encoder and decoder sides.
[0226] According to one embodiment, the method involves combining two or more predictions to obtain a final prediction of the current block.
[0227] Consequently, the final block prediction can be obtained by combining two or more different predictions. For example, two distinct BVs can be used to derive two distinct cross-component models. Then, two predictions are generated using the models and the final prediction is obtained by combining the predictions. The weights of combinations can be predefined, signaled in the bitstream, or can be determined on the decoder side.
[0228] According to one modality, multiple predictions can be obtained for combined predictions based on one or more of the following: - one or more cross-component predictions can be obtained by Petition 870250083300, dated 09 / 16 / 2025, page 71 / 298 62 / 79 block vectors determined from IBC-coded blocks in colocalized regions in a reference channel; - One or more cross-component predictions can be obtained by block vectors determined from IBC-encoded blocks in the vicinity of the current block in the current channel; - one or more predictions obtained by copying reference block samples into the current channel using block vectors determined from IBC-encoded blocks in colocalized regions in the reference channel; - one or more predictions obtained by copying reference block samples into the current channel using block vectors determined from IBC-encoded blocks in the neighborhood of the current block in the current channel.
[0229] According to one embodiment, some or all of the reference channel samples may be resampled before using them in the parameter derivation process for cross-component prediction.
[0230] According to one embodiment, some or all of the reference channel and / or current channel samples may be filtered before using them in the parameter derivation process for cross-component prediction.
[0231] According to one embodiment, the type of cross-component model is inferred from the reference area or block.
[0232] For example, if the BV points used for a reference block are coded in CCLM or CCCM method, then the same type of cross-component prediction can be used for the current block.
[0233] According to one embodiment, when the colocalized reference channel block is IBC-encoded, the corresponding IBC BV for the current channel (scaled depending on the relationship between the channels) can be used to check the reference area in the current channel to determine if there are one or more blocks encoded with cross-component prediction. If so, the cross-component prediction parameters of these blocks can be directly inherited to the current block using a location-based scan order or some other priority criterion. Alternatively, Petition 870250083300, dated 09 / 16 / 2025, page 72 / 298 63 / 79 These parameters can be added to a list of candidate cross-component predictors for the current mode from which the encoder can select and signal the index.
[0234] According to a modality, one or more of the parameters inherited from the corresponding reference area / block, determined according to the previous modality, can be modified before using them in the current block. For example, the bias term of the prediction model can be recalculated based on the neighboring reference samples of the current block.
[0235] According to one embodiment, the method comprises obtaining the final cross-component prediction parameters of the current block using a two-step derivation, wherein a first-step derivation is performed using the reference block in the reference channel pointed to by the block vector and a second-step derivation is performed in the neighborhood of the current block.
[0236] Thus, the final cross-component prediction parameters of the current block can be obtained using a two-step derivation. First-step derivation is performed using the reference block, to which the IBC BV points in the reference channel, and second-step derivation operates in the neighborhood of the current block. Figure 17e shows an example of the process according to the modality. Parameters derived in the first step are used to predict the neighborhood of current blocks in the current channel, and then second-step parameters are derived to predict the difference between the neighborhood of the current block and its prediction. Final prediction is the combination of the predictions obtained by applying the first- and second-step parameters to the colocalized block in the reference channel.
[0237] According to one embodiment, the reference block and the current block can be divided into sub-blocks with each sub-block having its own CCLM or CCCM model. Overlapping sub-blocks can be used to smooth the prediction obtained.
[0238] According to one embodiment, the CCLM or CCCM model derived in the mold area of the current block can be mixed with the model Petition 870250083300, dated 09 / 16 / 2025, page 73 / 298 64 / 79 derived in the reference block (as indicated by BV). For example, the prediction can merge the outputs of the two models. The weights used in the merging can vary as a function of location within the block, for example, giving more weight to the mold-based model in the upper left region and less to the lower right. The merging weights can be predefined, signaled in the bitstream, or can be determined on the decoder side.
[0239] According to one embodiment, one or more of the intra-block prediction modes of the reference block to which the BV is pointing can be used to predict the current block. Alternatively, one or more intra-block prediction modes of the reference block pointed to by BV are combined with cross-component prediction obtained from reference block samples and / or cross-component prediction obtained from current block neighborhood reference samples.
[0240] According to one embodiment, the method comprises using a co-located intrablock copy-encoded block residue as a segmentation for template derivation.
[0241] Thus, the residual of the colocalized IBC-coded block can be used as a segmentation to guide CCLM or CCCM model derivation. For example, only samples corresponding to significant residual activity are used in model derivation, as shown in Figure 17f, where the darker fragmented area within the blocks illustrates significant residual. In this context, significant can mean residual with a magnitude greater than zero, or a magnitude greater than some inferred, fixed, predetermined, or signaled threshold value. In another embodiment, model derivation can give more weight to samples defined by the area of significant residual activity.
[0242] According to one embodiment, samples not belonging to the area defined by significant residual activity by the previous embodiment can be predicted using block copy and the samples in the current channel as determined Petition 870250083300, dated 09 / 16 / 2025, page 74 / 298 65 / 79 by BV (possibly staggered) or vice versa.
[0243] According to one embodiment, the reference block and colocalized block can be divided into several regions based on connected component labeling of residual. For example, two blobs can be used to derive two models. Connected component labeling can be based, for example, on residual magnitude with blobs considered separate only if their corresponding residual values are significantly different (possibly indicating two different objects).
[0244] An apparatus according to one aspect comprises means for receiving an image block unit from a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; means for reconstructing samples of said luminance channel from the image block unit;means for determining a reference area for predicting target samples of at least one color channel of the image block unit, wherein said reference area comprises one or more reference samples in an actual block or colocalized in an actual color channel / frame or in a reference color channel / frame encoded using an intrablock copy (IBC) method; means for determining a block vector indicating a spatial distance from a target sample area to the reference area; and means for predicting said target samples of at least one color channel of the image block unit using the cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
[0245] Depending on one embodiment, the cross-component prediction method is either a cross-component linear model (CCLM) or a cross-component convolutional model (CCCM).
[0246] According to one embodiment, the apparatus comprises means for deriving parameters for the crossed component model using samples from a current-channel reference block and a channel-channel reference block Petition 870250083300, dated 09 / 16 / 2025, page 75 / 298 66 / 79 reference
[0247] According to one embodiment, the apparatus comprises means for inferring the block vector of a block colocalized in the reference channel; and means for scaling the block vector to correspond to a sampling density of the current channel.
[0248] According to one embodiment, the block vector inferred from the reference channel is at least partially misaligned with colocalized block coordinates.
[0249] According to one embodiment, the apparatus comprises means for identifying the nearest match to the reference area in the current channel using a template match-based search in the current channel.
[0250] According to one embodiment, the apparatus comprises means for calculating cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - one or more samples of the reference block in the current channel; - one or more samples of the reference block in the reference channel; - one or more samples of the block colocalized in the reference channel; - one or more neighborhood samples of the reference block in the current channel; - one or more neighborhood samples of the reference block in the reference channel; - one or more neighborhood samples of the current block in the current channel; - one or more neighborhood samples of the colocalized block in the reference channel.
[0251] According to one embodiment, the apparatus comprises means for calculating cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - Horizontal and / or vertical coordinates of said samples; - Directional gradient values of said samples.
[0252] According to one embodiment, the apparatus comprises means for Petition 870250083300, dated 09 / 16 / 2025, page 76 / 298 67 / 79 Use the block vector of the colocalized block inferred from the reference channel as an initial block vector for the current block.
[0253] According to one embodiment, the apparatus comprises means for generating a list of candidates for a block vector of encoded blocks using the intrablock copy (IBC) method in the reference channel in the colocalized region and / or in the neighborhood of the current block in the current channel.
[0254] According to one embodiment, the apparatus comprises means for obtaining the final cross-component prediction parameters of the current block using a two-stage derivation, wherein a first-stage derivation is performed using the reference block in the reference channel pointed to by the block vector and a second-stage derivation is performed in the vicinity of the current block.
[0255] As a further aspect, an apparatus is provided comprising: at least one processor and at least one memory, said at least one memory stored with code therein which, when executed by said at least one processor, causes the apparatus to perform at least: receiving an image block unit of a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; reconstructing samples of said luminance channel of the image block unit; determining a reference area for predicting target samples of at least one color channel of the image block unit, wherein said reference area comprises one or more reference samples in an actual block or colocalized in an actual color channel / frame or in a reference color channel / frame encoded using an intrablock copy (IBC) method;Determine a block vector indicating a spatial distance from a target sample area to the reference area; and predict said target samples from at least one color channel of the image block unit using the cross-component prediction model based on the reference samples in said reference area indicated by said block vector. Petition 870250083300, dated 09 / 16 / 2025, page 77 / 298 68 / 79
[0256] Depending on one embodiment, the cross-component prediction method is either a cross-component linear model (CCLM) or a cross-component convolutional model (CCCM).
[0257] According to one embodiment, the apparatus comprises code configured to make the apparatus derive the parameters for the cross-component model using samples from a current-channel reference block and a reference-channel reference block.
[0258] According to one embodiment, the apparatus comprises code configured to make the apparatus infer the block vector of a block colocalized in the reference channel; and scale the block vector to correspond to a sampling density of the current channel.
[0259] According to one embodiment, the block vector inferred from the reference channel is at least partially misaligned with colocalized block coordinates.
[0260] According to one embodiment, the device comprises code configured to enable the device to identify the nearest match to the reference area in the current channel using a template match-based search in the current channel.
[0261] According to one embodiment, the device comprises code configured to enable the device to calculate cross-component model parameters and apply the calculated model for prediction based on one or more of the following: - one or more samples of the reference block in the current channel; - one or more samples of the reference block in the reference channel; - one or more samples of the block colocalized in the reference channel; - one or more neighborhood samples of the reference block in the current channel; - one or more neighborhood samples of the reference block in the reference channel; - one or more neighborhood samples of the current block in the current channel; Petition 870250083300, dated 09 / 16 / 2025, page 78 / 298 69 / 79 - one or more neighborhood samples of the colocalized block in the reference channel.
[0262] According to one embodiment, the device comprises code configured to enable the device to calculate cross-component model parameters and apply the calculated model for prediction based on one or more of the following: - Horizontal and / or vertical coordinates of said samples; - Directional gradient values of said samples.
[0263] According to one embodiment, the device comprises code configured to make the device use the block vector of the colocalized block inferred from the reference channel as an initial block vector for the current block.
[0264] According to one embodiment, the device comprises code configured to make the device generate a list of candidates for a block vector of blocks encoded using the intrablock copy (IBC) method in the reference channel in the colocalized region and / or in the neighborhood of the current block in the current channel.
[0265] According to one embodiment, the apparatus comprises code configured to make the apparatus obtain the final cross-component prediction parameters of the current block using a two-step derivation, wherein a first-step derivation is performed using the reference block in the reference channel pointed to by the block vector and a second-step derivation is performed in the neighborhood of the current block.
[0266] Such devices may comprise, for example, the functional units revealed in any of Figures 1, 2, 4a and 4b to implement the modalities.
[0267] Such apparatus further comprises code, stored in said at least one memory which, when executed by said at least one processor, causes the apparatus to perform one or more of the modes disclosed herein.
[0268] Figure 18 is a graphical representation of a system of Petition 870250083300, dated 09 / 16 / 2025, page 79 / 298 70 / 79 exemplary multimedia communication within which various modalities can be implemented. A 1510 data source provides a source signal in an analog, uncompressed digital, or compressed digital format, or any combination of these formats. A 1520 encoder may include or be connected to pre-processing, such as data format conversion and / or filtering of the source signal. The 1520 encoder encodes the source signal into a coded media bitstream. It should be noted that a bitstream to be decoded can be received directly or indirectly from a remote device located on virtually any type of network. Additionally, the bitstream can be received from local hardware or software. The 1520 encoder may be capable of encoding more than one media type, such as audio and video, or more than one 1520 encoder may be required to encode different media types from the source signal.The 1520 encoder can also take synthetically produced input, such as graphics and text, or it may be capable of producing encoded bitstreams of synthetic media. Herein, only the processing of a single encoded media bitstream of one media type is considered for the sake of simplicity. It should be noted, however, that real-time broadcast services typically comprise multiple streams (typically at least one audio, video, and text captioning stream). It should also be noted that the system may include many encoders, but in the figure only one 1520 encoder is represented for the sake of simplicity without loss of generality. It should further be understood that, while the text and examples contained herein may specifically describe one encoding process, a person skilled in the art would understand that the same concepts and principles also apply to the corresponding decoding process and vice versa.
[0269] The encoded media bitstream can be transferred to a 1530 storage. The 1530 storage can comprise any type of mass storage for storing the encoded media bitstream. The format of the encoded media bitstream in the 1530 storage can be Petition 870250083300, dated 09 / 16 / 2025, page 80 / 298 71 / 79 an elementary self-contained bitstream format, or one or more encoded media bitstreams may be encapsulated in a container file, or the encoded media bitstream may be encapsulated in a segment format suitable for DASH (or a similar streaming system) and stored as a sequence of Segments. If one or more media bitstreams are encapsulated in a container file, a file generator (not shown in the figure) may be used to store one more media bitstream in the file and create file format metadata, which may also be stored in the file. The 1520 encoder or 1530 storage may comprise the file generator, or the file generator is operationally connected to the 1520 encoder or 1530 storage. Some systems operate “live,” that is, they omit storage and transfer the encoded media bitstream from the 1520 encoder directly to the 1540 sender.The encoded media bitstream can then be transferred to the 1540 sender, also known as the server, as needed. The format used in the transmission can be an independent elementary bitstream format, a packet stream format, a segment format suitable for DASH (or a similar streaming system), or one or more encoded media bitstreams can be encapsulated in a container file. The 1520 encoder, the 1530 storage, and the 1540 server can reside on the same physical device or can be included on separate devices.The 1520 encoder and the 1540 server can operate with live content in real time; in such a case, the encoded media bitstream is typically not stored permanently, but rather stored temporarily for short periods of time in the 1520 content encoder and / or the 1540 server to smooth out variations in processing delay, transfer delay, and encoded media bitrate.
[0270] Server 1540 sends the encoded media bitstream using a communication protocol stack. The stack may include, but is not limited to, one or more of the following: Real-Time Transport Protocol (RTP), Petition 870250083300, dated 09 / 16 / 2025, page 81 / 298 72 / 79 User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Transmission Control Protocol (TCP), and Internet Protocol (IP). When the communication protocol stack is packet-oriented, the 1540 server encapsulates the encoded media bitstream into packets. For example, when RTP is used, the 1540 server encapsulates the encoded media bitstream into RTP packets according to an RTP payload format. Typically, each media type has a dedicated RTP payload format. It should be noted again that a system may contain more than one 1540 server, but for simplicity, the following description considers only one 1540 server.
[0271] If the media content is encapsulated in a container file for storage 1530 or for data input at the sender 1540, the sender 1540 may understand or be operationally attached to a sending file parser (not shown in the figure). In particular, if the container file is not transmitted as such, but at least one of the encoded media bitstreams contained within is encapsulated for transport by a communication protocol, a sending file parser locates appropriate portions of the encoded media bitstream to be transmitted by the communication protocol. The sending file parser may also help create the correct format for the communication protocol, such as packet headers and payloads. The multimedia container file may contain encapsulation instructions, such as hint trails in ISOBMFF, for encapsulation of at least one of the media bitstreams contained in the communication protocol.
[0272] The 1540 server may or may not be connected to a 1550 gateway via a communication network, which may be, for example, a combination of a CDN, the Internet and / or one or more access networks. The gateway may also or alternatively be called an intermediate box. For DASH, the gateway may be an edge server (of a CDN) or a web proxy. It should be noted that the system may generally comprise any Petition 870250083300, dated 09 / 16 / 2025, page 82 / 298 73 / 79 number of gateways or similar, but for simplicity, the following description considers only a 1550 gateway. The 1550 gateway can perform different types of functions, such as translating a packet stream according to one communication protocol stack to another communication protocol stack, embedding and branching data streams, and manipulating the data stream according to the capabilities of the downlink and / or receiver, such as controlling the bit rate of the forwarded stream according to the prevailing conditions of the downlink network. The 1550 gateway can be a server entity in several modes.
[0273] The system includes one or more 1560 receivers, typically capable of receiving, demodulating, and decapsulating the transmitted signal into an encoded media bitstream. The encoded media bitstream can be transferred to a 1570 recording storage. The 1570 recording storage can comprise any type of mass storage for storing the encoded media bitstream. The 1570 recording storage can alternatively or additionally comprise computing memory, such as random access memory. The format of the encoded media bitstream in the 1570 recording storage can be an elementary self-contained bitstream format, or one or more encoded media bitstreams can be encapsulated in a container file.If there are multiple encoded media bitstreams, such as an audio stream and a video stream, associated with each other, a container file is typically used, and the 1560 receiver comprises or is attached to a container file generator that produces a container file from the input streams. Some systems operate “live,” that is, they omit the 1570 recording storage and transfer the encoded media bitstream from the 1560 receiver directly to the 1580 decoder. In some systems, only the most recent portion of the recorded stream, for example, the most recent 10-minute segment of the recorded stream, is retained in the 1570 recording storage, while any previously recorded data is discarded from the 1570 recording storage. Petition 870250083300, dated 09 / 16 / 2025, page 83 / 298 74 / 79
[0274] The encoded media bitstream can be transferred from the 1570 recording storage to the 1580 decoder. If there are many encoded media bitstreams, such as an audio stream and a video stream, associated with each other and encapsulated in a container file, or if a single media bitstream is encapsulated in a container file, for example, to facilitate access, a file parser (not shown in the figure) will be used to decapsulate each encoded media bitstream from the container file. The 1570 recording storage or a 1580 decoder may comprise the file parser, or the file parser is attached to the 1570 recording storage or the 1580 decoder. It should also be noted that the system may include many decoders, but here only one 1570 decoder is discussed to simplify the description without lack of generality.
[0275] The encoded media bitstream can be further processed by a 1570 decoder, whose output is one or more uncompressed media streams. Finally, a 1590 renderer can play back the uncompressed media streams with a loudspeaker or a monitor, for example. The 1560 receiver, the 1570 recording storage, the 1570 decoder, and the 1590 renderer can reside in the same physical device or can be included in separate devices.
[0276] A 1540 sender and / or a 1550 gateway can be configured to perform switching between different representations, for example, to switch between different viewing windows of 360-degree video content, view switching, bitrate adaptation and / or fast initialization, and / or a 1540 sender and / or a 1550 gateway can be configured to select the transmitted representation(s). Switching between different representations can occur for various reasons, such as to respond to requests from the 1560 receiver or prevailing conditions, such as throughput, of the network over which the bitstream is transmitted. In other words, the 1560 receiver can initiate switching between representations. A request Petition 870250083300, dated 09 / 16 / 2025, page 84 / 298 75 / 79 from the receiver could be, for example, a request for a Segment or Subsegment of a different representation than the previous one, a request for a change in scalability layers and / or transmitted sublayers, or a change to a rendering device with different capabilities compared to the previous one. A request for a Segment could be an HTTP GET request. A request for a Subsegment could be an HTTP GET request with a byte range. Additionally or alternatively, bitrate adjustment or bitrate adaptation can be used, for example, to provide so-called fast startup in streaming services, where the bitrate of the transmitted stream is lower than the bitrate of the channel after the start, or random access to the streaming to start playback immediately and achieve a buffer occupancy level that tolerates occasional packet delays and / or retransmissions.Bit rate adaptation can involve multiple representations or layer-up and layer-down switching operations, occurring in various orders.
[0277] A 1580 decoder can be configured to perform switching between different representations, for example, to switch between different 360-degree video content viewing windows, view switching, bitrate adaptation and / or fast initialization, and / or a 1580 decoder can be configured to select the transmitted representation(s). Switching between different representations can occur for various reasons, such as to obtain faster decoding operation or to adapt the transmitted bitstream, for example, in terms of bitrate, to the prevailing conditions, such as the data transfer rate, of the network over which the bitstream is transmitted. Faster decoding operation may be required, for example, if the device, including the 1580 decoder, is multitasking and uses computing resources for purposes other than decoding the video bitstream.In another example, a faster decoding operation may be necessary when the content is played back at a faster rate than the speed. Petition 870250083300, dated 09 / 16 / 2025, page 85 / 298 76 / 79 normal playback, for example, two or three times faster than the conventional real-time playback rate.
[0278] Above, some modes have been described with reference to and / or using HEVC and / or VVC terminology. It is important to understand that the modes can be implemented similarly with any video encoder and / or video decoder.
[0279] Above, where example modes have been described with reference to an encoder, it is necessary to understand that the resulting bitstream and the decoder may have corresponding elements. Similarly, when example modes have been described with reference to a decoder, it is necessary to understand that the encoder may have a computer structure and / or program to generate the bitstream to be decoded by the decoder. For example, some modes have been described relating to the generation of a prediction block as part of the encoding. Modes can be performed similarly by generating a prediction block as part of the decoding, with the difference that the encoding parameters, such as horizontal and vertical deviation, are decoded from the bitstream and not determined by the encoder.
[0280] The embodiments of the invention described above describe the codec in terms of separate encoder and decoder apparatus in order to aid in understanding the processes involved. However, it would be appreciated if the apparatus, structures and operations could be implemented as a single encoder-decoder apparatus / structure / operation. Additionally, it is possible that the encoder and decoder share some or all of the common elements.
[0281] Although the above examples describe embodiments of the invention operating within a codec inside an electronic device, it would be appreciated that the invention, as defined in the claims, can be implemented as part of any video codec. Thus, for example, embodiments of the invention can be implemented in a video codec that Petition 870250083300, dated 09 / 16 / 2025, page 86 / 298 77 / 79 can implement video encoding via fixed or wired communication paths.
[0282] Thus, the user equipment may comprise a video codec such as those described in the embodiments of the invention above. It should be noted that the term user equipment is intended to encompass any suitable type of wireless user equipment, such as mobile phones, portable data processing devices, or portable browsers.
[0283] Additionally, elements of a public terrestrial mobile network (PLMN) may also comprise video codecs, as described above.
[0284] In general, the various embodiments of the invention can be implemented in special-purpose hardware or circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although the invention is not limited thereto. Although various aspects of the invention may be illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it is well understood that those blocks, apparatus, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controller, or other computing devices, or some combination thereof.
[0285] The embodiments of this invention can be implemented by computer software executable by a mobile device data processor, as in the processor entity, or by hardware, or by a combination of software and hardware. Additionally, in this regard, it should be noted that any blocks of the logic flow, as in the Figures, can represent program steps, or logic circuits, blocks, and functions. Petition 870250083300, dated 09 / 16 / 2025, page 87 / 298 78 / 79 interconnected, or a combination of program steps and logic circuits, blocks and functions. Software can be stored on such physical media as memory chips or memory blocks implemented in the processor, magnetic media such as hard drives or floppy disks, and optical media such as, for example, DVDs and data variants thereof, CDs.
[0286] Memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Data processors may be of any type suitable to the local technical environment and may include one or more general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), and multi-core processor architecture-based processors, as non-limiting examples.
[0287] Modalities of the inventions can be practiced in various components, such as integrated circuit modules. The design of integrated circuits is, to a large extent, a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed onto a semiconductor substrate.
[0288] Programs, such as those provided by Synopsys, Inc. of Mountain View, California, and Cadence Design of San Jose, California, automatically route conductors and locate components on a semiconductor chip using well-established design rules as well as libraries of pre-stored design modules. Once the design of a semiconductor circuit has been completed, the resulting design, in a standardized electronic format (e.g., Opus, GDSII, or similar), can be transmitted to a semiconductor fabrication facility or fab for manufacturing. Petition 870250083300, dated 09 / 16 / 2025, page 88 / 298 79 / 79
[0289] The preceding description has provided, by means of non-limiting illustrative examples, a complete and informative description of the illustrative embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant art, in view of the preceding description, when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such modifications and similar modifications of the teachings of this invention will still fall within the scope of this invention. Petition 870250083300, dated 09 / 16 / 2025, page 89 / 298
Claims
1 / 4 CLAIMS 1. Apparatus characterized by comprising means for receiving an image block unit from a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; means for reconstructing samples of said luminance channel from the image block unit; means for determining a reference area for predicting target samples of at least one color channel of the image block unit, wherein said reference area comprises one or more reference samples in a current block or colocalized in a current color channel or in a current frame or in a reference color channel or in a reference frame encoded using an intrablock copy (IBC) method; means for determining a block vector indicating a spatial distance from a target sample area to the reference area;and means for predicting said target samples of at least one color channel of the image block unit using a cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
2. Apparatus, according to claim 1, characterized in that the cross-component prediction method is a linear cross-component model (CCLM) or a convolutional cross-component model (CCCM).
3. Apparatus, according to claim 1 or 2, characterized by comprising means for deriving the parameters for the cross-component model using samples from a reference block in the current channel and a reference block in the reference channel.
4. Apparatus, according to any of the preceding claims, characterized by comprising means for inferring the block vector of the colocalized block in the reference channel; and means for scaling the block vector to correspond to a sampling density of the current channel.
5. Apparatus, according to claim 4, characterized in that the block vector inferred from the reference channel is at least partially misaligned with colocalized block coordinates.
6. Apparatus, according to any of the preceding claims, characterized by comprising means for identifying the nearest match to the reference area in the current channel using a template match-based search in the current channel.
7. Apparatus, according to any of the preceding claims, characterized by comprising means for calculating cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - one or more samples of the reference block in the current channel; - one or more samples of the reference block in the reference channel; - one or more samples of the colocalized block in the reference channel; - one or more neighborhood samples of the reference block in the current channel; - one or more neighborhood samples of the reference block in the reference channel; - one or more neighborhood samples of the current block in the current channel; - one or more neighborhood samples of the colocalized block in the reference channel.
8. Apparatus, according to claim 7, characterized by Petition 870250083300, dated 09 / 16 / 2025, page 91 / 298 3 / 4 comprising means for calculating cross-component model parameters and applying the calculated model for prediction based on one or more of the following: - Horizontal and / or vertical coordinates of said samples; - Directional gradient values of said samples.
9. Apparatus, according to any of the preceding claims, characterized by comprising means for using the block vector of the colocalized block inferred from the reference channel as an initial block vector for the current block.
10. Apparatus, according to any of the preceding claims, characterized by comprising means for generating a list of candidates for a block vector of encoded blocks using the intrablock copy (IBC) method in the reference channel in the colocalized block and / or in the neighborhood of the current block in the current channel.
11. Apparatus, according to any of the preceding claims, characterized by comprising means for obtaining the final cross-component prediction parameters of the current block using a two-stage derivation, wherein a first-stage derivation is performed using the reference block in the reference channel pointed to by the block vector and a second-stage derivation is performed in the vicinity of the current block.
12. A method characterized by comprising receiving an image block unit from a frame, the image block unit comprising samples in color channels, wherein the color channels comprise at least one chrominance channel and one luminance channel; reconstructing samples of said luminance channels from the image block unit; determining a reference area for predicting target samples of Petition 870250083300, dated 09 / 16 / 2025, p.92 / 298 4 / 4 at least one color channel of the image block unit, wherein said reference area comprises one or more reference samples in a current block or colocalized in a current color channel or in a current frame or in a reference color channel or in a reference frame encoded using an intrablock copy (IBC) method; determine a block vector indicating a spatial distance from a target sample area to the reference area; and predict said target samples from at least one color channel of the image block unit using a cross-component prediction model based on the reference samples in said reference area indicated by said block vector.
13. Method according to claim 12, characterized in that the cross-component prediction method is either a cross-component linear model (CCLM) or a cross-component convolutional model (CCCM).
14. Method, according to claim 12 or 13, characterized by comprising deriving the parameters for the cross-component model using samples from a reference block in the current channel and a reference block in the reference channel.
15. A method, according to any one of claims 12-14, characterized by comprising inferring the block vector of the colocalized block in the reference channel; and scaling the block vector to correspond to a sampling density of the current channel. Petition 870250083300, dated 09 / 16 / 2025, pp. 93 / 298