Fusion of video prediction modes
Patent Information
- Application Number
- JP2024536106
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-23
- Filing Date
- 2023-01-04
- Publication Date
- 2026-01-28
AI Technical Summary
Existing video coding standards face challenges in achieving optimal compression efficiency for chroma components, leading to suboptimal coding performance due to the limited number of intra prediction modes and inter-component redundancy.
The proposed method fuses multiple chroma intra prediction modes using a weighted sum to generate a more accurate predicted chroma sample, incorporating linear models and decoder-side derived modes to improve coding efficiency.
Enhances the coding efficiency of chroma components by improving prediction accuracy, thereby reducing bit rates and enhancing overall video quality.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 63 / 296,533, filed January 5, 2022, and to U.S. Publication No. 18 / 146,172, filed December 23, 2022, each of which is incorporated by reference in its entirety into this specification.
[0002] Technical Field FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to video processing, and more particularly, to methods and systems for fusing chroma intra prediction modes. [Background technology]
[0003] background
[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, that specify specific video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of the new video coding standards becomes higher. Summary of the Invention
[0004] overview
[0004] Embodiments of the present disclosure are directed to fusing chroma intra prediction modes. In some embodiments, an exemplary method for processing video data includes generating a plurality of predicted chroma samples associated with a pixel by using a plurality of chroma intra prediction modes respectively, and determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples.
[0005]
[0005] An embodiment of the present disclosure provides an apparatus for processing video data, the system including a memory storing a set of instructions and one or more processors, the one or more processors being configured to execute the set of instructions to cause the apparatus to generate a plurality of predicted chroma samples associated with a pixel by respectively using a plurality of chroma intra prediction modes, and determine a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples.
[0006]
[0006] An embodiment of the present disclosure further provides a non-transitory computer-readable medium for storing a video bitstream for processing based on a method, the method including generating a plurality of predicted chroma samples associated with a pixel by respectively using a plurality of chroma intra prediction modes, and determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples.
[0007]
[0007] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for processing video data, the method including generating a plurality of predicted chroma samples associated with a pixel by respectively using a plurality of chroma intra prediction modes, and determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples.
[0008]
[0008] An embodiment of the present disclosure further provides a computer program product including computer program instructions that enable a computer to perform a method including generating a plurality of predicted chroma samples associated with a pixel by respectively using a plurality of chroma intra prediction modes, and determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples.
[0009]
[0009] An embodiment of the present disclosure further provides a computer program enabling a computer to perform a method including generating a plurality of predicted chroma samples associated with a pixel by respectively using a plurality of chroma intra prediction modes, and determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples.
[0010] BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief description of the drawings]
[0011] [Figure 1]
[0011] FIG. 1 is a schematic diagram illustrating an example structure of a video sequence according to some embodiments of the present disclosure. [Figure 2A]
[0012] 1 is a schematic diagram illustrating an example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 2B]
[0013] 1 is a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3A]
[0014] 1 is a schematic diagram illustrating an example decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3B]
[0015] 1 is a schematic diagram illustrating another example decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 4]
[0016] 1 is a block diagram of an example device for encoding or decoding video in accordance with some embodiments of the present disclosure. [Diagram 5]
[0017] 1 illustrates 67 intra-prediction modes according to some embodiments of this disclosure. [Figure 6]
[0018] 1 illustrates that a chroma coded block corresponds to multiple luma coded blocks in an I slice, according to some embodiments of this disclosure. [Figure 7]
[0019] 1 is a flowchart of a method for processing video data according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Description of the embodiments
[0020] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numerals in different figures represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of devices and methods consistent with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with a term and / or definition incorporated by reference, the term and definition provided herein shall control.
[0013]
[0021] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265, but using half the bandwidth.
[0014]
[0022] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing technology beyond HEVC using the Joint Search Model (JEM) reference software. Because coding technology has been incorporated into JEM, JEM has achieved much higher coding performance than HEVC.
[0015]
[0023] The VVC standard is a recent development and continues to incorporate more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in recent video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0016]
[0024] A video is a set of still pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display such pictures in chronological order. Furthermore, in some applications, a video capture device can transmit captured videos in real time to a video playback device (e.g., a computer with a monitor) for surveillance, conferencing, live broadcast, etc.
[0017]
[0025] To reduce the storage space and transmission bandwidth required by such applications, the video may be compressed before storage and transmission, and decompressed before display. This compression and decompression may be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. A module for compression is generally called an "encoder," and a module for decompression is generally called a "decoder." The encoder and decoder may be collectively referred to as a "codec." The encoder and decoder may be implemented as various suitable hardware, software, or combinations thereof. For example, a hardware implementation of the encoder and decoder may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. A software implementation of the encoder and decoder may include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be called a "transcoder."
[0018]
[0026] A video coding process can identify and keep useful information that can be used to reconstruct a picture and ignore information that is not important for the reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, then such a coding process can be called "lossy". Otherwise, such a coding process can be called "lossless". Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.
[0019]
[0027] Useful information of the picture being coded (called the "current picture") includes changes with respect to a reference picture (e.g. a previously coded and reconstructed picture). Such changes may include pixel position changes, luminance changes or color changes, of which position changes are the most important. Position changes of pixels representing an object may reflect the object's motion between the reference picture and the current picture.
[0020]
[0028] A picture that is coded without reference to another picture (i.e., the picture is its own reference picture) is called an "I-picture". If some or all of the blocks in a picture (e.g., blocks that generally refer to a portion of a video picture) are predicted using intra-prediction or inter-prediction with one reference picture (e.g., uni-prediction), the picture is called a "P-picture". If at least one block in the picture is predicted with two reference pictures (e.g., bi-prediction), the picture is called a "B-picture".
[0021]
[0029] 1 illustrates an example structure of a video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface for receiving video from a video content provider (e.g., a video broadcast transceiver).
[0022]
[0030] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture whose reference picture is picture 102 itself. Picture 104 is a P-picture whose reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture whose reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately preceding or following that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of the reference pictures to the examples shown in FIG.
[0023]
[0031] Typically, a video codec does not encode or decode an entire picture at once because such a task is computationally complex. Rather, a video codec may divide a picture into elementary segments and encode or decode a picture segment by segment. In this disclosure, such elementary segments are referred to as basic processing units ("BPUs"). For example, structure 110 in FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, a picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have a variable size within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of a basic processing unit may be selected for a picture based on a balance between efficiency of coding and the level of detail one wishes to preserve within the basic processing unit.
[0024]
[0032] A basic processing unit may be a logical unit that may include various types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit, where the luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated on each of its luma and chroma components.
[0025]
[0033] Video coding has multiple operation stages, examples of which are shown in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process, and therefore may be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing sub-units are also logical units that may include various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated on each of its luma and chroma components. It should be noted that such division can be performed on further levels, depending on the processing needs. It should also be noted that different stages can use different schemes to divide the basic processing unit.
[0026]
[0034] For example, in a mode decision stage (one example of which is shown in FIG. 2B), an encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder may split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.
[0027]
[0035] As another example, in the prediction stage (one example of which is shown in FIG. 2A-FIG. 2B), the encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.
[0028]
[0036] As another example, in the transform stage (one example of which is shown in FIG. 2A-2B), the encoder can perform a transform operation on the residual elementary processing sub-unit (e.g., CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder can further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0029]
[0037] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.
[0030]
[0038] In some implementations, to provide video encoding and decoding with parallel processing and error resilience capabilities, a picture can be divided into regions for processing, so that for a region of a picture, the encoding or decoding process can be independent of information of any other region of the picture. In other words, each region of a picture can be processed independently. In this way, a codec can process different regions of a picture in parallel, thus increasing the efficiency of coding. Furthermore, if data of a region is corrupted in processing or lost in network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience capabilities. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that various pictures of the video sequence 100 can have different partitioning schemes for dividing the picture into regions.
[0031]
[0039] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure is not limited to such embodiments.
[0032]
[0040] FIG. 2A illustrates a schematic diagram of an example of an encoding process 200A consistent with embodiments of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As illustrated in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 may be divided by the encoder into elementary processing units, elementary processing sub-units, or regions for processing. In some embodiments, the encoder may perform the process 200A at the level of the elementary processing units for each original picture of the video sequence 202. For example, the encoder may perform the process 200A in an iterative manner, and the encoder may encode a elementary processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform process 200A in parallel for each original picture region of video sequence 202 (eg, regions 114-118).
[0033]
[0041] In FIG. 2A , an encoder may feed a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, the encoder may feed quantized transform coefficients 216, after quantization stage 214, to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224 that is used in the prediction stage 204 of the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0034]
[0042] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0035]
[0043] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any manner of inputting data.
[0036]
[0044] In the prediction step 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224 and perform a prediction operation to generate predicted data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting the predicted data 206, which can be used to reconstruct the original BPU from the predicted data 206 and the prediction reference 224 as the predicted BPU 208.
[0037]
[0045] Ideally, the predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such differences, the encoder may generate the predicted BPU 208 and then subtract it from the original BPU to generate the residual BPU 210. For example, the encoder may subtract the values (e.g., grayscale values or RGB values) of pixels of the predicted BPU 208 from the values of corresponding pixels of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and the predicted BPU 208, each pixel of the residual BPU 210 may have a residual value. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant loss of quality.
[0038]
[0046] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," where each basis pattern is associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reproduced from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.
[0039]
[0047] Different transform algorithms may use different basis patterns. Different transform algorithms may be used in transform stage 212, for example, discrete cosine transform, discrete sine transform, etc. The transform in transform stage 212 is invertible. That is, the encoder may restore residual BPU 210 by the inverse operation of the transform (referred to as "inverse transform"). For example, to restore pixels of residual BPU 210, the inverse transform may be to multiply the values of corresponding pixels of the basis pattern by the associated respective coefficients and add the products to result in a weighted sum. In a video coding standard, both the encoder and the decoder may use the same transform algorithm (and thus the same basis pattern). Thus, the encoder may record only the transform coefficients, and the decoder may reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they may be used to reconstruct residual BPU 210 without significant loss of quality. Therefore, the residual BPU 210 is further compressed.
[0040]
[0048] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore the information of high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization scale factor") and rounding the quotient to its nearest neighbor. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and the transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, which further compresses the transform coefficients. The quantization process is also invertible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "inverse quantization").
[0041]
[0049] The quantization stage 214 may be lossy because the encoder ignores the remainder of such a division in a rounding operation. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different levels of information loss, the encoder may use different values of the quantization parameter or any other parameter of the quantization process.
[0042]
[0050] In the binary coding stage 226, the encoder may use a binary coding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm, to encode the prediction data 206 and the quantized transform coefficients 216. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transformation in the transformation stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. The encoder may use the output data of the binary coding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0043]
[0051] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may apply the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0044]
[0052] It should be noted that other variations of the process 200A can be used to encode the video sequence 202. In some embodiments, an encoder can perform the stages of the process 200A in a different order. In some embodiments, one or more stages of the process 200A can be combined into a single stage. In some embodiments, a single stage of the process 200A can be separated into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, the process 200A can include additional stages. In some embodiments, the process 200A can omit one or more stages in FIG. 2A.
[0045]
[0053] 2B shows a schematic diagram of another example encoding process 200B consistent with an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0046]
[0054] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels of one or more neighboring BPUs already coded in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions of one or more pictures already coded to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0047]
[0055] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that are encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or linear interpolation, polynomial extrapolation, polynomial interpolation, etc. In some embodiments, the encoder may perform extrapolation at a pixel level, such as by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction specified within the video coding standard used. In intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0048]
[0056] As another example, in the temporal prediction stage 2044, the encoder may perform inter prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform an operation of "motion estimation" to look for a matching region for a certain range (called a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine the region as a match region. The match region may have different (e.g., smaller, equal, larger, or differently shaped) dimensions than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the match region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder can look for a match region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the match region of each matching reference picture.
[0049]
[0057] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, scaling, etc. In inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the match region, a motion vector associated with the match region, a number of reference pictures, weights associated with the reference pictures, etc.
[0050]
[0058] To generate the predicted BPU 208, the encoder may perform an operation of "motion compensation." Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move the matching regions of the reference picture according to the motion vectors, so that the encoder can predict the original BPU of the current picture. If multiple reference pictures are used (e.g., as in picture 106 of FIG. 1), the encoder may move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values of the matching regions. In some embodiments, if the encoder assigns weights to the pixel values of the matching regions of the respective matching reference pictures, the encoder may add a weighted sum of the pixel values of the moved matching regions.
[0051]
[0059] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter predicted picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter predicted picture in which reference pictures (e.g., pictures 104 and 108) are in both temporal directions relative to picture 104.
[0052]
[0060] Continuing to refer to the forward path of the process 200B, after the spatial prediction step 2042 and the temporal prediction step 2044, in a mode decision step 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of the process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the prediction mode selected, the encoder may generate a corresponding prediction BPU 208 and prediction data 206.
[0053]
[0061] In the reconstruction path of the process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU being encoded and reconstructed in the current picture), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). The encoder can feed the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused during the coding of the prediction reference 224. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset (SAO), adaptive loop filter (ALF), etc. The loop filtered reference picture may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., for use as an inter-predicted reference picture for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in a binary coding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.
[0054]
[0062] FIG. 3A shows a schematic diagram of an example of a decoding process 300A consistent with an embodiment of the present disclosure. The process 300A may be a decompression process corresponding to the compression process 200A of FIG. 2A. In some embodiments, the process 300A may be similar to the reconstruction path of the process 200A. A decoder may decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression process (e.g., the quantization stage 214 of FIGS. 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to the processes 200A and 200B of FIGS. 2A-2B, the decoder may perform the process 300A at the level of a basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, and the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region (e.g., regions 114-118) of each picture encoded in video bitstream 228.
[0055]
[0063] In FIG. 3A, the decoder may feed a portion of the video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a "coded BPU") to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.
[0056]
[0064] The decoder may iteratively perform the process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0057]
[0065] In the binary decoding stage 302, the decoder may inverse the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, the prediction mode, parameters of the prediction operation, the type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder may depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.
[0058]
[0066] 3B shows a schematic diagram of another example of a decoding process 300B consistent with an embodiment of the present disclosure. The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., H.26x series). Compared to the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0059]
[0067] In the process 300B, for a coded elementary processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPU, parameters of extrapolation, the direction of the neighboring BPU relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors respectively associated with the matching regions, etc.
[0060]
[0068] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. As described in FIG. 3A, the decoder may add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.
[0061]
[0069] In the process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of the process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture to which all BPUs are decoded), the decoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., for use as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, the prediction data includes loop filter parameters if the prediction mode indicator of the prediction data 206 indicates that inter prediction was used to encode the current BPU.
[0062]
[0070] FIG. 4 is a block diagram of an example of a device 400 for encoding or decoding video consistent with an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 may become a dedicated machine for encoding or decoding video. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any combination of any number of central processing units ("CPUs"), graphics processing units ("GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), application specific integrated circuits (ASICs), and the like. In some embodiments, processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in FIG. 4, processor 402 may include multiple processors including processor 402a, processor 402b, and processor 402n.
[0063]
[0071] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in a process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or process the data for processing. The memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, and the like. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical entity.
[0064]
[0072] Bus 410 , such as an internal bus (eg, a CPU memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), etc., may be a communication device that transfers data between components within device 400 .
[0065]
[0073] For ease of explanation and without ambiguity, the present disclosure refers to the processor 402 and other data processing circuitry collectively as the "data processing circuitry." The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module, or may be fully or partially combined within any other component of the device 400.
[0066]
[0074] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0067]
[0075] In some embodiments, device 400 may further optionally include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), and the like.
[0068]
[0076] It should be noted that a video codec (e.g., a codec that executes processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all of the stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).
[0069]
[0077] Next, the luma intra prediction mode and chroma intra prediction mode used in VVC will be described. In VVC, the coding tree scheme supports separate block tree structures for each of the luma and chroma components. A CTU may include three CTBs (coding tree blocks), namely one luma CTB (Y) and two chroma CTBs (Cb and Cr). For P slices and B slices, the luma CTB and chroma CTB in one CTU share the same coding tree structure. Meanwhile, for I slices, the luma CTB and chroma CTB may have separate block tree structures. When separate block tree structures are applied, the luma CTB is divided into CUs using one coding tree structure, and the chroma CTB is divided into chroma CUs using another coding tree structure. This means that a CU in an I slice can contain either a coded block of the luma component or a coded block of the two chroma components, and a CU in a P or B slice always contains coded blocks of all three color components unless the image is monochrome (meaning there is no chroma component).
[0070]
[0078] Consistent with the disclosed embodiments, VVC supports the following intra prediction modes for the luma component: planar mode, DC mode, angular intra prediction mode, multi-reference line (MRL) prediction mode, intra sub-partition (ISP) mode, and matrix-based intra prediction (MIP) mode. These intra prediction modes are described in detail below.
[0071]
[0079] Angular intra prediction is a directional intra prediction method supported by HEVC and VVC. To capture arbitrary edge directions that appear in natural video, the number of angular intra prediction modes in VVC is extended from 33 used in HEVC to 65. In Figure 5, the new angular intra prediction modes not found in HEVC are depicted with dotted arrows.
[0072]
[0080] Similar to HEVC, VVC also supports two non-angular intra prediction modes: DC mode and planar mode. Thus, Figure 5 shows a total of 67 intra prediction modes. DC intra prediction mode uses the average sample value of reference samples for a block for prediction generation. VVC only uses reference samples along the long side of a rectangular block to calculate the average value, while for a square block, reference samples from both the left and top sides are used. In planar mode, the predicted sample value is obtained as a weighted average of four reference sample values. Here, the reference samples in the same row or column as the current sample and the reference samples in the bottom-left and top-right positions for the block are used.
[0073]
[0081] For MRL mode, in addition to the directly neighboring lines of adjacent samples, one of the two non-neighboring reference lines may contain the input for intra prediction in VVC.
[0074]
[0082] The ISP mode divides the luma intra prediction block into two or four sub-partitions vertically or horizontally depending on the block size. For each sub-partition, prediction and transform coding operations are performed separately, but the intra prediction mode is shared across all sub-partitions.
[0075]
[0083] MIP mode is a new intra prediction technique added to VVC. To predict samples of a block of width W and height H, MIP takes as input one line of H reconstructed adjacent boundary samples to the left of the block and one line of W reconstructed adjacent boundary samples above the block. The generation of the prediction signal is based on three steps: downsampling of the reference samples, matrix-vector multiplication, and upsampling of the result by linear interpolation.
[0076]
[0084] In the Enhanced Compression Model (ECM), several video compression techniques beyond VVC are considered. Two luma intra prediction modes are proposed: Decoder-side Intra Mode Derivation (DIMD) mode and Template-Based Intra Mode Derivation (TIMD) mode. When DIMD is applied, two intra prediction modes from 65 angle modes are derived from the reconstructed neighboring samples, and these two predictors are combined with the planar mode predictor by using weights derived from the gradient. When TIMD is applied, for each intra prediction mode in the list, the SATD between the template prediction sample and the reconstructed sample is calculated. The first two intra prediction modes with the smallest SATD are selected and fused with weights derived from the SATD.
[0077]
[0085] Next, chroma intra prediction modes are described. The intra prediction modes available for chroma components in VVC are three cross-component linear model (CCLM) modes including CCLM_LT, CCLM_L, and CCLM_T, a direct mode (DM), and four default intra prediction modes. In the following, these chroma intra prediction modes are described in detail.
[0078]
[0086] To reduce the redundancy of cross-components, three CCLM prediction modes are used in VVC, where the chroma components of a block may be predicted from the collocated reconstructed luma samples by a linear model whose parameters are derived from already reconstructed luma samples and chroma samples neighboring the block. The chroma samples may be predicted according to Equation 1.
number
number
[0079]
[0087] In VVC, three CCLM modes are specified: CCLM_LT mode, CCLM_L mode, and CCLM_T mode. These three modes differ from each other with respect to the location of the reconstructed neighboring samples used to derive the linear model parameters. Samples from the top boundary are involved in CCLM_T mode, and samples from the left boundary are involved in CCLM_L mode. In CCLM_LT mode, samples from both the top and left boundaries are used.
[0080]
[0088] In order to match the chroma sample positions of video sequences with 4:2:0 or 4:2:2 color formats, two kinds of downsampling filters can be applied to the luma samples, both of which have a downsampling ratio of 2 to 1 in the horizontal and vertical directions. Based on the flag information at the SPS level, a two-dimensional 6-tap (f1 in Equation 2) or 5-tap (f2 in Equation 2) filter is applied to the luma samples in the current block and its neighboring luma samples.
number
[0081]
[0089] The linear model parameters α and β are derived based on the reconstructed adjacent luma and chroma samples at both the encoder and decoder sides to avoid signaling overhead. The first adopted version of the CCLM mode used a linear minimum mean squared error (LMMSE) estimator for parameter derivation.
number
[0082]
[0090] However, in the final design, only four samples are included to reduce the computational complexity. For an M×N chroma block, the four samples used in CCLM_LT mode are those at M / 4 and 3M / 4 on the top boundary and N / 4 and 3N / 4 on the left boundary. In CCLM_T and CCLM_L modes, the top and left boundaries are expanded to a size of (M+N) samples, and the four samples used to derive the model parameters are at (M+N) / 8, 3(M+N) / 8, 5(M+N) / 8, and 7(M+N) / 8. After the four samples are selected, four comparison operations are used to determine the two smallest and two largest luma sample values among them. L max represents the average of the two largest luma sample values, and L min Let C denote the average of the two smallest luma sample values. max and C min Let denote the average of the corresponding chroma sample values. Then, the parameters of the linear model are obtained.
number
[0083]
[0091] ECM extends CCLM included in VVC by adding three Multi-Model LM (MMLM) modes: MMLM_LT, MMLM_L, and MMLM_T. In each MMLM mode, the reconstructed neighboring samples are classified into two classes using a threshold that is the average of the luma reconstructed neighboring samples. A linear model for each class is derived using the LMMSE method. To improve the prediction accuracy, ECM also includes two variants of the LM mode: Convolutional Cross-Component Model (CCCM) mode and Gradient Linear Model (GLM) mode. In CCCM mode, neighboring downsampled luma samples are also used to predict the current chroma sample, which extends the number of model parameters to seven. In GLM, the luma gradient is used to predict the current chroma sample.
[0084]
[0092] As described above, the valid intra prediction modes for the chroma components of VVC include the DM mode. When the DM mode is used, the intra prediction mode of the corresponding luma block determines the chroma intra mode as follows: If the corresponding luma block uses planar, DC, or angular mode, the same mode is used. If the corresponding luma block is coded using intra block copy (IBC) or palette mode, DC mode is used. If the corresponding luma block is coded using block DPCM (BDPCM) mode, either horizontal or vertical intra prediction mode is used depending on the BDPCM orientation. If the corresponding luma block uses MIP and the chroma color format is 4:4:4 and a single partition tree is applied, the same MIP mode is applied for the chroma block, otherwise the planar mode is applied. For B and P slices, the corresponding luma block represents the luma block at the same position as the current chroma block. For an I slice, a chroma coded block may correspond to multiple luma coded blocks since separate block partitioning structures are available for the luma and chroma components. For example, Figure 6 shows that a chroma coded block 610 has multiple co-located luma coded blocks 620 in an I slice. The corresponding luma block 625 is a luma coded block that includes a centrally located luma sample 630.
[0085]
[0093] As explained above, the chroma intra prediction modes in VVC also include four default intra prediction modes. If CCLM and DM modes are not used, the four default non-DM modes are given by the list {planar mode, vertical mode, horizontal mode, DC mode}. If a DM mode already belongs to the list (i.e., the DM mode is the same as one of the four modes), that mode in the list is replaced by the angular mode with mode index 66.
[0086]
[0094] In signaling a chroma intra mode, first, a flag cclm_mode_flag is signaled to indicate whether CCLM is applied or not. If cclm_mode_flag is signaled as true, an index cclm_mode_idx is signaled to indicate which of the three CCLM modes is applied. In the case of non-CCLM, a syntax intra_chroma_pred_mode is signaled to indicate which of the DM mode and the four default non-DM modes is applied. The binarization process of intra_chroma_pred_mode and the corresponding chroma intra prediction modes are shown in Table 1. If the first bin of intra_chroma_pred_mode is equal to 0, it means that the DM mode is applied. If the first bin of intra_chroma_pred_mode is equal to 1, it means that one of the four default non-DM modes is applied. Therefore, the first bin of intra_chroma_pred_mode can be considered as a DM flag indicating whether the DM mode is applied or not. When the DM mode is not used, an index ranging from 0 to 3 is binarized in 2-bit increments using a fixed-length codeword to determine which of the four non-DM modes is used.
[0087] [Table 1]
[0088]
[0095] Consistent with the disclosed embodiments, decoder-side derived chroma modes can also be used for chroma intra prediction, which can be derived based on texture gradients of collocated reconstructed luma samples or reconstructed chroma samples that are neighboring the current chroma block at both the encoder and decoder sides.
[0089]
[0096] Consistent with the disclosed embodiments, a new chroma intra prediction mode called CCLM-angular prediction mode may also be used for chroma intra prediction. If the chroma intra prediction of the current chroma block is in angular mode, a flag is signaled to indicate whether CCLM-angular mode is applied or not. If CCLM-angular mode is applied, both angular intra prediction mode and MMLM_LT mode prediction are performed, and the final prediction is set to be the average of angular intra prediction and MMLM_LT mode prediction.
[0090]
[0097] In the above embodiment, the number of intra prediction modes available for chroma blocks is less than the number of intra prediction modes available for luma blocks. Therefore, the intra prediction results for chroma blocks may not be as accurate as those for luma blocks. Chroma intra prediction has some prediction modes that use inter-component correlation and some prediction modes that use spatial correlation. Therefore, the accuracy of chroma intra prediction can be improved by fusing different chroma intra prediction modes.
[0091]
[0098] In this disclosure, in order to improve the coding efficiency of chroma intra prediction, it is proposed to merge some chroma intra prediction modes, i.e., the predicted sample values obtained by different chroma intra prediction modes are weighted to obtain a more accurate predicted sample.
[0092]
[0099] In some embodiments, n different chroma intra prediction modes are used to obtain n different predictors to perform intra prediction on the current chroma block, and the n predictors are weighted to obtain a final predictor. A predictor refers to a predicted sample value of a coded block. Different predictors can be fused according to Equation 6. pred(i,j)=w0*pred0(i,j)+w1*pred1(i,j)+···+w n-1 *pred n-1 (i,j) (Equation 6) where pred(i,j) represents the final predicted chroma sample in the current chroma block, and pred k (i,j) represents the predicted chroma sample of the current chroma block obtained by the k-th chroma intra prediction mode, and w k represents the weight of the k-th chrominance intra prediction mode (0≦k <n)。
[0093]
[0100] The n chroma intra prediction modes involved in the fusion may be any n modes of intra prediction modes available for the chroma block, where the value of n is greater than 1 and less than or equal to the number of available chroma intra prediction modes. For example, the n chroma intra prediction modes may be any n modes among CCLM_LT mode, CCLM_L mode, CCLM_T mode, MMLM_LT mode, MMLM_L mode, MMLM_T mode, DM mode, four default modes, and a decoder-side derived chroma mode. Also, the value of n may be any integer value between 2 and 12. The weights of the n intra prediction modes may be the same or different. The weights are positive values, and the sum of the n weights is 1.
[0094]
[0101] In some embodiments, the n chroma intra-prediction modes involved in the fusion include at least one LM mode and at least one non-LM mode. The LM mode may be one or more of a CCLM_LT mode, a CCLM_L mode, a CCLM_T mode, a MMLM_LT mode, a MMLM_L mode, a MMLM_T mode, a CCCM mode, and a GLM mode. The non-LM mode may be one or more of a DM mode, four default modes, and a decoder-side derived chroma mode.
[0095]
[0102] In some embodiments, the n chroma intra prediction modes involved in the fusion include an MMLM_LT mode and at least one non-LM mode.
[0096]
[0103] In some embodiments, the n chroma intra prediction modes involved in the merging include a decoder-side derived chroma mode and at least one LM mode.
[0097]
[0104] In some embodiments, the n chroma intra prediction modes involved in the fusion are different based on the type of slice of the current picture. In one example, the fusion chroma intra prediction mode is only used for I slices and is not used for B and P slices. In another example, the n chroma intra prediction modes involved in the fusion include an MMLM_LT mode and one non-LM mode. For I slices, the non-LM mode is one of a DM mode, four default modes, and a decoder-side derived chroma mode, and for B and P slices, the non-LM mode is a decoder-side derived chroma mode.
[0098]
[0105] Consistent with the disclosed embodiments, the weights of the n chroma intra prediction modes involved in the fusion may be determined according to different methods.
[0099]
[0106] In some embodiments, constant weights are used to weight different chroma intra prediction modes. In one example, equal weights are used. For example, if there are two chroma intra prediction modes used for fusion, the weight of each mode is equal to 1 / 2. In another example, unequal weights are used. In another example, two chroma intra prediction modes are used for fusion, one mode is a LM mode and the other mode is a non-LM mode. In this case, the weight of the LM mode is equal to 1 / 4 and the weight of the non-LM mode is equal to 3 / 4. Or the weight of the LM mode is equal to 3 / 4 and the weight of the non-LM mode is equal to 1 / 4.
[0100]
[0107] In some embodiments, multiple sets of weights can be selected for the fused chroma intra prediction mode, and an index is used to indicate which selected set of weights is used for the current coded block. Specifically, the encoder selects the best weight index through a rate-distortion optimization (RDO) decision and signals the index to the bitstream, and the decoder can obtain the index from the bitstream. In one example, there are two chroma intra prediction modes used for fusion, and three sets of weights can be selected: {w0=1 / 4, w1=3 / 4; w0=1 / 2, w1=1 / 2; w0=3 / 4, w1=1 / 4}. An index from 0 to 2 is signaled to indicate which set of weights is used for the current block.
[0101]
[0108] In some embodiments, the sets of weights may be selected based on the chroma prediction modes of the neighboring blocks. In one example, there are two chroma intra prediction modes used for blending, one mode is LM mode and the other mode is non-LM mode. In that case, the weights of each mode may be determined based on the chroma prediction modes of the upper neighboring chroma block and the left neighboring chroma block. Specifically, if both the upper neighboring and the left neighboring are LM mode coded, the weights of the LM mode and the non-LM mode are {3 / 4, 1 / 4}, if both the upper neighboring and the left neighboring are non-LM mode coded, the weights of the LM mode and the non-LM mode are {1 / 4, 3 / 4}, and if one of the upper neighboring and the left neighboring is LM mode coded and the other is non-LM mode coded, the weights of the LM mode and the non-LM mode are {1 / 2, 1 / 2}. In particular, if the chroma intra prediction mode of the neighbor is the proposed fusion mode, it is considered as a LM mode or a non-LM mode for the weight derivation of the current chroma block, or the weight of the neighboring chroma block is directly used for the current chroma block.
[0102]
[0109] In some embodiments, the set of weights may be selected based on the chroma prediction mode used for fusion and the sample position of the current chroma block. For example, there are two chroma intra prediction modes used for fusion, one mode is a non-angular mode and the other mode is an angular mode. The non-angular mode may be one or more of a planar mode, a DC mode, and a LM mode. In that case, the weights for each mode may be determined based on the angular mode and the sample position, as shown in the following example.
[0103]
[0110] In one example, a chroma block is first divided vertically (for horizontal mode) or horizontally (for vertical mode) into four equal-area regions, and different weights are used for each region. For example, four weight sets (w0,w1)=(6 / 8,2 / 8), (w0,w1)=(5 / 8,3 / 8), (w0,w1)=(3 / 8,5 / 8) and (w0,w1)=(2 / 8,6 / 8) are used for the four regions, respectively. w0 is used for angular modes and w1 is used for non-angular modes. Horizontal modes correspond to angular modes with mode index less than 34, and vertical modes correspond to angular modes with mode index greater than 34.
[0104]
[0111] In another example, the weight of each mode is related to the distance between the current sample and the reference sample based on the angular mode direction: w0, used for angular modes, is proportional to the distance, and w1, used for non-angular modes, is inversely proportional to the distance.
[0105]
[0112] In some embodiments, the above embodiments for selecting a set of weights can be combined. For example, there are two chroma intra prediction modes used for fusion, one mode is a LM mode and the other mode is a non-LM mode. If the non-LM mode is a non-angular mode, the weights can be selected based on the chroma prediction modes of the neighboring blocks, and if the non-LM mode is an angular mode, the weights can be selected based on the chroma prediction mode used for fusion and the sample position of the current chroma block.
[0106]
[0113] In some embodiments, the weights may differ based on the type of slice of the current picture, for example, for I slices, a method of selecting weights based on the intra-prediction modes of neighboring blocks is used, while for B slices and P slices, equal weights are used.
[0107]
[0114] This disclosure also provides a method for signaling the chroma intra-prediction mode involved in the fusion. When a chroma block is selected to use a proposed fusion mode for intra prediction, the chroma intra-prediction mode involved in the fusion can be obtained by explicit signaling, can be implicitly derived from information of the current block, or can be a combination of the two methods.
[0108]
[0115] In some embodiments, an explicit signaling method is used to determine which modes are involved in merging. For example, only two modes are allowed to be weighted, one mode is one of DM mode, four default modes, and decoder-side derived chroma mode, and the other mode is one of CCLM_LT mode and MMLM_LT mode. Then, after signaling DM mode, four default modes, or decoder-side derived chroma mode, a flag is signaled to indicate whether to merge with the other mode. If this flag is true, another flag is signaled to indicate whether to merge with CCLM_LT mode or MMLM_LT mode.
[0109]
[0116] In some embodiments, an implicit method is used to determine which modes are involved in merging. For example, only two modes are allowed to be weighted, one mode is one of DM mode, four default modes, and decoder-side derived chroma mode, and the other mode is one of CCLM_LT mode, CCLM_L mode, and CCLM_T mode. Then, after signaling DM mode, four default modes, or decoder-side derived chroma mode, a flag is signaled to indicate whether to merge with the other mode. If the flag is true, the other mode is determined based on the first mode. If the first mode is a non-angular mode, the second mode is determined to be CCLM_LT mode, if the first mode is a horizontal mode, the second mode is determined to be CCLM_L mode, and if the first mode is a vertical mode, the second mode is determined to be CCLM_T mode.
[0110]
[0117] In some embodiments, explicit signaling and implicit methods are combined to determine which modes are involved in merging. For example, only two modes are allowed to be weighted, one mode is one of DM mode, four default modes, and decoder-side derived chroma mode, and the other mode is one of CCLM_LT mode, CCLM_L mode, CCLM_T mode, MMLM_LT mode, MMLM_L mode, and MMLM_T mode. Then, after signaling DM mode, four default modes, or decoder-side derived chroma mode, a flag is signaled to indicate whether to merge with the other mode. If this flag is true, another flag is signaled to indicate whether to merge with CCLM mode or MMLM mode. Then, the other mode is further determined based on the first mode. If the first mode is a non-angle mode, the second mode is determined to be a CCLM_LT / MMLM_LT mode, if the first mode is a horizontal mode, the second mode is determined to be a CCLM_L / MMLM_L mode, and if the first mode is a vertical mode, the second mode is determined to be a CCLM_T / MMLM_T mode.
[0111]
[0118] The embodiments provided by the present disclosure can be freely combined. In one example, there are two chroma intra prediction modes used for fusion, one mode is a non-LM mode and the other mode is an MMLM_LT mode. The non-LM mode can be one of the DM mode, the four default modes, and the decoder-side derived chroma mode. And the weights of the non-LM mode and the MMLM_LT mode are determined based on the chroma prediction modes of the neighboring blocks. Specifically, if both the upper neighboring block and the left neighboring block are coded in LM mode, the weights of the MMLM_LT mode and the non-LM mode are {3 / 4, 1 / 4}, if both the upper neighboring block and the left neighboring block are coded in non-LM mode, the weights of the MMLM_LT mode and the non-LM mode are {1 / 4, 3 / 4}, if one of the upper neighboring block and the left neighboring block is coded in LM mode and the other is coded in non-LM mode, the weights of the MMLM_LT mode and the non-LM mode are {1 / 2, 1 / 2}. In particular, if the chroma intra prediction mode of the neighboring block is the proposed fusion mode, it is considered as a non-LM mode. Also, after signaling the DM mode, the four default modes, or the decoder-side derived chroma mode, a flag is signaled to indicate whether to merge with the MMLM_LT mode. In another example, there are two chroma intra prediction modes used for fusion, one mode is a non-LM mode and the other mode is an MMLM_LT mode. The non-LM mode used for the I slice may be one of the DM mode, the four default modes, and the decoder-side derived chroma mode, and the non-LM modes used for the B slice and the P slice are the decoder-side derived chroma modes. Also, the weights of the non-LM mode and the MMLM_LT mode used for the I slice are determined based on the chroma prediction modes of the neighboring blocks, and the weights used for the B slice and the P slice are equal weights.
[0112]
[0119] The above embodiments may be performed as part of a video data processing process, such as encoding process 200A (FIG. 2A) or 200B (FIG. 2B), or decoding process 300A (FIG. 3A) or 300B (FIG. 3B). FIG. 7 is a flowchart of a method 700 for processing video data, according to some embodiments of the present disclosure. The method 700 is used to combine (i.e., fuse) multiple different chroma intra prediction modes to generate an intra predicted chroma block. The method 700 may be performed by an encoder or decoder in predicting a chroma block, or may be performed by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, one or more processors (e.g., processor 402 of FIG. 4) may perform the method 700. In some embodiments, the method 700 may be implemented by a computer program product embodied in a computer-readable medium, the computer program product including computer executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). As shown in FIG. 7, the method 700 includes the following steps 710 to 740.
[0113]
[0120] In step 710, a processor (e.g., processor 402 of FIG. 4) determines a plurality of chroma intra prediction modes. The plurality of chroma intra prediction modes may be applied to a target chroma block to obtain a plurality of predicted chroma blocks, respectively. The obtained plurality of predicted chroma blocks may be combined to generate a final predicted chroma block. In some embodiments, the plurality of chroma intra prediction modes may include at least one LM mode and at least one non-LM mode. The at least one LM mode may include one or more of a CCLM_LT mode, a CCLM_L mode, a CCLM_T mode, a MMLM_LT mode, a MMLM_L mode, or a MMLM_T mode. The at least one non-LM mode may include one or more of a DM mode, a default mode, or a decoder-side derived chroma mode. In some embodiments, the plurality of chroma intra prediction modes may include an MMLM_LT mode and at least one non-LM mode. In some embodiments, the plurality of chroma intra prediction modes may include a decoder-side derived chroma mode and at least one non-LM mode.
[0114]
[0121] In some embodiments, the multiple chroma intra prediction modes may be signaled by one or more flags in the bitstream. For example, the multiple chroma intra prediction modes may include two modes, referred to herein as a first mode and a second mode. When the method 700 is performed by a decoder (e.g., process 300A of FIG. 3A or process 300B of FIG. 3B), the processor decodes the first flag signaled in the bitstream and determines whether to select a first mode from a first set of modes including a DM mode, four default modes, and a decoder-side derived chroma mode based on the value of the first flag. The processor also decodes a second flag signaled in the bitstream and determines whether to select a second mode from a second set of modes including a CCLM_LT mode and an MMLM_LT mode based on the value of the second flag. Correspondingly, when method 700 is performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), a processor encodes in a bitstream one or more flags associated with the plurality of chroma intra prediction modes. For example, the one or more flags may include a first flag indicating whether one of the plurality of chroma intra prediction modes is selected from a first set of modes including a DM mode, four default modes, and a decoder-side derived chroma mode. Also, the one or more flags may include a second flag indicating whether another one of the plurality of chroma intra prediction modes is selected from a second set of modes including a CCLM_LT mode and an MMLM_LT mode.
[0115]
[0122] In some embodiments, one or more of the multiple chroma intra prediction modes may be implicitly determined. For example, the multiple chroma intra prediction modes may include two modes, referred to herein as a first mode and a second mode. When the method 700 is performed by a decoder (e.g., process 300A of FIG. 3A or process 300B of FIG. 3B), the processor decodes a flag signaled in the bitstream and determines whether to select a first mode from a first set of modes including a DM mode, four default modes, and a decoder-side derived chroma mode based on the value of the flag. Then, based on the mode selected from the first set of modes, the processor selects a second mode from a second set of modes including a CCLM_LT mode, a CCLM_L mode, and a CCLM_T mode. For example, if the mode selected from the first set of modes is a non-angular mode, the second mode is determined to be a CCLM_LT mode; if the mode selected from the first set of modes is a horizontal mode, the second mode is determined to be a CCLM_L mode; and if the mode selected from the first set of modes is a vertical mode, the second mode is determined to be a CCLM_T mode.
[0116]
[0123] In some embodiments, explicit signaling and implicit methods are combined to determine one or more of the multiple chroma intra prediction modes. For example, the multiple chroma intra prediction modes may include two modes, referred to herein as a first mode and a second mode. When method 700 is performed by a decoder (e.g., process 300A of FIG. 3A or process 300B of FIG. 3B), the processor decodes a first flag signaled in the bitstream and determines whether to select a first mode from a first set of modes including a DM mode, four default modes, and a decoder-side derived chroma mode based on the value of the first flag. The processor also decodes a second flag signaled in the bitstream and determines whether to select a second mode from a second set of modes including a CCLM_LT mode, a CCLM_L mode, a CCLM_T mode, a MMLM_LT mode, a MMLM_L mode, and a MMLM_T mode based on the value of the second flag. Then, based on the type of the mode selected from the first set of modes, the processor determines which mode in the second set of modes should be selected as the second mode. For example, if the mode selected from the first set of modes is a non-angle mode, the second mode is determined to be one of the CCLM_LT mode or the MMLM_LT mode, if the mode selected from the first set of modes is a horizontal mode, the second mode is determined to be one of the CCLM_L mode or the MMLM_L mode, and if the mode selected from the first set of modes is a vertical mode, the second mode is determined to be one of the CCLM_T mode or the MMLM_T mode.
[0117]
[0124] In some embodiments, the processor determines the plurality of chroma intra prediction modes based on a type of slice associated with the target chroma block. For example, in response to the target chroma block being associated with an I slice, the processor determines that the plurality of chroma intra prediction modes include at least one of a DM mode, a default mode, or a decoder-side derived chroma mode, or in response to the target chroma block being associated with a B slice or a P slice, the processor determines that the plurality of chroma intra prediction modes include a decoder-side derived chroma mode.
[0118]
[0125] In step 720, the processor generates a number of predicted chroma samples associated with a pixel, the pixel being within a target chroma block, by respectively using a number of chroma intra prediction modes.
[0119]
[0126] In step 730, the processor determines a plurality of weights for the plurality of chroma intra prediction modes, respectively. The plurality of weights are used to determine a weighted sum of the plurality of predicted chroma samples. The sum of the plurality of weights is equal to 1. The plurality of weights may be equal or unequal. For example, if the plurality of chroma intra prediction modes has two modes, the processor may set the weight of each of the two modes to be equal to 1 / 2. As another example, if it is determined that the plurality of chroma intra prediction modes includes an LM mode and a non-LM mode (step 710), the processor may assign unequal weights to the LM mode and the non-LM mode, such as 1 / 4 to the LM mode and 3 / 4 to the non-LM mode.
[0120]
[0127] In some embodiments, an index pointing to multiple weights may be signaled in the bitstream. Specifically, a look-up table including multiple sets of weights and their respective indexes is pre-determined. By signaling one of the indexes in the bitstream, the video decoder can know which set of weights to use. More specifically, when method 700 is performed by a decoder (e.g., process 300A of FIG. 3A or process 300B of FIG. 3B), the processor decodes the index signaled in the bitstream and determines the set of weights associated with the index based on the look-up table. Correspondingly, when method 700 is performed by an encoder (e.g., process 200A of FIG. 2A or process 200B of FIG. 2B), the processor selects an index from the look-up table associated with the set of weights used to determine a weighted sum of multiple predicted chroma samples. The processor then encodes the selected index in the bitstream.
[0121]
[0128] In some embodiments, the processor may determine the weights based on the intra-prediction modes used to predict the neighboring chroma blocks of the target chroma block. For example, the neighboring chroma blocks include a neighboring block above the target chroma block and a neighboring block to the left of the target chroma block. For example, the multiple chroma intra-prediction modes are determined to include LM mode and non-LM mode (step 710). Then, if both the upper and left neighboring blocks are LM mode coded, the processor determines the weights of the LM mode and non-LM mode to be 3 / 4 and 1 / 4, respectively, if both the upper and left neighboring blocks are non-LM mode coded, the processor determines the weights of the LM mode and non-LM mode to be 1 / 4 and 3 / 4, respectively, if one of the upper and left neighboring blocks is LM mode coded and the other is non-LM mode coded, the processor determines the weights of the LM mode and non-LM mode to be equal, i.e., 1 / 2. In some embodiments, the determination of the weights based on the intra-prediction modes used to predict neighboring chroma blocks of the target chroma block may be triggered by certain predefined conditions. For example, if the chroma intra-prediction modes determined in step 710 include an LM mode and a non-LM mode, and the non-LM mode is a non-angular mode, the processor triggers a method for determining the weights based on the intra-prediction modes used to predict neighboring chroma blocks of the target chroma block.
[0122]
[0129] In some embodiments, the processor may determine the weights based on the sample location in the target chroma block. For example, the chroma intra prediction modes determined in step 710 include a non-angular mode and an angular mode. The non-angular mode may be one of a planar mode, a DC mode, or a LM mode. The processor may then determine the weights for the angular mode and the non-angular mode as follows: the processor first divides the target chroma block into four equal-area regions vertically (for a horizontal mode) or horizontally (for a vertical mode), and then assigns a different set of weights to each region. As another example, the processor may set the weights associated with the angular mode to be proportional to the distance from the pixel location of the current sample to the pixel location of the reference sample, while setting the weights associated with the non-angular mode to be inversely proportional to the distance. In some embodiments, the determination of the weights based on the sample location in the target chroma block may be triggered by a certain predefined condition. For example, if the multiple chroma intra-prediction modes determined in step 710 include an LM mode and a non-LM mode, and the non-LM mode is an angular mode, the processor triggers a method for determining multiple weights based on a sample position within the target chroma block.
[0123]
[0130] In some embodiments, the processor may determine the multiple weights based on a type of slice associated with the target chroma block. For example, in response to the first predicted chroma sample being associated with an I slice, the processor determines the multiple weights based on a chroma intra prediction mode used to predict one or more neighboring blocks of the target chroma block, or in response to the first predicted chroma sample being associated with a B slice or a P slice, the processor determines that the multiple weights have equal values.
[0124]
[0131] In step 740, the processor determines a first predicted chroma sample associated with the pixel based on a weighted sum of the multiple predicted chroma samples. The first predicted chroma sample is calculated according to Equation 6. The processor repeats this calculation for each pixel in the target chroma block, and the resulting predicted chroma samples form the final predicted chroma block.
[0125]
[0132] In some embodiments, a non-transitory computer-readable storage medium is also provided that stores a bitstream for processing according to the above method. For example, the bitstream includes encoded syntax elements, such as flags, indexes, etc., that indicate the chroma intra-prediction modes to be fused or weights associated with the chroma intra-prediction modes, respectively. In some embodiments, a non-transitory computer-readable storage medium is also provided that includes instructions, which may be executed by an apparatus (such as the disclosed encoders and decoders) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions of the above. The apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0126]
[0133] In some embodiments, a computer program product is provided, the program product including computer program instructions that enable a computer to perform the steps of the method according to any of the embodiments of the present disclosure.
[0127]
[0134] In some embodiments, a computer program is provided, the computer program enabling a computer to carry out the steps of the method according to any of the embodiments of the present disclosure.
[0128]
[0135] It should be noted that relative terms such as "first" and "second" herein are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the terms "comprise," "have," "contain," and "include," as well as other similar forms of terms, are intended to be equivalent in meaning and are non-limiting in that the items following any one of these terms are not intended to be an exhaustive listing of such items or to be limited to only the items listed.
[0129]
[0136] The embodiments may be further described using the following clauses.
[0130]
[0137] 1. A video processing method comprising: generating a plurality of predicted chroma samples associated with the pixel by respectively using a plurality of chroma intra prediction modes; determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples; A video processing method comprising:
[0131]
[0138] 2. The method of claim 1, wherein the multiple chrominance intra-prediction modes include at least one linear model (LM) mode and at least one non-LM mode.
[0132]
[0139] 3. The at least one LM mode includes one or more of a cross-component linear model LT (CCLM_LT) mode, a CCLM_L mode, a CCLM_T mode, a multi-model linear model LT (MMLM_LT) mode, a MMLM_L mode, a MMLM_T mode, a convolutional cross-component model (CCCM) mode, or a gradient linear model (GLM) mode; and 3. The method of claim 2, wherein the at least one non-LM mode includes one or more of a direct mode (DM) mode, a default mode, or a decoder-side derived chroma mode.
[0133]
[0140] 4. The method of claim 1, wherein the multiple chroma intra prediction modes include a multi-model linear model LT (MMLM_LT) mode and at least one non-LM mode.
[0134]
[0141] 5. The method of claim 1, wherein the multiple chroma intra-prediction modes include a decoder-side derived chroma mode and at least one LM mode.
[0135]
[0142] 6. The method of claim 1, further comprising determining a plurality of weights for a plurality of chrominance intra prediction modes, respectively, wherein a sum of the plurality of weights is equal to one.
[0136]
[0143] 7. The method of claim 1, wherein the multiple chroma intra prediction modes are equally weighted in determining the first predicted chroma sample.
[0137]
[0144] 8. The plurality of chroma intra prediction modes includes a linear model (LM) mode and a non-LM mode; and 2. The method of claim 1, wherein the LM mode and the non-LM mode are unequally weighted in determining the first predicted chroma sample.
[0138]
[0145] 9. The method is used to decode a bitstream; and determining a plurality of weights based on an index signaled in the bitstream; determining a weighted sum of the predicted chroma samples using the determined weights; and 2. The method of claim 1, further comprising:
[0139]
[0146] 10. The method is used to encode video data; and Encoding in the bitstream an index associated with a plurality of weights used to determine a weighted sum of a plurality of predicted chroma samples. 2. The method of claim 1, further comprising:
[0140]
[0147] 11. The first predicted chroma sample is part of a first predicted chroma block, and the method further comprises: determining a plurality of weights based on a chroma intra prediction mode used to predict one or more neighboring blocks of the first predicted chroma block; determining a weighted sum of the predicted chroma samples using the determined weights; 2. The method of claim 1, further comprising:
[0141]
[0148] 12. The method of claim 11, wherein the one or more neighboring blocks include a neighboring block above the first predicted chroma block and a neighboring block to the left of the first predicted chroma block.
[0142]
[0149] 13. The method of claim 11, wherein the multiple chrominance intra prediction modes include a linear model (LM) mode and a non-LM mode, and the determination of multiple weights based on the chrominance intra prediction mode used to predict one or more neighboring blocks is triggered by a determination that the non-LM mode is a non-angular mode.
[0143]
[0150] 14. Determining a plurality of weights based on a pixel location of the first predicted chroma sample; determining a weighted sum of the predicted chroma samples using the determined weights; 2. The method of claim 1, further comprising:
[0144]
[0151] 15. The plurality of predicted chroma samples includes an angular mode, and the method further comprises: determining a weight for the angle mode based on a pixel location of the first predicted chroma sample within the predicted chroma block; 15. The method of claim 14, further comprising:
[0145]
[0152] 16. The plurality of predicted chroma samples includes an angular mode or a non-angular mode, and the method further comprises: determining at least one of an angular mode weight or a non-angular mode weight based on a distance from the first predicted chroma sample to a reference sample; 15. The method of claim 14, further comprising:
[0146]
[0153] 17. The method of claim 16, wherein the weight of the angle mode is proportional to the distance from the first predicted chroma sample to the reference sample.
[0147]
[0154] 18. The method of claim 16, wherein the weight of the non-angular mode is inversely proportional to the distance from the first predicted chroma sample to the reference sample.
[0148]
[0155] 19. The method of claim 14, wherein the multiple chroma intra prediction modes include a linear model (LM) mode and a non-LM mode, and the determination of the multiple weights based on the position of the first predicted chroma sample is triggered by a determination that the non-LM mode is an angular mode.
[0149]
[0156] 20. The method is used for decoding a bitstream, and Determining multiple chrominance intra-prediction modes based on one or more flags signaled in the bitstream 2. The method of claim 1, further comprising:
[0150]
[0157] 21. The plurality of chrominance intra-prediction modes includes a first mode and a second mode; the bitstream signals a first set of chroma intra prediction modes and a second set of chroma intra prediction modes; and The method comprises: determining whether to select a first mode from a first set of chrominance intra-prediction modes based on a value of a first flag signaled in the bitstream; determining whether to select a second mode from the second set of chrominance intra-prediction modes based on a value of a second flag signaled in the bitstream; 21. The method of claim 20, further comprising:
[0151]
[0158] 22. The first set of chroma intra-prediction modes includes a DM mode, a default mode, and a decoder-side derived chroma mode; and 22. The method of claim 21, wherein the second set of chrominance intra prediction modes includes a CCLM_LT mode and an MMLM_LT mode.
[0152]
[0159] 23. Selecting one of a first set of chrominance intra prediction modes as a first mode; selecting one of a second set of chrominance intra prediction modes as a second mode based on the mode selected from the first set of chrominance intra prediction modes; 22. The method of claim 21, further comprising:
[0153]
[0160] 24. A second set of chroma intra prediction modes includes a CCLM_LT mode, a CCLM_L mode, a CCLM_T mode, a MMLM_LT mode, a MMLM_L mode, and a MMLM_T mode; and Selecting one of the second set of chrominance intra prediction modes as the second mode includes: selecting one of the CCLM_LT mode or the MMLM_LT mode as a second mode in response to the first mode being a non-angle mode; selecting one of the CCLM_L mode or the MMLM_L mode as the second mode in response to the first mode being the horizontal mode; or selecting one of the CCLM_T mode or the MMLM_T mode as a second mode in response to the first mode being a vertical mode. 24. The method according to claim 23, comprising:
[0154]
[0161] 25. The plurality of chrominance intra-prediction modes include a first mode and a second mode; the bitstream signals a first set of chroma intra prediction modes; and The method comprises: selecting one of a first set of chrominance intra-prediction modes as a first mode based on a value of a first flag signaled in the bitstream; determining a second mode based on the selected mode; and 21. The method of claim 20, further comprising:
[0155]
[0162] 26. Determining a second mode based on a selected mode includes: determining a second mode as a CCLM_LT mode in response to the selected mode being a non-angle mode; determining the second mode to be a CCLM_L mode in response to the selected mode being a horizontal mode; or determining a second mode as a CCLM_T mode in response to the selected mode being a vertical mode; 26. The method according to claim 25, comprising:
[0156]
[0163] 27. The method is used to encode video data; and Encoding in the bitstream one or more flags associated with multiple chrominance intra-prediction modes 2. The method of claim 1, further comprising:
[0157]
[0164] 28. The method of clause 1, further comprising determining a plurality of chroma intra prediction modes based on a type of slice associated with the first predicted chroma sample.
[0158]
[0165] 29. In response to the first predicted chroma sample being associated with an I slice, determining that the plurality of chroma intra prediction modes includes at least one of a DM mode, a default mode, or a decoder-side derived chroma mode; or determining, in response to the first predicted chroma sample being associated with a B slice or a P slice, that the plurality of chroma intra prediction modes includes a decoder-side derived chroma mode; 29. The method of claim 28, further comprising:
[0159]
[0166] 30. Determining a plurality of weights based on a type of slice associated with the first predicted chroma sample; determining a weighted sum of the predicted chroma samples using the determined weights; 2. The method of claim 1, further comprising:
[0160]
[0167] 31. In response to the first predicted chroma sample being associated with an I slice, determining a plurality of weights based on a chroma intra prediction mode used to predict one or more neighboring blocks of the first predicted chroma block; or determining, in response to the first predicted chroma sample being associated with a B slice or a P slice, that the weights have equal values; 31. The method of claim 30, further comprising:
[0161]
[0168] 32. An apparatus comprising: a memory configured to store a set of instructions; and one or more processors, the one or more processors comprising: generating a plurality of predicted chroma samples associated with the pixel by respectively using a plurality of chroma intra prediction modes; determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples; The apparatus is configured to execute a set of instructions to cause the apparatus to perform the
[0162]
[0169] 33. A non-transitory computer-readable medium for storing a video bitstream for processing according to a method, the method comprising: generating a plurality of predicted chroma samples associated with the pixel by respectively using a plurality of chroma intra prediction modes; determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples; A non-transitory computer readable medium comprising:
[0163]
[0170] 34. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for processing video data, the method comprising: generating a plurality of predicted chroma samples associated with the pixel by respectively using a plurality of chroma intra prediction modes; determining a first predicted chroma sample based on a weighted sum of the plurality of predicted chroma samples; A non-transitory computer readable medium comprising:
[0164]
[0171] As used herein, unless otherwise specified, the word "or" includes all possible combinations unless impracticable. For example, if a database is stated to include A or B, the database can include A, B, A, and B, unless otherwise specified or impracticable. As a second example, if a database is stated to include A, B, or C, the database can include A, B, C, A and B, A and C, B and C, A and B and C, unless otherwise specified or impracticable.
[0165]
[0172] It will be understood that the above described embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0166]
[0173] In the above specification, the embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications to the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the present disclosure being indicated by the appended claims. The order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. As such, one skilled in the art can appreciate that the steps may be performed in different orders while implementing the same method.
[0167]
[0174] Although illustrative embodiments have been disclosed in the drawings and herein, many variations and modifications thereto may be made, and therefore, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A video decoding method, comprising: decoding a first flag from the bitstream; determining a first chroma prediction mode based on the first flag; determining whether predicted chroma samples are generated by the first chroma prediction mode or by the first chroma prediction mode and a cross-component prediction mode; reconstructing a picture based on the predicted chroma samples; A video decoding method comprising:
2. the first chroma prediction mode includes one or more of a direct mode (DM) mode, a default mode, or a decoder-side derived chroma mode; and 2. The method of claim 1, wherein the cross-component prediction mode comprises one or more of a cross-component linear model LT (CCLM_LT) mode, a CCLM_L mode, a CCLM_T mode, a multi-model linear model LT (MMLM_LT) mode, a MMLM_L mode, a MMLM_T mode, a convolutional cross-component model (CCCM) mode, or a gradient linear model (GLM) mode.
3. determining whether the predicted chroma samples are generated by the first chroma prediction mode or by the first chroma prediction mode and the cross-component prediction mode based on a second flag; and The method of claim 1 , further comprising: decoding the second flag from the bitstream in response to the first chroma prediction mode being direct mode (DM).
4. determining whether the predicted chroma samples are generated by the first chroma prediction mode or by the first chroma prediction mode and the cross-component prediction mode based on a second flag; and The method of claim 1 , further comprising: decoding the second flag from the bitstream in response to the first chroma prediction mode being a decoder-side derived chroma mode.
5. selecting the first chroma prediction mode from a list of chroma prediction modes based on the first flag; decoding a second flag from the bitstream in response to the selection of a predetermined chroma prediction mode from the list; and determining, based on the second flag, whether the predicted chroma sample is generated by the first chroma prediction mode or by the first chroma prediction mode and the cross-component prediction mode; The method of claim 1 further comprising:
6. If the predicted chroma sample is associated with an I slice, the first chroma prediction mode is selected from a non-linear mode or a linear mode; or The method of claim 1 , wherein if the predicted chroma sample is associated with a B slice or a P slice, the first chroma prediction mode is selected from a decoder-side intra-mode derived chroma mode or a linear mode.
7. generating the predicted chroma samples according to the first chroma prediction mode and the cross-component prediction mode includes: generating a first predicted chroma sample according to the first chroma prediction mode; generating a second predicted chroma sample according to the cross-component chroma prediction mode; and generating a third predicted chroma sample based on a weighted sum of the first predicted chroma sample and the second predicted chroma sample; The method of claim 1 , comprising:
8. 8. The method of claim 7, further comprising determining weights for each of the first and second predicted chroma samples based on coding information of neighboring chroma blocks of the first and second predicted chroma samples.
9. The method of claim 8 , wherein the weights of the first and second predicted chroma samples are 3:1 or 1:
3.
10. 1. A video encoding method, comprising: encoding a first flag in the bitstream indicating a first chroma prediction mode; generating a predicted chroma sample according to the first chroma prediction mode or according to the first chroma prediction mode and a cross-component prediction mode; encoding the predicted chroma samples; A video encoding method comprising:
11. the first chroma prediction mode includes one or more of a direct mode (DM) mode, a default mode, or a decoder-side derived chroma mode; and 11. The method of claim 10, wherein the cross-component prediction mode comprises one or more of a cross-component linear model LT (CCLM_LT) mode, a CCLM_L mode, a CCLM_T mode, a multi-model linear model LT (MMLM_LT) mode, a MMLM_L mode, a MMLM_T mode, a convolutional cross-component model (CCCM) mode, or a gradient linear model (GLM) mode.
12. 11. The method of claim 10, further comprising: in response to the first chroma prediction mode being a direct mode (DM), encoding within the bitstream a second flag indicating whether the predicted chroma samples are generated by the first chroma prediction mode or by the first chroma prediction mode and the cross-component prediction mode.
13. 11. The method of claim 10, further comprising: in response to the first chroma prediction mode being a decoder-side derived chroma mode, encoding, into the bitstream, a second flag indicating whether the predicted chroma samples are generated by the first chroma prediction mode or by the first chroma prediction mode and the cross-component prediction mode.
14. selecting the first chroma prediction mode from a list of chroma prediction modes based on the first flag; in response to the selection of a predetermined chroma prediction mode from the list, encoding within the bitstream a second flag indicating whether the predicted chroma samples are generated by the first chroma prediction mode or by the first chroma prediction mode and the cross-component prediction mode; The method of claim 10 further comprising:
15. If the predicted chroma sample is associated with an I slice, the first chroma prediction mode is selected from a non-linear mode or a linear mode; or The method of claim 10 , wherein if the predicted chroma sample is associated with a B slice or a P slice, the first chroma prediction mode is selected from a decoder-side intra-mode derived chroma mode or a linear mode.
16. generating the predicted chroma samples according to the first chroma prediction mode and the cross-component prediction mode includes: generating a first predicted chroma sample according to the first chroma prediction mode; generating a second predicted chroma sample according to the cross-component chroma prediction mode; and generating a third predicted chroma sample based on a weighted sum of the first predicted chroma sample and the second predicted chroma sample; The method of claim 10, comprising:
17. 17. The method of claim 16, further comprising determining weights for each of the first and second predicted chroma samples based on coding information of neighboring chroma blocks of the first and second predicted chroma samples.
18. The method of claim 17 , wherein the weights of the first and second predicted chroma samples are 3:1 or 1:
3.
19. 1. A method for storing a video bitstream, comprising: generating predicted chroma samples according to a first chroma prediction mode or according to the first chroma prediction mode and a cross-component prediction mode; generating a bitstream; storing the bitstream in a non-transitory computer readable medium; and the bitstream comprises: a first flag indicating the first chroma prediction mode; and coded information indicative of the predicted chroma sample; A method comprising:
20. the first chroma prediction mode includes one or more of a direct mode (DM) mode, a default mode, or a decoder-side derived chroma mode; and 20. The method of claim 19, wherein the cross-component prediction mode comprises one or more of a cross-component linear model LT (CCLM_LT) mode, a CCLM_L mode, a CCLM_T mode, a multi-model linear model LT (MMLM_LT) mode, a MMLM_L mode, a MMLM_T mode, a convolutional cross-component model (CCCM) mode, or a gradient linear model (GLM) mode.