Method and system for performing composite inter and intra prediction
CIIP addresses the inefficiencies in advanced video coding standards by employing a template-based intra-mode derivation method to enhance inter and intra prediction, resulting in improved compression efficiency and reduced bandwidth usage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ALIBABA DAMO (HANGZHOU) TECH CO LTD
- Filing Date
- 2022-06-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing video coding standards face challenges in achieving high compression efficiency, particularly with the development of advanced standards like VVC/H.266, where existing prediction methods are not optimized for combined inter and intra prediction, leading to suboptimal coding performance.
The implementation of composite inter and intra prediction (CIIP) using a template-based intra-mode derivation (TIMD) method to determine intra prediction modes, weighting intra and inter predictors, and signaling flags for decoder usage, enhancing the prediction process.
CIIP improves coding efficiency by optimizing prediction techniques, allowing for better compression performance and reduced bandwidth requirements without significant quality degradation.
Smart Images

Figure 0007894893000006 
Figure 0007894893000007 
Figure 0007894893000008
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications [1] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 215,474, filed Jun. 27, 2021, U.S. Provisional Patent Application No. 63 / 235,110, filed Aug. 19, 2021, and U.S. Patent Application Publication No. 17 / 808,212, filed Jun. 22, 2022, which are hereby incorporated by reference in their entireties.
[0002] Technical Field [2] This disclosure generally relates to video processing, and more particularly, to methods and systems for performing combined inter and intra prediction.
Background Art
[0003] Background [3] Video is a collection of static pictures (or "frames") that capture visual information. To reduce memory and transmission bandwidth, video can be compressed before being stored or transmitted and restored before being displayed. The compression process is usually called encoding, and the restoration process is usually called decoding. Most commonly, there are various video coding formats that use standardized video coding techniques based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standards such as the High - Efficiency Video Coding (HEVC / H.265) standard, which specifies a particular video coding format, the Versatile Video Coding (VVC / H.266) and the AVS standard have been developed by standardization organizations. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards becomes higher.
Summary of the Invention
Means for Solving the Problems
[0004] Summary of the Disclosure [4] Embodiments of the present disclosure provide a method for performing composite inter and intra prediction (CIIP). The method includes determining that CIIP is effective for a target block, determining a first intra prediction mode for the target block using a template-based intra mode derivation (TIMD) method, generating intra predictors for the target block in the first intra prediction mode, and obtaining a final predictor for the target block by weighting the intra predictors and inter predictors of the target block.
[0005] [5] Embodiments of the present disclosure provide a method for performing composite inter and intra prediction (CIIP). The method includes determining a first intra prediction mode for a block of interest using a template-based intra-mode derivation (TIMD) method; generating intra predictors for the block of interest using the first intra prediction mode; obtaining a final predictor for the block of interest by weighting the intra predictors and inter predictors of the block of interest; and signaling a flag indicating that CIIP is enabled and an index indicating that the TIMD method is used to determine the first intra prediction mode for the block of interest.
[0006] [6] Embodiments of the present disclosure provide a non-temporary computer-readable storage medium for storing a bitstream, the bitstream comprising a flag and a first index relating to encoded video data, the flag indicating that inter and intra predictive (CIIP) is used for encoded video data, and the first index indicating a template-based intra-mode derivation (TIMD) method used for CIIP, the flag and the index causing a decoder to use the TIMD method to determine a first intra predictive mode for a target block, generate an intra predictor for the target block in the intra predictive mode, and obtain a final predictor for the target block by weighting the intra predictor and inter predictor for the target block.
[0007] Brief explanation of the drawing [7] Embodiments and various aspects of the present disclosure are shown in the following detailed description and accompanying drawings. The various features shown are not drawn to scale. [Brief explanation of the drawing]
[0008] [Figure 1] [8] This is a schematic diagram showing the structure of an example of a video sequence according to some embodiments of the present disclosure. [Figure 2A] [9] A schematic diagram illustrating an exemplary encoding process of a hybrid video encoding system consistent with embodiments of the present disclosure. [Figure 2B]
[10] A schematic diagram showing another exemplary encoding process for a hybrid video encoding system consistent with embodiments of the present disclosure. [Figure 3A]
[11] A schematic diagram illustrating an exemplary decoding process of a hybrid video encoding system consistent with embodiments of the present disclosure. [Figure 3B]
[12] A schematic diagram showing another exemplary decoding process for a hybrid video encoding system consistent with embodiments of the present disclosure. [Figure 4]
[13] A block diagram of an exemplary device for encoding or decoding video according to some embodiments of the present disclosure. [Figure 5]
[14] An angle intra-prediction mode in VVC according to some embodiments of the present disclosure. [Figure 6]
[15] Exemplary adjacency blocks used in deriving a general most likely mode (MPM) list according to some embodiments of the present disclosure are shown. [Figure 7]
[16] An example of a pixel used to calculate the gradient in a decoder-side intra-mode derivation (DIMD) according to some embodiments of the present disclosure is shown. [Figure 8]
[17] Predictive blending of DIMD according to some embodiments of the present disclosure. [Figure 9]
[18] Exemplary templates and reference samples used in template-based intra-mode derivation (TIMD) according to some embodiments of the present disclosure are shown. [Figure 10]
[19] The upper and left adjacent blocks used to derive the weights of composite inter and intra prediction (CIIP) according to some embodiments of the present disclosure are shown. [Figure 11]
[20] An exemplary flowchart of an extended CIIP mode using a location-dependent intra-predictive combination (PDPC) according to some embodiments of the present disclosure is shown. [Figure 12]
[21] An exemplary flowchart of a method for generating intra predictors in CIIP according to some embodiments of the present disclosure is shown. [Figure 13]
[22] Another exemplary flowchart of a method for generating intra predictors in CIIP according to some embodiments of the present disclosure is shown. [Figure 14]
[23] Another exemplary flowchart of a method for generating intra predictors in CIIP according to some embodiments of the present disclosure is shown. [Figure 15]
[24] Another exemplary flowchart of a method for generating intra predictors in CIIP according to some embodiments of the present disclosure is shown. [Figure 16]
[25] Another exemplary flowchart of a method for generating intra predictors in CIIP according to some embodiments of the present disclosure is shown. [Figure 17A]
[26] An exemplary method for vertically dividing a coded block is shown according to some embodiments of the present disclosure. [Figure 17B]
[26] An exemplary method for horizontally dividing a coded block is shown according to some embodiments of the present disclosure. [Figure 18]
[27] An exemplary method for dividing coded blocks in angle intra-prediction mode is shown according to some embodiments of the present disclosure. [Figure 19]
[28] An exemplary flowchart of a method for determining a prediction mode according to some embodiments of the present disclosure is shown.
Best Mode for Carrying Out the Invention
[0009] Detailed Description
[29] Here, reference is made in detail to exemplary embodiments shown by way of example in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise specified, the same numbers in different figures represent the same or similar elements. The implementation forms described in the following description of the exemplary embodiments do not represent all implementation forms that conform to the present invention. Rather, those implementation forms are merely examples of devices and methods that conform to aspects related to the present invention as recited in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the case of conflict with terms and / or definitions incorporated by reference, the terms and definitions shown in this specification shall prevail.
[0010]
[30] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0011]
[31] In order to achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. Since the coding technology was incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.
[0012]
[32] The VVC standard is a recent development and continues to include more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0013]
[33] Images are a collection of static pictures (or "frames") arranged in chronological order to store visual information. An image capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and an image playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with display capabilities) can be used to display such pictures in chronological order. In some applications, the image capture device can transmit the captured images in real time to an image playback device (e.g., a computer with a monitor) for surveillance, conferencing, or live broadcasting.
[0014]
[34] To reduce the memory space and bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a general-purpose computer processor) or dedicated hardware. Modules for compression are generally called “encoders,” and modules for decompression are generally called “decoders.” Encoders and decoders can be collectively called “codecs.” Encoders and decoders can be implemented as any of a variety of appropriate hardware, software, or a combination thereof. For example, hardware implementations of encoders and decoders may include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), rewritable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer executable instructions, firmware, or algorithms or processes implemented by any appropriate computer fixed in computer-readable media. Video compression and decompression can be implemented using various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, and the H.26x series. In some applications, a codec can decompress video from a first encoding standard and then recompress the decompressed video using a second encoding standard; in this case, the codec can be called a "transcoder."
[0015]
[35] A video encoding process can identify and retain useful information that can be used to reconstruct the picture, and ignore information that is not important for reconstruction. If the ignored non-important information cannot be fully reconstructed, such an encoding process can be called “lossy.” Otherwise, such an encoding process can be called “lossy.” Most encoding processes are lossy, which is a trade-off to reduce the required memory space and transmission bandwidth.
[0016]
[36] Useful information in the encoded picture (referred to as the “current picture”) includes changes to the reference picture (e.g., a picture previously encoded and reconstructed). Such changes may include changes in the position, brightness, or color of pixels, of which position changes are the most important. Changes in the position of a group of pixels representing an object may reflect the movement of the object between the reference picture and the current picture.
[0017]
[37] A picture that is coded without referencing another picture (i.e., such a picture is its own reference picture) is called an "I picture". A picture is called a "P picture" if some or all of the blocks within it (e.g., a block that generally refers to a portion of a video picture) are predicted using intra-prediction or inter-prediction with one reference picture (e.g., unidirectional prediction). A picture is called a "B picture" if at least one block within it is predicted using two reference pictures (e.g., bidirectional prediction).
[0018]
[38] Figure 1 shows the structure of an example of a video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video) or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0019]
[39] As shown in Figure 1, the video sequence 100 may include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are sequential, with more pictures between picture 106 and picture 108. In Figure 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and as indicated by the arrow, its reference picture is picture 102. Picture 106 is a B picture, and as indicated by the arrow, its reference pictures are pictures 104 and 108. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures 102-106 are merely examples, and this disclosure is not limited to the examples shown in Figure 1 of the examples of the reference pictures.
[0020]
[40] Typically, a video codec does not encode or decode the entire picture at once, because such a task is computationally complex. Rather, a video codec can divide the picture into basic segments and encode or decode the picture segment by segment. In this disclosure, such basic segments are referred to as basic processing units ("BPUs"). For example, structure 110 in Figure 1 shows an example of the structure of a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, basic processing units may be referred to as "macroblocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) and as "encoded tree units" ("CTUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). Basic processing units, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels, can have variable sizes within a picture. The size and shape of the basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail to be maintained within the basic processing unit.
[0021]
[41] A basic processing unit can be a logical unit that may contain various types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color picture may include a luminance component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, and the lumina and chroma components may have basic processing units of the same size. In some video encoding standards (e.g., H.265 / HEVC or H.266 / VVC), the lumina and chroma components may be called a “coded tree block” (“CTB”). Any operation performed on a basic processing unit can be repeated on its lumina and chroma components, respectively.
[0022]
[42] The encoding of video involves several operational stages, examples of which are shown in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process and can therefore be further divided into segments, which are referred to in this disclosure as “basic processing subunits.” In some embodiments, a basic processing subunit may be referred to as a “block” within some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as an “encoded unit” (“CU”) within other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than a basic processing unit. Like a basic processing unit, a basic processing subunit is a logical unit which may contain various types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated on its luma and chroma components, respectively. It should be noted that such divisions may be carried out to further levels depending on the processing requirements. It should also be noted that various stages can divide the basic processing unit using different methods.
[0023]
[43] For example, in the mode determination stage (one example of which is shown in Figure 2B), the encoder can determine which prediction mode (e.g., intrapicture prediction or interpicture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each basic processing subunit.
[0024]
[44] In another example (an example of which is shown in Figures 2A and 2B), during the prediction stage, the encoder can perform prediction operations at the level of a basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called “prediction blocks” or “PB” in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.
[0025]
[45] In another example, during the conversion stage (an example of which is shown in Figures 2A and 2B), the encoder can perform conversion operations on residual sub-processing units (e.g., CUs). However, in some cases, the sub-processing units may still be too large to process. The encoder can further divide the sub-processing units into smaller segments (e.g., called "conversion blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform conversion operations at that level. It should be noted that the division method of the same sub-processing unit may differ between the prediction and conversion stages. For example, in H.265 / HEVC or H.266 / VVC, the prediction and conversion blocks of the same CU may have different sizes and numbers.
[0026]
[46] In the structure 110 of Figure 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, and the boundaries between them are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.
[0027]
[47] In some implementations, a picture can be divided into processing regions to give parallel processing and error tolerance capabilities to the encoding and decoding of the video, so that the encoding or decoding process does not have to depend on information from any other region of the picture for any region of the picture. In other words, each region of the picture can be processed independently. In this way, the codec can process different regions of the picture in parallel and thus increase the efficiency of encoding. Furthermore, if the data of a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without depending on the corrupted or lost data, thus providing error tolerance. Some video encoding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that the various pictures of video sequence 100 may have various division schemes for dividing the picture into regions.
[0028]
[48] For example, in Figure 1, structure 110 is divided into three regions 114, 116 and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 contains four basic processing units. Regions 116 and 118 each contain six basic processing units. It should be noted that the basic processing units, basic sub-units and regions of structure 110 in Figure 1 are merely examples and this disclosure does not limit its embodiments.
[0029]
[49] Figure 2A shows a schematic diagram of an example of an encoding process 200A consistent with embodiments of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As shown in Figure 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 in Figure 1, the video sequence 202 may include a set of pictures (referred to as “original pictures”) arranged in chronological order. Similar to the structure 110 in Figure 1, each original picture in the video sequence 202 may be divided by the encoder into a basic processing unit, a basic processing subunit, or a region for processing. In some embodiments, the encoder can perform process 200A at the level of the basic processing unit with respect to each original picture in the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, and the encoder can encode a basic processing unit in a single iteration of process 200A. In some embodiments, the encoder can execute process 200A in parallel for each region of the original picture in the video sequence 202 (e.g., regions 114-118).
[0030]
[50] In Figure 2A, the encoder can feed the basic processing unit of the original picture of the video sequence 202 (called the "original BPU") to the prediction stage 204 to generate the predicted data 206 and the predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to generate the residual BPU 210. The encoder can feed the residual BPU 210 to the conversion stage 212 and the quantization stage 214 to generate the quantized conversion coefficients 216. The encoder can feed the predicted data 206 and the quantized conversion coefficients 216 to the binary encoding stage 226 to generate the video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226 and 228 can be called the "forward path". During process 200A, the encoder may, after the quantization stage 214, feed the quantized transformation coefficients 216 to the inverse quantization stage 218 and the inverse transformation stage 220 to generate a reconstructed residual BPU 222. The encoder may then use the reconstructed residual BPU 222, along with the predicted BPU 208, to generate a prediction criterion 224 to be used in the prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the “reconstruction path”. The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0031]
[51] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a prediction criterion 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.
[0032]
[52] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video acquisition device (e.g., a camera). As used herein, the term “receive” may mean receiving, inputting, acquiring, retrieving, obtaining, reading, accessing or any action in any manner for the purpose of inputting data.
[0033]
[53] In prediction stage 204, the encoder receives the original BPU and prediction criterion 224 in the current iteration and can perform prediction operations to generate prediction data 206 and the predicted BPU 208. The prediction criterion 224 may be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which may be used to reconstruct the original BPU as the predicted BPU 208 from the prediction data 206 and prediction criterion 224.
[0034]
[54] Ideally, the predicted BPU208 may be identical to the original BPU. However, due to less-than-ideal prediction and reconstruction operations, the predicted BPU208 is generally slightly different from the original BPU. To record such differences, the encoder can generate the predicted BPU208 and then subtract it from the original BPU to generate the residual BPU210. For example, the encoder can subtract the pixel values (e.g., grayscale values or RGB values) of the predicted BPU208 from the corresponding pixel values of the original BPU. As a result of such subtraction between the original BPU and the corresponding pixels of the predicted BPU208, each pixel of the residual BPU210 may have a residual value. Compared to the original BPU, the predicted data 206 and residual BPU210 may have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading quality. Thus the original BPU is compressed.
[0035]
[55] In order to further compress the residual BPU210, in the transformation step 212, the encoder can reduce the spatial redundancy of the residual BPU210 by decomposing it into a set of two-dimensional "basis patterns," each basis pattern being associated with "transformation coefficients." The basis patterns can have the same size (e.g., the size of the residual BPU210). Each basis pattern can represent the fluctuation frequency (e.g., luminance fluctuation frequency) component of the residual BPU210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis pattern. In other words, the decomposition can decompose the fluctuations of the residual BPU210 into the frequency domain. Such decomposition is analogous to the discrete Fourier transform of a function, the basis patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transformation coefficients are analogous to the coefficients associated with the basis functions.
[0036]
[56] Various transformation algorithms can be used with various basis patterns. For example, various transformation algorithms can be used in transformation step 212, such as discrete cosine transform and discrete sine transform. The transformation in transformation step 212 is inversely transformable. That is, the encoder can reconstruct the residual BPU 210 by the inverse operation of the transformation (called the "inverse transform"). For example, to reconstruct the pixels of the residual BPU 210, the inverse transform may be to multiply the values of the corresponding pixels in the basis pattern by the respective coefficients in question, and then add the products to obtain a weighted sum. In the video coding standard, both the encoder and the decoder can use the same transformation algorithm (and therefore the same basis pattern). Thus, the encoder may record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis pattern from the encoder. The transformation coefficients may have fewer bits than the residual BPU 210, but these transformation coefficients can be used to reconstruct the residual BPU 210 without significantly degrading the quality. Therefore, the residual BPU210 is further compressed.
[0037]
[57] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, various basis patterns can represent various fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore information on high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized conversion coefficients 216 by dividing each conversion coefficient by an integer value (called the "quantization scale factor") and rounding the quotient to its nearest integer. After such an operation, some conversion coefficients of the high-frequency basis patterns can be converted to zero, and the conversion coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-value quantized conversion coefficients 216, thereby further compressing the conversion coefficients. The quantization process is also inversely convertible, and the quantized conversion coefficients 216 can be reconstructed into conversion coefficients in the inverse operation of quantization (called "inverse quantization").
[0038]
[58] Because the encoder ignores the remainder of such division in the rounding operation, the quantization stage 214 may be irreversible. Typically, the quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized conversion coefficients 216 may need. To obtain different levels of information loss, the encoder may use different values of the quantization parameters or any other parameters of the quantization process.
[0039]
[59] In the binary encoding stage 226, the encoder may encode the predicted data 206 and the quantized conversion coefficients 216 using a binary encoding technique such as entropy encoding, variable-length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the predicted data 206 and the quantized conversion coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transformation in the transformation stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may use the output data from the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0040]
[60] Referring to the reconstruction path of process 200A, in the inverse quantization step 218, the encoder may perform inverse quantization on the quantized transformation coefficients 216 to generate reconstructed transformation coefficients. In the inverse transformation step 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transformation coefficients. The encoder may generate a prediction criterion 224 to be used in the next iteration of process 200A, in addition to the reconstructed residual BPU 222 and the predicted BPU 208.
[0041]
[61] It should be noted that other variations of process 200A may be used to encode the video sequence 202. In some embodiments, the encoder may perform the steps of process 200A in a different order. In some embodiments, one or more steps of process 200A may be combined into a single step. In some embodiments, a single step of process 200A may be divided into multiple steps. For example, the transformation step 212 and the quantization step 214 may be combined into a single step. In some embodiments, process 200A may include additional steps. In some embodiments, process 200A may omit one or more steps in Figure 2A.
[0042]
[62] Figure 2B shows a schematic diagram of another example 200B of an encoding process consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video encoding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0043]
[63] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or “intra prediction”) can use one or more already coded pixels of neighboring BPUs within the same picture to predict the current BPU. That is, the prediction criterion 224 in spatial prediction may include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or “inter prediction”) can use one or more already coded regions of a picture to predict the current BPU. That is, the prediction criterion 224 in temporal prediction may include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0044]
[64] Referring to process 200B, in the forward path, the encoder performs prediction operations in spatial prediction stage 2042 and temporal prediction stage 2044. For example, in spatial prediction stage 2042, the encoder may perform intra-prediction. With respect to the original BPU of the picture being encoded, the prediction criterion 224 may include one or more adjacent BPUs within the same picture that are encoded (in the forward path) and reconstructed (in the reconstruction path). The encoder may generate a predicted BPU 208 by extrapolating adjacent BPUs. Extrapolation techniques may include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the values of the corresponding pixels for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or in any direction specified within the video encoding standard used. In intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the extrapolation parameters, and the orientation of the adjacent BPUs used relative to the original BPU.
[0045]
[65] In another example, in the temporal prediction stage 2044, the encoder may perform interpretation. With respect to the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may generate reconstructed BPUs in addition to the predicted BPU 208, with the reconstructed residual BPU 222. Once all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed pictures as reference pictures. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference picture (referred to as the "search window"). The location of the search window in the reference picture may be determined based on the location of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (for example, by using the PEL recursive algorithm, block matching algorithm, etc.), the encoder can determine that region as a match region. The match region may have different dimensions from the original BPU (for example, smaller, equal to, larger, or different shape). Since the reference picture and the current picture are separated in time in the timeline (for example, as shown in Figure 1), the match region can be considered to "move" to the original BPU's position over time. The encoder can record the direction and distance of such movement as a "motion vector". If multiple reference pictures are used (for example, as picture 106 in Figure 1), the encoder can search for a match region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the match region for each matching reference picture.
[0046]
[66] Motion estimation can be used to identify various types of motion, such as translation, rotation, and scaling. In interpretation, the prediction data 206 may include, for example, the location of the matching region (e.g., coordinates), motion vectors associated with the matching region, the number of reference pictures, and weights associated with the reference pictures.
[0047]
[67] To generate a predicted BPU208, the encoder may perform a “motion compensation” operation. Motion compensation can be used to reconstruct a predicted BPU208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder may move a matching region of a reference picture according to a motion vector, in which case the encoder can predict the original BPU of the current picture. If multiple reference pictures are used (e.g., picture 106 in Figure 1), the encoder may move the matching region of each reference picture according to an individual motion vector and average the pixel values of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of each matching reference picture, the encoder may add up the weighted sum of the pixel values of the moved matching region.
[0048]
[68] In some embodiments, interpretation may be unidirectional or bidirectional. Unidirectional interpretation may use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in Figure 1 is a unidirectional interpretation picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional interpretation may use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in Figure 1 is a bidirectional interpretation picture in which reference pictures (e.g., pictures 104 and 108) are in both temporal directions relative to picture 104.
[0049]
[69] Continuing to refer to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of the cost function depending on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate the corresponding predicted BPU 208 and predicted data 206.
[0050]
[70] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current BPU being encoded and reconstructed in the current picture), the encoder may directly feed the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). The encoder may feed the prediction criterion 224 to the loop filtering stage 232, in which the encoder may apply loop filtering to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused during the encoding of the prediction criterion 224. Various loop filtering techniques may be applied in the loop filtering stage 232, e.g., deblocking, sample-adaptive offset, adaptive loop filtering, etc. The loop-filtered reference picture may be stored in the buffer 234 (or “decoded picture buffer”) for later use (e.g., to be used as an inter-prediction criterion picture for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode the loop filter parameters (e.g., the strength of the loop filter) along with the quantized transformation coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.
[0051]
[71] Figure 3A shows a schematic diagram of an example of a decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A-2B, the decoder may execute process 300A at the level of basic processing units (BPUs) for each picture encoded in the video bitstream 228. For example, the decoder can execute process 300A in an iterative manner, and the decoder can decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder can execute process 300A in parallel for each region of picture (e.g., regions 114-118) to be encoded within the video bitstream 228.
[0052]
[72] In Figure 3A, the decoder can feed a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode the portion into prediction data 206 and quantized transformation coefficients 216. The decoder can feed the quantized transformation coefficients 216 to the inverse quantization stage 218 and the inverse transformation stage 220 to generate the reconstructed residual BPU 222. The decoder can feed the prediction data 206 to the prediction stage 204 to generate the predicted BPU 208. The decoder can generate a prediction criterion 224 in addition to the reconstructed residual BPU 222 and the predicted BPU 208. In some embodiments, the prediction criterion 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can feed the prediction criterion 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.
[0053]
[73] The decoder can iteratively perform process 300A to decode each encoding BPU of the encoded picture and generate a predictive criterion 224 for encoding the next encoding BPU of the encoded picture. After decoding all encoding BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0054]
[74] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary encoding technique used by the encoder (e.g., entropy encoding, variable-length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other arbitrary lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized conversion coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as, for example, the prediction mode, parameters of the prediction operation, type of conversion, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder can depacketize the video bitstream 228 and then feed it to the binary decoding stage 302.
[0055]
[75] Figure 3B shows a schematic diagram of another example 300B of a decoding process consistent with embodiments of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0056]
[76] In process 300B, with respect to the encoded basic processing unit ("current BPU") of the encoded picture being decoded ("current picture"), the prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction was used by the encoder to encode the current BPU, the prediction data 206 may include prediction mode indicators (e.g., flag values) indicating the intra-prediction, parameters of the intra-prediction operation, etc. Parameters of the intra-prediction operation may include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as reference, the size of the adjacent BPU, extrapolation parameters, the orientation of the adjacent BPU relative to the original BPU, etc. In another example, if inter-prediction was used by the encoder to encode the current BPU, the prediction data 206 may include prediction mode indicators (e.g., flag values) indicating the inter-prediction, parameters of the inter-prediction operation, etc. The parameters for the interpretation operation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with each reference picture, the locations (e.g., coordinates) of one or more matching regions within each reference picture, and one or more motion vectors associated with each matching region.
[0057]
[77] Based on the prediction mode indicator, the decoder can decide whether to perform a spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or a temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of the execution of such spatial or temporal prediction are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder can generate the predicted BPU 208. As shown in Figure 3A, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate the prediction criterion 224.
[0058]
[78] In process 300B, the decoder may feed the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation within the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., the decoded current BPU), the decoder may feed the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., the reference picture from which all BPUs have been decoded), the decoder may feed the prediction criterion 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction criterion 224 in the manner shown in Figure 2B. The loop-filtered reference picture can be stored in buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data includes loop filter parameters.
[0059]
[79] Figure 4 is a block diagram of an example of a device 400 for encoding or decoding video, consistent with embodiments of the present disclosure. As shown in Figure 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 can be a dedicated machine for encoding or decoding video. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any combination of any number of central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), composite programmable logic units (CPLDs), rewritable gate arrays (FPGAs), systems on a chip (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, the processor 402 may also be a set of processors grouped together as a single logical component. For example, as shown in Figure 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0060]
[80] The device 400 may also include a memory 404 configured to store data (e.g., sets of instructions, computer code, intermediate data, etc.). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for implementing stages in processes 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and processing data (e.g., via the bus 410), execute the program instructions, and perform calculations or operations on the processing data. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random-access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash® (CF) cards, etc. Memory 404 can also be a group of memories (not shown in Figure 4) that are grouped together as a single logical component.
[0061]
[81] Buses 410, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a universal serial bus port, a peripheral component interconnection express port), may be communication devices for transferring data between components within the device 400.
[0062]
[82] For simplicity of explanation without causing ambiguity, the processor 402 and other data processing circuits are collectively referred to as “data processing circuits.” The data processing circuits may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single, independent module, or may be fully or partially combined within any other component of the device 400.
[0063]
[83] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth® adapters, infrared adapters, near-field communication ("NFC") adapters, cellular network chips, etc.
[0064]
[84] In some embodiments, the device 400 may optionally further include a peripheral device interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display or a light-emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), and the like.
[0065]
[85] It should be noted that a video codec (for example, a codec that runs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. In another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of the device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0066]
[86] VVC offers multiple intra-prediction modes. Figure 5 shows angular intra-prediction modes in VVC according to some embodiments of the present disclosure. As shown in Figure 5, to capture any edge direction shown in natural image, the number of angular intra-prediction modes in VVC is extended from 33 used in HEVC to 65, with directional modes not present in HEVC shown as dotted arrows.
[0067]
[87] The VVC standard implements two non-angle intra-prediction modes: DC mode and planar mode (as found in HEVC). In DC intra-prediction mode, the average sample value of the reference samples for a block is used for prediction generation. In VVC, only the reference samples along the long side of a rectangular block are used to calculate the average, while in the case of a square block, both the left and top reference samples are used. In planar mode, the predicted sample value is obtained as a weighted average of four reference sample values: the reference sample in the same row or column as the current sample and the reference samples in the lower left and upper right positions relative to the current block. The 65 angle modes and the two non-angle modes can be called normal intra-prediction modes.
[0068]
[88] In some embodiments, a most likely mode list (MPM) is proposed. As discussed above, VVC has 67 angular modes. If the prediction mode of each block were encoded separately, 7 bits would be required to encode the 67 modes. Therefore, VVC employs a method for constructing a most likely mode (MPM) list. In image and video encoding, adjacent blocks are usually strongly correlated, so the intra-prediction modes of adjacent blocks are likely to be the same or similar. Thus, the MPM list is constructed based on the intra-prediction modes of the left adjacent block and the adjacent block above. In VVC, the length of the MPM list is 6. To keep the complexity of generating the MPM list low, an intra-prediction mode encoding method is used that uses 6 MPMs derived from 2 available adjacent intra-prediction modes.
[0069]
[89] Regardless of whether the MRL (Multiple Reference Lines) coding tool and the ISP (Intra Subpartition) coding tool are applied, intra blocks use a unified 6-MPM list, also called the primary MPM (PMPM) list. The MPM list is constructed based on the intra modes of the left adjacent block and the upper adjacent block. Assuming that the intra mode of the left block is denoted as "left" and the intra mode of the upper block is denoted as "up", the unified 6-MPM list is constructed as follows: If no adjacent block is available, the intra prediction mode is set to "planar" by default. If both the left mode and the upper mode are non-angled modes, the MPM list is set to {planar,DC,V,H,V-4,V+4}, where V is called the vertical mode and H is called the horizontal mode. If either the left mode or the up mode is an angle mode and the other is not an angle mode, set the mode to Max as the larger of the left and up modes, and set the MPM list to {Plane, Max, Max-1, Max+1, Max-2, Max+2}. If both left and up are angles and they are different, set the mode to "Max" as the larger of the left and up modes, and set the mode to "Min" as the smaller of the left and up modes. If Max-Min is equal to 1, the MPM list is set to {Plane, Left, Up, Min-1, Max+1, Min-2}; otherwise, if Max-Min is 62 or greater, the MPM list is set to {Plane, Left, Up, Min+1, Max-1, Min+2}. If Max-Min is equal to 2, the MPM list is set to {Plane, Left, Up, Min+1, Min-1, Max+1}; otherwise, the MPM list is set to {Plane, Left, Up, Min-1, Min+1, Max-1}. If both Left and Up are angles and they are the same, the MPM list is set to {Plane, Left, Left-1, Left+1, Left-2, Left+2}. Furthermore, the first bin of the MPM index codeword is context-coded using CABAC (Context-Based Adaptive Binary Arithmetic Coding). A total of three contexts are used, corresponding to whether the current intrablock is MRL-enabled, ISP-enabled, or a regular intrablock.For entropy coding of 61 non-MPM modes, TBC (Truncated Binary Code) is used.
[0070]
[90] In some embodiments, a secondary MPM method may be used. The primary MPM (PMPM) list consists of six entries, and the secondary MPM (SMPM) list contains sixteen entries. First, a general MPM list having 22 entries is created, of which the first six entries are included in the PMPM list and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. Next, the intra-prediction modes of the adjacent blocks are added to the list. Figure 6 shows exemplary adjacent blocks used to derive the general MPM list according to some embodiments of the present disclosure. As shown in Figure 6, the intra-prediction modes of the left (L), top (A), bottom left (BL), top right (AR), and top left (AL) adjacent blocks are used. If the CU block is vertical, the order of the adjacent blocks is A, L, BL, AR, AL. Otherwise, i.e., if the CU block is horizontal, the order of the adjacent blocks is L, A, BL, AR, AL. Next, two decoder-side intra-prediction modes are added to the list. Next, angle modes derived from the first two available angle modes in the list are added to the list by adding an offset. Finally, if the list is incomplete, default modes are added until the list is complete, i.e., until it has 22 entries. According to some embodiments of this disclosure, the default mode list is defined as {DC,V,H,V-4,V+4,14,22,42,58,10,26,38,62,6,30,34,66,2,48,52,16}.
[0071]
[91] In the decoder, the PMPM flag is parsed first. If the PMPM flag is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list is selected; otherwise, the SMPM flag is parsed to determine whether to parse the SMPM index for the remaining modes.
[0072]
[92] In some embodiments, a position-dependent intra-prediction combination (PDPC) is provided. In VVC, the intra-prediction result is further modified by the PDPC method. PDPC is applied without signaling to the following intra-prediction modes: planar, DC, intra-angle modes below horizontal mode and intra-angle modes above vertical mode. PDPC is not applied if the current block is in BDPCM (block-based differential pulse code modulation) mode or if the MRL index is greater than 0.
[0073]
[93] The predicted sample pred(x',y') is given by the following equation using a linear combination of the intra-prediction mode (e.g., DC, planar, or angular mode) and the reference sample: pred(x',y')=Clip(0,(1<<BitDepth)-1,(wL×R-1,y’+wT×Rx’,-1+(64-wL-wT)×pred(x’,y’)+32)> >6) The prediction is made according to the formula, where Rx',-1 and R-1,y' represent the reference samples located above and to the left of the current sample (x',y'), respectively. The PDPC weights and scale factors depend on the prediction mode and block size.
[0074]
[94] Furthermore, a decoder-side intra-predictive mode derivation (DIMD) is provided. In the decoder-side intra-predictive mode derivation method, the intra-predictive mode is not transmitted by the bitstream. Instead, a texture gradient process is performed to derive two best modes. The same format is used on the encoder and decoder sides. Predictors for the two derived modes and the planar mode are successfully computed, and the weighted average of the three predictors is used as the final predictor for the current block.
[0075]
[95] DIMD mode is used as an alternative intra-prediction mode, and a flag is signaled for each block to indicate whether DIMD mode should be used. If the flag is true (e.g., the flag is equal to 1), DIMD mode is used for the current block, and the BDPCM flag, MIP (matrix-weighted intra-prediction) flag, ISP flag, and MRL index are assumed to be 0. In this case, all parsing of intra-prediction modes is also skipped. If the flag is false (e.g., the flag is equal to 0), DIMD mode is not used for the current block, and parsing of other intra-prediction modes proceeds normally.
[0076]
[96] A histogram is constructed by performing a texture gradient process to derive two intra-prediction modes and determine the weights for each mode.
[0077]
[97] Figure 7 shows an exemplary sample used to calculate gradients in DIMD according to some embodiments of the present disclosure. As shown in Figure 7, gradient analysis is performed on sample 710 of an L-shaped template of a second adjacent line surrounding a block in order to construct a DIMD histogram of the block. For each available reconstruction sample of the template, horizontal gradient Gx and vertical gradient Gy are performed by applying horizontal and vertical Sobel filters as follows:
number
[0078]
[98] For each sample in the template in which the horizontal gradient Gx and vertical gradient Gy are calculated, the gradient strength (G) and direction (O) are further calculated using Gx and Gy as follows:
number
[0079]
[99] The gradient direction O is converted to the nearest intra-angle prediction mode and used to index the histogram, which is initially initialized to zero. The histogram values in that intra-angle prediction mode are incremented by G. Once all samples in the template have been processed, the histogram contains the cumulative gradient intensity for each intra-angle prediction mode. For the subsequent prediction fusion process, two modes with the largest and second largest amplitude values are selected and marked as M1 and M2, respectively. If the maximum amplitude value of the histogram is 0, the planar mode is selected as the intra-prediction mode for the current block.
[0080]
[0100] In DIMD, two intra-prediction angle modes M1 and M2, corresponding to two maximum histogram amplitude values, are combined with the planar mode to generate the final predicted value for the current block.
[0081]
[0101] The predictive blend is applied as a weighted average of the three predictors described above. The weights for the planar modes are fixed at 21 / 64 (approximately equal to 1 / 3), proportional to the amplitude values of M1 and M2, with the remaining weights 43 / 64 (approximately equal to 2 / 3) shared between M1 and M2. Figure 8 shows the predictive blend process for DIMD according to some embodiments of the present disclosure. As shown in Figure 8, ampl(M1) and ampl(M2) represent the amplitude values of M1 and M2, respectively.
[0082]
[0102] The DIMD mode is used only for luma blocks. If the current luma block selects the DIMD mode, the intra-predictive mode of the current block is stored as M1 for selecting the low-frequency non-separated transform (LFNST) set of the current block, deriving the most probable mode (MPM) list of adjacent luma blocks, and deriving the direct mode (DM) of the chroma block at the same location.
[0083]
[0103] Furthermore, in some embodiments, a different decoder-side intra-prediction mode derivation method, template-based intra-mode derivation (TIMD) using MPMs, can be used. Instead of signaling, the intra-prediction modes of the CU are derived in a template-based manner on both the encoder and decoder sides. Candidates are constructed from an MPM list, and the candidate modes may be 67 intra-prediction modes, as with VVCs, or they may be extended to 131 intra-prediction modes. Figure 9 shows exemplary templates and reference samples used in TIMD according to some embodiments of the present disclosure. As shown in Figure 9, for each candidate mode, a reference sample 920 of the template is used to generate a prediction sample of template 910. A value is calculated as the sum of absolute transformation differences (SATD) between the prediction sample and the reconstructed sample of the template. The intra-prediction mode with the minimum SATD is selected as the TIMD mode and used for the intra-prediction of the current CU.
[0084]
[0104] TIMD mode is used as an additional intra-prediction method for CU. A flag is signaled in the Sequence Parameter Set (SPS) to enable / disable TIMD. If the flag is true (e.g., the flag is equal to 1), the CU level flag is signaled to indicate whether TIMD is used. The TIMD flag is signaled after the MIP flag. If the TIMD flag is true (e.g., the TIMD flag is equal to 1), all remaining syntax elements related to the intra-prediction mode, including MRL and ISP, and the normal parsing stage for the intra-prediction mode are skipped.
[0085]
[0105] In TIMD, the number of intra-prediction modes has been expanded to 131, so when remembering the intra-prediction modes for the current block, a table is used to map the 131 modes to the original 67 intra-prediction modes of VVC.
[0086]
[0106] In accordance with the disclosed embodiments, the fusion of two intra-predictive modes can be derived from a TIMD method called the TIMD fusion method. Instead of selecting only one mode having the minimum SATD value, the TIMD method is used to select two modes having the first two minimum SATD values, and then the predictors of the two selected modes are blended to generate the final predictor for the current block. The weights of the two modes are inversely proportional to the SATD values of the two modes.
[0087]
[0107] During the construction of the MPM list, if adjacent blocks are intercoded, the intra-prediction mode of adjacent blocks is derived as a planar mode. In accordance with the disclosed embodiments, to improve the accuracy of the MPM list, if adjacent blocks are intercoded, the propagated intra-prediction mode is derived using the motion vector and reference picture of the adjacent blocks, and the propagated intra-prediction mode is used in the construction of the MPM list. Specifically, with respect to an intercoded block, a reference block can be determined according to its own motion vector and reference picture. If a reference block is intracoded, the propagated intra-prediction mode of the current block is set to the intra-prediction mode of the reference block. The propagated intra-prediction mode of the current block can then be used in the construction of the MPM list of adjacent blocks.
[0088]
[0108] Composite Inter and Intra Prediction (CIIP) is provided. In VVC, when a CU is coded in merge mode, an additional flag is signaled to indicate whether CIIP mode applies to the current CU if the CU contains at least 64 luma samples (i.e., CU width × CU height is 64 or greater) and if both the CU width and CU height are less than 128 luma samples. CIIP prediction combines inter predictors with intra predictors. Inter predictor P in CIIP mode. interThis is derived using the same interprediction process applied to normal merge mode, and the intrapredictor P intra The weights are derived according to a normal intra-prediction process in planar mode. Figure 10 shows the upper and left adjacent blocks used to derive the weights of the CIIP in some embodiments of the present disclosure. The intra-predictors and inter-predictors are combined using a weighted average, and the weight values are calculated according to the coding mode of the upper and left adjacent blocks (as shown in Figure 10).
[0089]
[0109] The weights (wIntra,wInter) of the intra predictor and inter predictor are set adaptively as follows: If both the upper and left neighbors are intracoded, (wIntra,wInter) is set to equal (3,1). If one of these blocks is intracoded, the weights are set to be the same, i.e., equal to (2,2). If neither the upper nor left neighbors are intracoded, the weights are set to equal (1,3). The CIIP predictor is formed as follows: P CIIP =(wInter*P inter +wIntra*P intra +2)>>2
[0090]
[0110] In some embodiments, multi-hypothesis predictions for intra-modes and inter-modes can be used. In merged CUs, one flag is signaled with respect to the merged mode to select an intra-prediction mode from the intra-candidate list if the flag is true. In luma components, the intra-candidate list is derived from four intra-prediction modes, including DC mode, planar mode, horizontal mode, and vertical mode. One intra-prediction mode selected by the intra-prediction mode index and one inter-prediction mode selected by the merged index are combined using a weighted average. The weights for combining predictions are described below. If DC or planar mode is selected, or if the width or height of the CU is less than 4, equal weights are applied. For CUs with a width and height of 4 or more, if horizontal / vertical mode is selected, one CU is first divided vertically or horizontally into four equally sized regions. .wIntrai,wInteri) is applied to the corresponding region, where i is 1 to 4, and (wIntra1,wInter1)=(6,2), (wIntra2,wInter2)=(5,3), (wIntra3,wInter3)=(3,5), and (wIntra4,wInter4)=(2,6). (wIntra1,wInter1) is for the region closest to the reference sample, and (wIntra4,wInter4) is for the region furthest from the reference sample. The composite predictor can be calculated by summing the two weighted predictors and right-shifting by a certain number of bits, which is obtained by the logarithm of the sum of the two weights. In this example, the composite predictor is obtained by summing the two weighted predictors and right-shifting by 3 bits. In some embodiments, if the sum of the two weights is equal to 1, the logarithm of 1 is 0, so the composite predictor can be obtained by directly summing the two weighted predictors; right-shifting is not necessary.For example, each weight set could be (wIntra1,wInter1)=(6 / 8,2 / 8), (wIntra2,wInter2)=(5 / 8,3 / 8), (wIntra3,wInter3)=(3 / 8,5 / 8), and (wIntra4,wInter4)=(2 / 8,6 / 8), and applied to the corresponding region. (wIntra1,wInter1) is for the region closest to the reference sample, and (wIntra4,wInter4) is for the region furthest from the reference sample.
[0091]
[0111] In some embodiments, the CIIP_PDPC mode can be used. In CIIP_PDPC, the prediction of the normal merge mode is refined using the upper (Rx,-1) and left (-1,Ry) reconstruction samples. This refinement inherits the position-dependent prediction combination (PDPC) scheme. Figure 11 shows an exemplary flowchart of CIIP_PDPC according to some embodiments of the present disclosure. Referring to Figure 11, WT and WL are weighted values that depend on the sample position within the block defined in the PDPC.
[0092]
[0112] The CIIP_PDPC mode is signaled along with the CIIP mode. If the CIIP flag is true, another flag, the CIIP_PDPC flag, is further signaled to indicate whether to use CIIP_PDPC.
[0093]
[0113] In current CIIP designs, only planar modes are used to derive the intra-prediction portion. Even in multi-hypothesis prediction using intra-mode and inter-mode methods, the intra-prediction modes used to generate the intra-prediction portion can only be selected from a maximum of four modes. Therefore, the intra-predictors in CIIP are not sufficiently accurate.
[0094]
[0114] Accuracy can be improved by allowing the intra-prediction portion of CIIP to select from a wider range of intra-prediction modes. However, this requires more bits to indicate which intra-prediction mode to use. Considering that DIMD and TIMD are two decoder-side intra-prediction mode derivation methods that can conserve signaling for intra-prediction modes, using intra-prediction modes derived by DIMD and / or TIMD to generate the intra-predictor for CIIP may improve coding performance.
[0095]
[0115] This disclosure provides several methods for generating CIIP predictors in order to improve coding performance.
[0096]
[0116] In some embodiments, a method for generating intra predictors in CIIP is proposed.
[0097]
[0117] In the first embodiment, the normal intra-prediction modes available for generating intra-predictors in CIIP are extended to up to 67 modes, which can be called CIIP_NormalModes, as shown in Figure 5. Figure 12 shows an exemplary flowchart of Method 1200 for generating intra-predictors in CIIP according to some embodiments of the present disclosure. Method 1200 may be executed by a decoder (e.g., by process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of a device (e.g., device 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may execute Method 1200. In some embodiments, Method 1200 may be implemented by a computer program product embodied in a computer-readable medium, which includes computer-executable instructions such as program code executed by a computer (e.g., device 400 in Figure 4). Referring to Figure 12, Method 1200 may include the following steps S1202 and S1206.
[0098]
[0118] In S1202, the intra-prediction mode used to generate intra-predictors in CIIP is selected from a plurality of intra-prediction modes. In some embodiments, the intra-prediction mode is selected from the angular intra-prediction mode shown in Figure 5, and two non-angular intra-prediction modes, the planar mode and the DC mode. Therefore, the number of intra-prediction modes used to generate intra-predictors in CIIP can be up to 67. Intra-predictors can be derived according to a normal intra-prediction process. In some embodiments, intra-predictors in CIIP can be generated using non-normal intra-prediction modes such as MIP mode, ISP mode, and MRL mode.
[0099]
[0119] In S1204, a parameter (e.g., an index) indicating the selected intra-prediction mode is decoded. For example, the index values may correspond to different intra-prediction modes (e.g., 0 to 66). In some embodiments, the index may correspond to MIP mode, ISP mode, MRL mode, etc. The index can be coded in a variety of ways, not limited to those specified herein.
[0100]
[0120] In step S1206, an intra predictor is generated in CIIP using the selected intra predictor mode.
[0101]
[0121] Instead of using only the planar mode, multiple intra-prediction modes can be used to generate intra-predictors to improve the accuracy of CIIP.
[0102]
[0122] In a second embodiment, a combination of DIMD and CIIP (e.g., CIIP_DIMD mode) is provided. Specifically, DIMD information is used to generate intra predictors in CIIP.
[0103]
[0123] In some embodiments, the intra-predictor in CIIP is generated using a DIMD blend prediction method. In this example, the final intra-predictor in CIIP is generated by weighting and averaging the predictors of two intra-prediction angle modes and a planar mode, each having the first two maximum amplitude values in the DIMD histogram.
[0104]
[0124] In some embodiments, a DIMD texture gradient process is performed on the current block. The intra predictor in CIIP is generated using an intra predictive mode that has the maximum amplitude value in the DIMD histogram.
[0105]
[0125] In some embodiments, intrapredictors in CIIP are generated using an intraprediction mode selected from several fixed intraprediction modes according to DIMD information. Specifically, the fixed intraprediction modes include planar mode, horizontal mode, and vertical mode. For example, if the intraprediction mode with the maximum amplitude value in the DIMD histogram is close to the horizontal mode (e.g., if the absolute difference between the derived mode index and the horizontal mode index is less than 10), the horizontal mode is used to generate the intrapredictor in CIIP. If the mode is close to the vertical mode (e.g., if the absolute difference between the derived mode index and the vertical mode index is less than 10), the vertical mode is used to generate the intrapredictor in CIIP. Otherwise, the planar mode is used to generate the intrapredictor in CIIP.
[0106]
[0126] In some embodiments, intra predictors in CIIP can be generated by selecting from a subset of intra predictor modes. Figure 13 shows an exemplary flowchart of a method 1300 for generating intra predictors in CIIP according to some embodiments of the present disclosure. Method 1300 may be executed by a decoder (e.g., by process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of a device (e.g., device 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may execute method 1300. In some embodiments, method 1300 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions such as program code executed by a computer (e.g., device 400 in Figure 4). Referring to Figure 13, method 1300 may include the following steps S1302 to S1306.
[0107]
[0127] In step S1302, a subset of intra-prediction modes is derived according to the DIMD histogram. The subset of intra-prediction modes includes N modes that have the first N maximum amplitude values in the DIMD histogram, where N is a positive integer.
[0108]
[0128] In step 1304, the parameters (e.g., index) that indicate the selected intra-prediction mode are determined.
[0109]
[0129] In step 1306, an intra predictor is generated in CIIP using the selected intra predictor mode.
[0110]
[0130] In some embodiments, if the intra-prediction modes used for blend prediction and the weights of each intra-prediction mode differ from the current DIMD design, the intra-predictors in CIIP can be generated using a blend prediction method, an intra-prediction mode with the largest weight, and a non-planar intra-prediction mode with the largest weight.
[0111]
[0131] In a third embodiment, a combination of TIMD and CIIP (e.g., CIIP_TIMD mode) is provided. Specifically, TIMD information is used to generate intra predictors in CIIP. For example, if it is determined that the CIIP mode is valid for the target block, the intra predictors in CIIP are generated using the intra prediction mode determined using the TIMD method.
[0112]
[0132] In some embodiments, intra-predictors in CIIP can be generated using extended intra-prediction modes derived by the TIMD method. In the current TIMD design, the indices of the derived extended intra-prediction modes range from 0 to 130. Therefore, the number of extended intra-prediction modes used to generate intra-predictors in CIIP can be up to 131.
[0113]
[0133] In some embodiments, an extended intra-prediction mode having an index ranging from 0 to 130 derived by the TIMD method can be mapped to a normal intra-prediction mode in a VVC having an index ranging from 0 to 66, and the mapped normal intra-prediction mode is used to generate an intra-predictor in CIIP.
[0114]
[0134] In some embodiments, one of several fixed intra-prediction modes is selected to generate intra-predictors in CIIP according to the intra-prediction mode derived by TIMD. For example, the fixed intra-prediction modes may include planar mode, horizontal mode, and vertical mode. If the derived mode is close to the horizontal mode (e.g., the absolute difference between the derived mode index and the horizontal mode index is less than 10), the horizontal mode is used to generate intra-predictors in CIIP. If the derived mode is close to the vertical mode (e.g., the absolute difference between the derived mode index and the vertical mode index is less than 10), the vertical mode is used to generate intra-predictors in CIIP. Otherwise, the planar mode is used to generate intra-predictors in CIIP.
[0115]
[0135] In some embodiments, a subset of intra-prediction modes is derived according to costs (e.g., SATD) calculated using the TIMD method, and the subset of intra-prediction modes includes N modes with small costs. A parameter (e.g., an index) is signaled to select one mode from the subset of intra-prediction modes, and an intra-predictor in CIIP is generated using the selected mode.
[0116]
[0136] As described above, a list of intra-prediction modes is constructed using the TIMD method, and one intra-prediction is selected from the list by the template-based derivation method. In some embodiments, when the TIMD method is used to derive the intra-prediction modes in CIIP, the list of intra-prediction modes used may differ from that of a normal TIMD.
[0117]
[0137] In some embodiments, intra predictors in CIIP are generated using intra predictive modes derived from a subset of the standard TIMD intra predictive mode list.
[0118]
[0138] In some embodiments, a subset of the normal TIMD intra-prediction mode list is a fixed TIMD intra-prediction mode list, such as a fixed TIMD intra-prediction mode list including planar mode, DC mode, horizontal mode, and vertical mode. The intra-predictor in CIIP is generated using the intra-prediction modes derived from the fixed TIMD intra-prediction mode list.
[0119]
[0139] In some embodiments, the final intra predictor in the CIIP is generated using the intra predictor mode having the smallest SATD value in the TIMD intra predictor mode list. Figure 14 shows an exemplary flowchart of a method 1400 for generating an intra predictor in the CIIP according to some embodiments of the present disclosure. Method 1400 may be executed by a decoder (e.g., by process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of a device (e.g., device 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may execute method 1400. In some embodiments, method 1400 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions such as program code executed by a computer (e.g., device 400 in Figure 4). Referring to Figure 14, method 1400 may include the following steps S1402 to S1406.
[0120]
[0140] In step S1402, the SATD value for each intra-predictive mode in the TIMD mode list is calculated.
[0121]
[0141] In step S1404, the intra-prediction mode with the minimum value is determined as the intra-prediction mode to be used for generating intra-predictors in CIIP.
[0122]
[0142] In step S1406, the determined intra-prediction mode is used to generate intra-predictors in CIIP.
[0123]
[0143] In some embodiments, the intra-prediction mode having the smallest SATD value in the TIMD intra-prediction mode list is further mapped to a normal intra-prediction mode in the VVC (e.g., having indices from 0 to 67), and the mapped normal intra-prediction mode is then used to generate the final intra-predictor in the CIIP.
[0124]
[0144] In a fourth embodiment, a combination of a TIMD fusion method and a CIIP method (e.g., CIIP_TIMD fusion mode) is provided. Specifically, information from the TIMD fusion method is used to generate intra predictors in CIIP. In some embodiments, the final intra predictor in CIIP is generated by weighting the predictors of two intra predictor modes having the first two minimum SATD values in the TIMD intra predictor mode list.
[0125]
[0145] In a fifth embodiment, it is proposed to modify the derivation method of a TIMD method or TIMD fusion method by scaling the SATD values of some intra-prediction modes in the TIMD intra-prediction mode list by a coefficient, and using the scaled SATD values to derive the intra-prediction mode to be used in the current block. Figure 15 shows an exemplary flowchart of a method 1500 for generating intra-predictors in a CIIP according to some embodiments of the present disclosure. Method 1500 may be executed by a decoder (e.g., by process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of a device (e.g., device 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may execute method 1500. In some embodiments, method 1500 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions such as program code executed by a computer (e.g., device 400 in Figure 4). Referring to Figure 15, Method 1500 may include the following steps S1502 to S1506.
[0126]
[0146] In step S1502, the SATD value for the template of each mode in the TIMD intra-prediction mode list is calculated.
[0127]
[0147] In step S1504, one or more SATD values are scaled by a coefficient. In some embodiments, the SATD value in planar mode is multiplied by the coefficient. The coefficient can be any positive value greater than 0 and less than 1. For example, the coefficient may be equal to 0.9.
[0128]
[0148] In some embodiments, the SATD value in planar mode is multiplied by a coefficient based on the block size of the current block, meaning that different coefficients are available for blocks of different sizes. For example, for blocks larger than 1024, the coefficient is equal to 0.8, and for blocks smaller than or equal to 1024, the coefficient is equal to 0.9.
[0129]
[0149] In some embodiments, the SATD value of an intra-prediction mode within a subset of the TIMD intra-prediction mode list is multiplied by a coefficient. The coefficient can be any positive value greater than 0 and less than 1. For example, the subset may be a list containing planar modes and DC modes, or a list containing planar modes, DC modes, horizontal modes, and vertical modes. In this example, the coefficients may differ for different intra-prediction modes. Even for the same intra-prediction mode, the coefficients may differ if the block size is different.
[0130]
[0150] In step S1506, an intra predictor is generated in CIIP using the intra prediction determined using the scaled SATD value.
[0131]
[0151] In a sixth embodiment, the derivation method of the TIMD method or TIMD fusion method can be modified based on the SATD value in planar mode. Figure 16 shows an exemplary flowchart of method 1600 for generating an intra predictor in CIIP according to some embodiments of the present disclosure. Method 1600 may be executed by a decoder (e.g., by process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of a device (e.g., device 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may execute method 1600. In some embodiments, method 1600 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions such as program code executed by a computer (e.g., device 400 in Figure 4). Referring to Figure 16, method 1600 may include the following steps S1602 to S1606.
[0132]
[0152] Step 1602 calculates the SATD value for the template of each mode in the TIMD intra-prediction mode list.
[0133]
[0153] In step 1604, the SATD value for the planar mode is compared to the minimum SATD value.
[0134]
[0154] In step 1606, if the SATD value in planar mode is sufficiently close to the minimum SATD value (for example, the SATD value in planar mode does not exceed 1.2 times the minimum SATD value), the planar mode is used to generate the intra predictor in CIIP.
[0135]
[0155] In some embodiments, methods 1500 and 1600 are used only for CIIP mode-coded blocks that use TIMD or TIMD fusion information.
[0136]
[0156] In the seventh embodiment, a combination of DIMD, TIMD, and CIIP methods is provided. Specifically, DIMD and TIMD information is used together to generate an intra predictor in CIIP.
[0137]
[0157] In some embodiments, to generate the final intra predictor in CIIP, an intra predictor generated by DIMD information (as described in the second embodiment providing a method for combining DIMD and CIIP) and an intra predictor generated by TIMD information (as described in the third embodiment providing a method for combining TIMD and CIIP) are blended.
[0138]
[0158] In some embodiments, N modes, a planar mode, and a DC mode, each having N maximum amplitude values in a DIMD histogram, are used as the input list for a TIMD method to derive intra-prediction modes used to generate intra-predictors in CIIP. The value of N can be any positive integer between 2 and 65.
[0139]
[0159] In some embodiments, the TIMD derivation mode is used to generate the intra predictor in CIIP only if the TIMD derivation mode is one of the N modes, the planar mode, and the DC mode, which have N maximum amplitude values in the TIMD histogram. Otherwise, the planar mode is used to generate the intra predictor in CIIP.
[0140]
[0160] For the sake of clarity, in this disclosure, the CIIP mode proposed in the above embodiment having a modified intra predictor generation method is referred to as CIIP_novel mode. It should be understood that CIIP_novel mode may include any modified forms such as CIIP_normal mode, CIIP_DIMD mode, CIIP_TIMD mode, and CIIP_TIMD fusion mode.
[0141]
[0161] In some embodiments, the weights of the interpreter and intrapredictor in CIIP can be modified.
[0142]
[0162] In the eighth embodiment, it is proposed that the weights of the inter-predictor and intra-predictor (wIntra,wInter) in the proposed CIIP_novel mode may be the same as or different from those in the current CIIP design.
[0143]
[0163] In some embodiments, the weights of the inter-predictor and intra-predictor in the proposed CIIP_novel mode are the same as in the current CIIP design. For example, if both the upper and left neighbors are intra-coded, (wIntra,wInter) is set to equal (3,1). If one of these blocks is intra-coded, their weights are identical, i.e., (2,2), otherwise the weights are set to equal (1,3).
[0144]
[0164] In some embodiments, the weights of the inter- and intra-predictors in the proposed CIIP_novel mode depend on the intra-prediction mode used to generate the intra-predictor. For example, if the intra-prediction mode is close to the horizontal or vertical mode (e.g., the absolute value of the difference between the intra-prediction mode index and the horizontal or vertical mode index is less than the threshold) and the width or height is not less than 4, then the multi-hypothesis prediction weights for the intra and inter-mode methods described above are used. For example, if the intra-prediction mode is one of the 67 modes of VVC, the threshold may range from 0 to 34. The threshold can be any valid positive integer.
[0145]
[0165] Figures 17A and 17B illustrate exemplary methods for dividing a coded block vertically and horizontally, respectively, according to some embodiments of the present disclosure. As shown in Figure 17A, in the near-horizontal mode (e.g., the intra-predictive mode index (i.e., angular mode index) is greater than or equal to 2 and less than 34), the current block is first divided vertically into four equally sized subblocks with subblock indices of 0 to 3 from left to right. As shown in Figure 17B, in the near-vertical mode (e.g., the intra-predictive mode index is greater than 34 and less than or equal to 66), the current block is first divided horizontally into four equally sized subblocks with subblock indices of 0 to 3 from top to bottom. In some embodiments, when the intra-predictive mode is extended, the near-horizontal mode refers to a mode where the angular mode index is greater than or equal to the angular mode index of the diagonal mode from bottom left to top right, and less than the angular mode index of the diagonal mode from top left to bottom right. Near-vertical mode refers to a mode in which the angular mode index is greater than or equal to the angular mode index of the diagonal mode from the top left to the bottom right, and less than or equal to the angular mode index of the diagonal mode from the top right to the bottom left.
[0146]
[0166] Different weights (wIntra,wInter) are used for different subblocks. For example, as shown in Table 1, four sets of weights (wIntra1,wInter1)=(6,2), (wIntra2,wInter2)=(5,3), (wIntra3,wInter3)=(3,5), and (wIntra4,wInter4)=(2,6) are used for four subblocks, respectively.
[0147] [Table 1]
[0148]
[0167] Otherwise, for example, if the intra-prediction mode index is 0 or 1, the same weights as the current CIIP design are used. That is, if both the upper and left adjacents are intra-coded, (wIntra,wInter) is set to equal (3,1). If one of these blocks is intra-coded, these weights are identical, i.e., (2,2). If neither the upper nor left adjacents are intra-coded, the weights are set to equal (1,3).
[0149]
[0168] Figure 18 shows another exemplary method for dividing a coded block in an angular intra-prediction mode according to some embodiments of the present disclosure. Referring to Figure 18, if the angular intra-prediction mode is close to the diagonal mode (for example, the absolute value of the difference between the angular intra-prediction mode index and the diagonal index (i.e., 34 in VVC) is less than a threshold), the current block is divided into four regions, with different weights (wIntra,wInter) used for the different regions. For example, four sets of weights (wIntra1,wInter1)=(6,2), (wIntra2,wInter2)=(5,3), (wIntra3,wInter3)=(3,5), and (wIntra4,wInter4)=(2,6) are used for the four regions, respectively.
[0150]
[0169] In some embodiments, a weight derivation method based on the coding mode of adjacent blocks and a weight derivation method based on sub-regions are combined. For example, if both the upper and left adjacent blocks are intracoded, four sets of weights (wIntra1,wInter1)=(7,1), (wIntra2,wInter2)=(6,2), (wIntra3,wInter3)=(4,4), and (wIntra4,wInter4)=(3,5) are used for the four regions. If only one of these blocks is intracoded, four sets of weights (wIntra1,wInter1)=(6,2), (wIntra2,wInter2)=(5,3), (wIntra3,wInter3)=(3,5), and (wIntra4,wInter4)=(2,6) are used for the four regions. If neither the upper adjacent nor the left adjacent is intracoded, then four sets of weights are used for the four regions: (wIntra1,wInter1)=(5,3), (wIntra2,wInter2)=(4,4), (wIntra3,wInter3)=(2,6), and (wIntra4,wInter4)=(1,7).
[0151]
[0170] In some embodiments, the intra-prediction mode used to generate the intra-predictor in CIIP is first converted to wide-angle mode in order to derive the weights (wIntra,wInter) by the method described above in the eighth embodiment.
[0152]
[0171] In some embodiments, the intra predictor is generated by blending several intra prediction modes. To derive (wIntra,wInter) by the method described above in the eighth embodiment, it is proposed to use a planar mode, or the intra prediction mode with the greatest weight, or the non-planar intra prediction mode with the greatest weight.
[0153]
[0172] In some embodiments, a method for signaling CIIP_novel mode is provided.
[0154]
[0173] In the ninth embodiment, the original CIIP mode, which blends the planar mode and the normal merge mode, is replaced by a proposed CIIP_novel mode that modifies the way intra predictors are generated.
[0155]
[0174] In the tenth embodiment, it is adaptively determined whether to replace the original CIIP mode (e.g., an intra predictor generated in planar mode only) with a proposed CIIP_novel mode (e.g., CIIP_normal mode, CIIP_DIMD mode, CIIP_TIMD mode, CIIP_TIMD fused mode, etc.).
[0156]
[0175] Figure 19 shows an exemplary flowchart of a method 1900 for determining a prediction mode according to some embodiments of the present disclosure. Method 1900 may be performed by a decoder (e.g., by process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of a device (e.g., device 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) may perform method 1900. In some embodiments, method 1900 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions such as program code executed by a computer (e.g., device 400 in Figure 4). Referring to Figure 19, method 1900 may include the following steps S1902 to S1906.
[0157]
[0176] In step S1902, the original CIIP mode or the new CIIP mode is determined for the current block. In some embodiments, this determination is based on the size of the current block. For example, the original CIIP mode is used for blocks with a size greater than a threshold, and the proposed new CIIP mode is used for blocks with a size less than or equal to the threshold. For example, the threshold may be equal to 1024.
[0158]
[0177] In some embodiments, this decision is based on the current block width and height. For example, for blocks with a width or height exceeding a threshold, the original CIIP mode is used; otherwise, the proposed CIIP_new mode is used. For example, for blocks with a long-side-to-short-side ratio exceeding a threshold, the original CIIP mode is used; otherwise, the proposed CIIP_new mode is used.
[0159]
[0178] In some embodiments, this decision is based on the relationship between the intra-prediction mode having the maximum amplitude value in the DIMD method and the intra-prediction mode having the minimum SATD value in the TIMD method. For example, if the two intra-prediction modes are similar (e.g., the absolute difference between the indices of the two modes is less than 4), the proposed CIIP_novel mode is used; otherwise, the original CIIP mode is used.
[0160]
[0179] In step S1904, a parameter (e.g., a flag) indicating whether the original CIIP mode or the new CIIP mode is being used is decoded.
[0161]
[0180] In step S1906, either the original CIIP mode or the CIIP_new mode is executed according to the parameters. In some embodiments, a further decision may be made to determine which of the CIIP_new modes is used based on the current block (for example, based on the size of the current block).
[0162]
[0181] In the eleventh embodiment, the proposed CIIP_novel_mode can be used as a new CIIP mode, and an explicit signaling method is used to determine which CIIP mode to use.
[0163]
[0182] For example, if only one new CIIP mode is added, a CIIP index of 0 to 2 will be signaled to determine which CIIP mode to use, as shown in Table 2.
[0164] [Table 2]
[0165]
[0183] For example, if two new CIIP modes are added, a CIIP index from 0 to 3 is signaled to determine which CIIP mode to use, as shown in Table 3, where CIIP_DIMD mode means the CIIP and DIMD combination described in the second embodiment, and CIIP_TIMD mode means the CIIP and TIMD combination described in the third embodiment.
[0166] [Table 3]
[0167]
[0184] In some embodiments, the disclosed CIIP mode can be used for encoding chroma samples.
[0168]
[0185] In current CIIP designs, the DM mode (intra-prediction mode for co-located luma blocks) is used to derive intra-predictors for chroma blocks coded with CIIP. This disclosure proposes that, for a chroma block, if co-located luma blocks are coded by the proposed CIIP_novel mode, the same intra-prediction mode derivation method is used for the current chroma block to generate intra-predictors. In some embodiments, the planar mode is always used for chroma blocks coded with the CIIP_novel mode to generate intra-predictors.
[0169]
[0186] In some embodiments, with respect to a block coded by the disclosed CIIP_novel mode, the prediction mode of the block can be stored as an inter-prediction mode, or planar mode, or intra-prediction mode used to generate an intra-predictor.
[0170]
[0187] In some embodiments, an intra-propagation prediction mode for CIIP is proposed.
[0171]
[0188] In some embodiments, it is proposed to modify the propagation intra-prediction mode of the block encoded by the original CIIP mode or the proposed CIIP_novel mode.
[0172]
[0189] In some embodiments, with respect to blocks coded by the original CIIP mode or the proposed CIIP_novel mode, the propagation intra-prediction mode is set to planar mode and used to construct the MPM list. For other intercoded blocks, the propagation intra-prediction mode is derived using motion vectors and reference pictures.
[0173]
[0190] In some embodiments, with respect to blocks coded by the original CIIP mode or the proposed CIIP_novel mode, the propagated intra-prediction mode is set to the intra-prediction mode used to generate the intra-predictor in the CIIP and used to construct the MPM list. In other intercoded blocks, the propagated intra-prediction mode is derived using motion vectors and reference pictures.
[0174]
[0191] In some embodiments, it is proposed to modify the propagation intra-prediction mode for blocks coded by CIIP_PDPC mode. For blocks coded by CIIP_PDPC mode, the propagation intra-prediction mode is set to planar mode and used for constructing the MPM list. For other intercoded blocks, the propagation intra-prediction mode is derived using motion vectors and reference pictures.
[0175]
[0192] In some embodiments, the two aforementioned methods for determining the propagation intra-prediction mode can be combined for blocks coded by the original CIIP mode or the proposed CIIP_novel mode and by the CIIP_PDPC mode. For example, in blocks coded by the original CIIP mode or the proposed CIIP_novel mode, the propagation intra-prediction mode is set to the intra-prediction mode used to generate the intra-predictor in CIIP. In blocks coded by the CIIP_PDPC mode, the propagation intra-prediction mode is set to the planar mode, and in other intercoded blocks, the propagation intra-prediction mode is derived using motion vectors and reference pictures.
[0176]
[0193] The embodiments described above can be combined in any combination.
[0177]
[0194] The embodiments can be further described using the following clauses: 1. A method for performing composite inter-intra prediction (CIIP), Determine that CIIP is valid for the target block. Determine the first intra-predictive mode of the target block using the template-based intra-mode derivation (TIMD) method. To generate an intra predictor for the target block in the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block. A method that includes this. 2. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the value of the sum of absolute transform differences (SATD) of target blocks associated with each of the multiple intra-prediction modes in the TIMD mode list, and The first intra-prediction mode for the target block is determined from among multiple intra-prediction modes to be the intra-prediction mode with the smallest SATD value. The method described in Clause 1, further including the following: 3. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the SATD value of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list. The intra-prediction mode with the smallest SATD value is determined from among multiple intra-prediction modes, and the intra-prediction mode with the smallest SATD value is mapped to a normal intra-prediction mode, and Determine the normal intra-prediction mode as the first intra-prediction mode for the target block. The method described in Clause 1, further including the following: 4. Obtaining the final predictor for the target block by weighting the intra predictor and the inter predictor of the target block is possible. In response to the first intra prediction mode being an angular mode, the intra weights and inter weights are determined based on the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block by the intra weight and inter weight, respectively. The method described in Clause 1, further including the following: 5. Determining intra weights and inter weights based on the first intra prediction mode is: The method involves dividing a target block into multiple subblocks, wherein the target block is divided vertically if the angular mode index of the first intra-prediction mode is less than a default value, or the target block is divided horizontally if the angular mode index of the first intra-prediction mode is greater than or equal to a default value, and Determining the sub-intra weights and sub-inter weights for each of the multiple sub-blocks. It further includes, The final predictor for the target block can be obtained by weighting the intra predictor and inter predictor of the target block. The process involves determining multiple sub-final predictors associated with multiple sub-blocks, wherein each of the multiple sub-final predictors is determined by weighting the intra predictor and inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, where the number of bits is obtained by the logarithm of the sum of the subintra weights and subinter weights. The method described in Clause 4, further including the following: 6. If the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the lower left to the upper right, and less than the angular mode index of the diagonal mode from the upper left to the lower right, the target block is divided vertically into four equally sized subblocks, or if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, and less than or equal to the angular mode index of the diagonal mode from the upper right to the lower left, the target block is divided horizontally into four equally sized subblocks. The four subblocks of equal size include the first subblock, the second subblock, the third subblock, and the fourth subblock, and the first, second, third, and fourth subblocks are arranged from left to right when the target block is divided vertically, or from top to bottom when the target block is divided horizontally, and The sub-intra weight and sub-inter weight of the first sub-block are 6 and 2, respectively. The sub-intra weight and sub-inter weight of the second sub-block are 5 and 3, respectively. The sub-intra weight and sub-inter weight of the third sub-block are 3 and 5, respectively, and The method according to Clause 5, wherein the sub-intra weight and sub-inter weight of the fourth sub-block are 2 and 6, respectively. 7. Determining the first intra-prediction mode of the target block using the TIMD method is: If the target block has a size below a threshold, the TIMD method is used to determine the first intra-prediction mode of the target block, and If the target block has a size exceeding the threshold, the first intra-prediction mode of the target block is determined to be the planar mode. The method described in any one of the clauses 1 to 6, further including the method described in any one of the clauses 1 to 6. 8. The threshold is equal to 1024, as described in Clause 7. 9. Set the planar mode as the propagation intra-prediction mode, and Construct a list of most likely modes (MPMs) of adjacent blocks to the target block in the propagation intra-prediction mode. The method described in any one of the clauses 1 to 8, further including the method described in any one of the clauses 1 to 8. 10. Equipment for performing composite inter- and intra-prediction (CIIP), Memory configured to store instructions, Includes one or more processors, and one or more processors Determine that CIIP is valid for the target block. Determine the first intra-predictive mode of the target block using the template-based intra-mode derivation (TIMD) method. To generate an intra predictor for the target block in the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block. A device configured to execute commands in order to cause another device to perform a certain action. 11. When determining the first intra-prediction mode of the target block using the TIMD method, one or more processors: Calculate the value of the sum of absolute transform differences (SATD) of target blocks associated with each of the multiple intra-prediction modes in the TIMD mode list, and The first intra-prediction mode for the target block is determined from among multiple intra-prediction modes to be the intra-prediction mode with the smallest SATD value. The equipment described in Clause 10, further configured to execute commands in order to cause the equipment to perform the following actions. 12. When determining the first intra-prediction mode of the target block using the TIMD method, one or more processors: Calculate the SATD value of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list. The intra-prediction mode with the smallest SATD value is determined from among multiple intra-prediction modes, and the intra-prediction mode with the smallest SATD value is mapped to a normal intra-prediction mode, and Determine the normal intra-prediction mode as the first intra-prediction mode for the target block. The equipment described in Clause 10, further configured to execute commands in order to cause the equipment to perform the following actions. 13. When obtaining the final predictor for a target block by weighting the intra predictor and inter predictor of the target block, one or more processors: In response to the first intra prediction mode being an angular mode, the intra weights and inter weights are determined based on the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block by the intra weight and inter weight, respectively. The equipment described in Clause 10, further configured to execute commands in order to cause the equipment to perform the following actions. 14. When determining intra weights and inter weights based on the first intra prediction mode, one or more processors: The method involves dividing a target block into multiple subblocks, wherein the target block is divided vertically if the angular mode index of the first intra-prediction mode is less than a default value, or the target block is divided horizontally if the angular mode index of the first intra-prediction mode is greater than or equal to a default value, and Determining the sub-intra weights and sub-inter weights for each of the multiple sub-blocks. It is further configured to execute commands to cause the device to perform the following actions: The final predictor for the target block can be obtained by weighting the intra predictor and inter predictor of the target block. The process involves determining multiple sub-final predictors associated with multiple sub-blocks, wherein each of the multiple sub-final predictors is determined by weighting the intra predictor and inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, where the number of bits is obtained by the logarithm of the sum of the subintra weights and subinter weights. The equipment described in Clause 13, further including the equipment described in Clause 13. 15. If the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the lower left to the upper right, and less than the angular mode index of the diagonal mode from the upper left to the lower right, the target block is divided vertically into four equally sized subblocks, or if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, and less than or equal to the angular mode index of the diagonal mode from the upper right to the lower left, The four subblocks of equal size include the first subblock, the second subblock, the third subblock, and the fourth subblock, and the first, second, third, and fourth subblocks are arranged from left to right when the target block is divided vertically, or from top to bottom when the target block is divided horizontally, and The sub-intra weight and sub-inter weight of the first sub-block are 6 and 2, respectively. The sub-intra weight and sub-inter weight of the second sub-block are 5 and 3, respectively. The sub-intra weight and sub-inter weight of the third sub-block are 3 and 5, respectively, and The sub-intra weight and sub-inter weight of the fourth sub-block are 2 and 6, respectively, as described in Clause 14. 16. When determining the first intra-prediction mode of the target block using the TIMD method, one or more processors: If the target block has a size below a threshold, the TIMD method is used to determine the first intra-prediction mode of the target block, and If the target block has a size exceeding the threshold, the first intra-prediction mode of the target block is determined to be the planar mode. A device as described in any one of clauses 10 to 15, further configured to execute commands in order to cause the device to perform the action. 17. The threshold is equal to 1024, as specified in Clause 16. 18. One or more processors, Set the planar mode as the propagation intra-prediction mode, and Construct a list of most likely modes (MPMs) of adjacent blocks to the target block in the propagation intra-prediction mode. A device as described in any one of clauses 10 to 17, further configured to execute commands in order to cause the device to perform the action. 19. A non-temporary computer-readable medium for storing a set of instructions, the set of instructions being executable by one or more processors of the device to initiate a method of performing composite inter-intra prediction (CIIP), the method being: Determine that CIIP is valid for the target block. Determine the first intra-predictive mode of the target block using the template-based intra-mode derivation (TIMD) method. To generate an intra predictor for the target block in the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block. Non-temporary computer-readable media, including [specific examples of such media]. 20. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the value of the sum of absolute transform differences (SATD) of target blocks associated with each of the multiple intra-prediction modes in the TIMD mode list, and The first intra-prediction mode for the target block is determined from among multiple intra-prediction modes to be the intra-prediction mode with the smallest SATD value. Non-temporary computer-readable media as defined in Clause 19, further including the above. 21. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the SATD value of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list. The intra-prediction mode with the smallest SATD value is determined from among multiple intra-prediction modes, and the intra-prediction mode with the smallest SATD value is mapped to a normal intra-prediction mode, and Determine the normal intra-prediction mode as the first intra-prediction mode for the target block. Non-temporary computer-readable media as defined in Clause 19, further including the above. 22. Obtaining the final predictor for a target block by weighting the intra predictor and the inter predictor of the target block is possible. In response to the first intra prediction mode being an angular mode, the intra weights and inter weights are determined based on the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block by the intra weight and inter weight, respectively. Non-temporary computer-readable media as defined in Clause 19, further including the above. 23. Determining intra weights and inter weights based on the first intra prediction mode is: The method involves dividing a target block into multiple subblocks, wherein the target block is divided vertically if the angular mode index of the first intra-prediction mode is less than a default value, or the target block is divided horizontally if the angular mode index of the first intra-prediction mode is greater than or equal to a default value, and Determining the sub-intra weights and sub-inter weights for each of the multiple sub-blocks. It further includes, The final predictor for the target block can be obtained by weighting the intra predictor and inter predictor of the target block. The process involves determining multiple sub-final predictors associated with multiple sub-blocks, wherein each of the multiple sub-final predictors is determined by weighting the intra predictor and inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, where the number of bits is obtained by the logarithm of the sum of the subintra weights and subinter weights. Non-temporary computer-readable media as defined in Clause 22, further including the above. 24. If the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the lower left to the upper right, and less than the angular mode index of the diagonal mode from the upper left to the lower right, the target block is divided vertically into four equally sized subblocks, or if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, and less than or equal to the angular mode index of the diagonal mode from the upper right to the lower left, The four subblocks of equal size include the first subblock, the second subblock, the third subblock, and the fourth subblock, and the first, second, third, and fourth subblocks are arranged from left to right when the target block is divided vertically, or from top to bottom when the target block is divided horizontally, and The sub-intra weight and sub-inter weight of the first sub-block are 6 and 2, respectively. The sub-intra weight and sub-inter weight of the second sub-block are 5 and 3, respectively. The sub-intra weight and sub-inter weight of the third sub-block are 3 and 5, respectively, and The non-temporary computer-readable medium described in Clause 23 has sub-intra weights and sub-inter weights of 2 and 6, respectively, for the fourth sub-block. 25. Determining the first intra-prediction mode of the target block using the TIMD method is: If the target block has a size below a threshold, the TIMD method is used to determine the first intra-prediction mode of the target block, and If the target block has a size exceeding the threshold, the first intra-prediction mode of the target block is determined to be the planar mode. Non-temporary computer-readable media as described in any one of clauses 19 to 24, further including the above. 26. The threshold is equal to 1024, for non-temporary computer-readable media as defined in Clause 25. 27. The method is, Set the planar mode as the propagation intra-prediction mode, and Construct a list of most likely modes (MPMs) of adjacent blocks to the target block in the propagation intra-prediction mode. Non-temporary computer-readable media as described in any one of clauses 19 to 26, further including the above. 28. A method for performing composite inter- and intra-prediction (CIIP), Determine the first intra-predictive mode of the target block using the template-based intra-mode derivation (TIMD) method. In the first intra prediction mode, generate intra predictors for the target block. The final predictor for the target block is obtained by weighting the intra predictor and the inter predictor of the target block, and Signaling a flag indicating that CIIP is enabled and an index indicating that the TIMD method is being used to determine the first intra-prediction mode for the target block. A method that includes this. 29. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the value of the sum of absolute transform differences (SATD) of target blocks associated with each of the multiple intra-prediction modes in the TIMD mode list, and The first intra-prediction mode for the target block is determined from among multiple intra-prediction modes to be the intra-prediction mode with the smallest SATD value. The method described in Article 28, further including the method described in Article 28. 30. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the SATD value of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list. The intra-prediction mode with the smallest SATD value is determined from among multiple intra-prediction modes, and the intra-prediction mode with the smallest SATD value is mapped to a normal intra-prediction mode, and Determine the normal intra-prediction mode as the first intra-prediction mode for the target block. The method described in Article 28, further including the method described in Article 28. 31. Obtaining the final predictor for a target block by weighting the intra predictor and the inter predictor of the target block is possible. In response to the first intra prediction mode being an angular mode, the intra weights and inter weights are determined based on the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block by the intra weight and inter weight, respectively. The method described in Article 28, further including the method described in Article 28. 32. Determining intra weights and inter weights based on the first intra prediction mode is: The method involves dividing a target block into multiple subblocks, wherein the target block is divided vertically if the angular mode index of the first intra-prediction mode is less than a default value, or the target block is divided horizontally if the angular mode index of the first intra-prediction mode is greater than or equal to a default value, and Determining the sub-intra weights and sub-inter weights for each of the multiple sub-blocks. It further includes, The final predictor for the target block can be obtained by weighting the intra predictor and inter predictor of the target block. The process involves determining multiple sub-final predictors associated with multiple sub-blocks, wherein each of the multiple sub-final predictors is determined by weighting the intra predictor and inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, where the number of bits is obtained by the logarithm of the sum of the subintra weights and subinter weights. The method described in Clause 31, further including the method described in Clause 31. 33. If the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the lower left to the upper right, and less than the angular mode index of the diagonal mode from the upper left to the lower right, the target block is divided vertically into four equally sized subblocks, or if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, and less than or equal to the angular mode index of the diagonal mode from the upper right to the lower left, The four subblocks of equal size include the first subblock, the second subblock, the third subblock, and the fourth subblock, and the first, second, third, and fourth subblocks are arranged from left to right when the target block is divided vertically, or from top to bottom when the target block is divided horizontally, and The sub-intra weight and sub-inter weight of the first sub-block are 6 and 2, respectively. The sub-intra weight and sub-inter weight of the second sub-block are 5 and 3, respectively. The sub-intra weight and sub-inter weight of the third sub-block are 3 and 5, respectively, and The method according to clause 32, wherein the sub-intra weight and sub-inter weight of the fourth sub-block are 2 and 6, respectively. 34. Determining the first intra-prediction mode of the target block using the TIMD method is: If the target block has a size below a threshold, the TIMD method is used to determine the first intra-prediction mode of the target block, and If the target block has a size exceeding the threshold, the first intra-prediction mode of the target block is determined to be the planar mode. The method described in any one of the clauses 28 to 33, further including the method described in any one of the clauses 28 to 33. 35. The threshold is equal to 1024, as described in Clause 34. 36. Setting the planar mode as the propagation intra-prediction mode, and Construct a list of most likely modes (MPMs) of adjacent blocks to the target block in the propagation intra-prediction mode. The method described in any one of the clauses 28 to 35, further including the method described in any one of the clauses 28 to 35. 37. A non-temporary computer-readable medium for storing a bitstream, wherein the bitstream includes a flag and a first index related to encoded video data, the flag indicating that inter- and intra-prediction (CIIP) is used for encoded video data, the first index indicating a template-based intra-mode derivation (TIMD) method used for CIIP, and the flag and index are, The TIMD method is used to determine the first intra-prediction mode of the target block. To generate intra predictors for the target block in intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and inter predictor of the target block. A non-temporary computer-readable medium that allows a decoder to perform the decoding process. 38. The bitstream further includes an angular mode index associated with the video data, and the angular mode index is, The method involves dividing a target block into multiple subblocks, wherein the target block is divided vertically if the angular mode index of the first intra-prediction mode is less than a default value, or the target block is divided horizontally if the angular mode index of the first intra-prediction mode is greater than or equal to a default value, and Determining the sub-intra weights and sub-inter weights for each of the multiple sub-blocks, The process involves determining multiple sub-final predictors associated with multiple sub-blocks, wherein each of the multiple sub-final predictors is determined by weighting the intra predictor and inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, where the number of bits is obtained by the logarithm of the sum of the subintra weights and subinter weights. A non-temporary computer-readable medium as described in Clause 37, which allows the decoder to perform the decoding.
[0178]
[0195] In some embodiments, non-temporary computer-readable storage media containing instructions are also provided, which can be executed by devices for performing the above-described methods (such as disclosed encoders and decoders). Common forms of non-temporary media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH®-EPROMs or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges and networked versions thereof. Devices may include one or more processors (CPUs), input / output interfaces, network interfaces and / or memory.
[0179]
[0196] It should be noted that relational terms such as “First” and “Second” in this specification are used solely to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between those entities or operations. Furthermore, “Include,” “Have,” “Contain,” and “Incorporate,” as well as other similar terms, are intended to be equivalent in meaning and are intended to be non-restrictive in that the item following any one of these terms is not intended to be an exhaustive list of such items, nor is it intended to be limited to only the items listed.
[0180]
[0197] When used herein, unless otherwise specified, the word “or” encompasses all possible combinations, except in cases where it is not feasible. For example, if it is stated that a database may contain A or B, then unless otherwise specified or it is not feasible, the database may contain A or B or A and B. As a second example, if it is stated that a database may contain A, B or C, then unless otherwise specified or it is not feasible, the database may contain A, B, C, A and B, A and C, B and C, A and B and C.
[0181]
[0198] It will be understood that the embodiments described above can be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, the software can be stored in the computer-readable medium described above. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. It will also be understood by those skilled in the art that multiple of the above modules / units can be combined into a single module / unit, and each of the above modules / units can be further divided into multiple submodules / subunits.
[0182]
[0199] This specification has described embodiments with respect to numerous specific details that may vary depending on the implementation. Specific adaptations and modifications may be made to the embodiments described. Other embodiments may become apparent to those skilled in the art by examining this specification and practicing the invention disclosed herein. This specification and examples are provided for illustrative purposes only, and the true scope and spirit of the invention are intended to be shown by the appended claims. The order of steps shown in the figures is for illustrative purposes only and is not intended to limit the order of steps to any particular set. Therefore, those skilled in the art will understand that these steps can be performed in different orders while implementing the same method.
[0183]
[0200] Exemplary embodiments have been disclosed in the drawings and this specification. However, many modifications and alterations can be made to those embodiments. Accordingly, although specific terms have been used, they are used only in a general and descriptive sense and are not intended to be limiting.
Claims
1. A method for performing composite inter- and intra-prediction (CIIP), To determine that the CIIP is valid for the target block, A template-based intra-mode derivation (TIMD) method is used to determine the first intra-prediction mode of the target block. To generate an intra predictor for the target block in the first intra prediction mode, and The final predictor for the target block is obtained by weighting and averaging the intra predictor and inter predictor of the target block. Includes, Determining the first intra-prediction mode of the target block using the TIMD method is: If the target block has a size less than or equal to a threshold, the TIMD method is used to determine the first intra prediction mode of the target block, and If the target block has a size exceeding the threshold, the first intra prediction mode of the target block is determined to be the planar mode. Methods that further include the above.
2. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the value of the sum of absolute transformation differences (SATD) of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list, and The first intra-prediction mode for the target block is determined from among the plurality of intra-prediction modes to have the smallest SATD value. The method according to claim 1, further comprising:
3. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the SATD value of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list. The intra-prediction mode having the smallest SATD value is determined from among the plurality of intra-prediction modes, and the intra-prediction mode having the smallest SATD value is mapped to a normal intra-prediction mode, and The normal intra prediction mode is determined as the first intra prediction mode for the target block. The method according to claim 1, further comprising:
4. Obtaining the final predictor of the target block by weighting the intra predictor and the inter predictor of the target block is: In response to the fact that the first intra prediction mode is an angular mode, the intra weights and inter weights are determined based on the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and the inter predictor of the target block by the intra weight and inter weight, respectively. The method according to claim 1, further comprising:
5. Determining the intra weights and inter weights based on the first intra prediction mode is: Dividing the target block into a plurality of subblocks, wherein the target block is divided vertically when the angular mode index of the first intra prediction mode is less than a default value, or the target block is divided horizontally when the angular mode index of the first intra prediction mode is equal to or greater than the default value, and Determine the sub-intra weight and sub-inter weight for each of the aforementioned subblocks. It further includes, Obtaining the final predictor of the target block by weighting the intra predictor and the inter predictor of the target block is: Determining a plurality of sub-final predictors associated with each of the plurality of sub-blocks, wherein each of the plurality of sub-final predictors is determined by weighting the intra predictor and the inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of the multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, wherein the number of bits is obtained by the logarithm of the sum of the subintra weights and the subinter weights. The method according to claim 4, further comprising:
6. The target block is vertically divided into four equally sized subblocks if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the lower left to the upper right, and less than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, or if the target block is horizontally divided into four equally sized subblocks if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, and less than or equal to the angular mode index of the diagonal mode from the upper right to the lower left. The four subblocks of equal size include a first subblock, a second subblock, a third subblock, and a fourth subblock, the first, second, third, and fourth subblocks being arranged from left to right when the target block is divided vertically, or from top to bottom when the target block is divided horizontally, and The sub-intra weight and sub-inter weight of the first sub-block are 6 and 2, respectively. The sub-intra weight and sub-inter weight of the second sub-block are 5 and 3, respectively. The subintra weight and subinter weight of the third subblock are 3 and 5, respectively, and The method according to claim 5, wherein the sub-intra weight and sub-inter weight of the fourth sub-block are 2 and 6, respectively.
7. The method according to claim 1, wherein the threshold is equal to 1024.
8. Setting the planar mode as the propagation intra-prediction mode, and Constructing a list of most likely modes (MPMs) of adjacent blocks to the target block in the aforementioned propagation intra-prediction mode. The method according to claim 1, further comprising:
9. A method for performing composite inter- and intra-prediction (CIIP), Determine the first intra-prediction mode of the target block using a template-based intra-mode derivation (TIMD) method. To generate an intra predictor for the target block in the first intra prediction mode, The final predictor for the target block is obtained by weighting the intra predictor and the inter predictor of the target block, and Signaling a flag indicating that the CIIP is valid and an index indicating that the TIMD method is used to determine the first intra-prediction mode of the target block. A method that includes this.
10. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the value of the sum of absolute transformation differences (SATD) of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list, and The first intra-prediction mode for the target block is determined from among the plurality of intra-prediction modes to have the smallest SATD value. The method according to claim 9, further comprising:
11. Determining the first intra-prediction mode of the target block using the TIMD method is: Calculate the SATD value of the target block associated with each of the multiple intra-prediction modes in the TIMD mode list. The intra-prediction mode having the smallest SATD value is determined from among the plurality of intra-prediction modes, and the intra-prediction mode having the smallest SATD value is mapped to a normal intra-prediction mode, and The normal intra prediction mode is determined as the first intra prediction mode for the target block. The method according to claim 9, further comprising:
12. Obtaining the final predictor of the target block by weighting the intra predictor and the inter predictor of the target block is: In response to the fact that the first intra prediction mode is an angular mode, the intra weights and inter weights are determined based on the first intra prediction mode, and The final predictor for the target block is obtained by weighting the intra predictor and the inter predictor of the target block by the intra weight and inter weight, respectively. The method according to claim 9, further comprising:
13. Determining the intra weights and inter weights based on the first intra prediction mode is: Dividing the target block into a plurality of subblocks, wherein the target block is divided vertically when the angular mode index of the first intra prediction mode is less than a default value, or the target block is divided horizontally when the angular mode index of the first intra prediction mode is equal to or greater than the default value, and Determine the sub-intra weight and sub-inter weight for each of the aforementioned subblocks. It further includes, Obtaining the final predictor of the target block by weighting the intra predictor and the inter predictor of the target block is: Determining a plurality of sub-final predictors associated with each of the plurality of sub-blocks, wherein each of the plurality of sub-final predictors is determined by weighting the intra predictor and inter predictor by the sub-intra weight and sub-inter weight of the respective sub-block. Determining the sum of the multiple sub-final predictors, and The final predictor is obtained by right-shifting by a certain number of bits, wherein the number of bits is obtained by the logarithm of the sum of the subintra weights and the subinter weights. The method according to claim 12, further comprising:
14. The target block is vertically divided into four equally sized subblocks if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the lower left to the upper right, and less than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, or if the target block is horizontally divided into four equally sized subblocks if the angular mode index of the first intra-prediction mode is greater than or equal to the angular mode index of the diagonal mode from the upper left to the lower right, and less than or equal to the angular mode index of the diagonal mode from the upper right to the lower left. The four subblocks of equal size include a first subblock, a second subblock, a third subblock, and a fourth subblock, the first, second, third, and fourth subblocks being arranged from left to right when the target block is divided vertically, or from top to bottom when the target block is divided horizontally, and The sub-intra weight and sub-inter weight of the first sub-block are 6 and 2, respectively. The sub-intra weight and sub-inter weight of the second sub-block are 5 and 3, respectively. The subintra weight and subinter weight of the third subblock are 3 and 5, respectively, and The method according to claim 13, wherein the sub-intra weight and sub-inter weight of the fourth sub-block are 2 and 6, respectively.
15. Determining the first intra-prediction mode of the target block using the TIMD method is: If the target block has a size less than or equal to a threshold, the TIMD method is used to determine the first intra prediction mode of the target block, and If the target block has a size exceeding a threshold, the first intra prediction mode of the target block is determined to be the planar mode. The method according to claim 9, further comprising:
16. The method according to claim 15, wherein the threshold is equal to 1024.
17. Setting the planar mode as the propagation intra-prediction mode, and Constructing a list of most likely modes (MPMs) of adjacent blocks to the target block in the aforementioned propagation intra-prediction mode. The method according to claim 9, further comprising:
18. A method for storing the bitstream of a video sequence, Receiving a video sequence, Encoding one or more pictures of the aforementioned video sequence, To generate a bitstream, and The bitstream is stored in a non-temporary computer-readable storage medium. Including the above encoding, Determine that composite inter- and intra-prediction (CIIP) is effective for the target block. Determine the first intra-prediction mode of the target block using a template-based intra-mode derivation (TIMD) method. To generate an intra predictor for the target block in the first intra prediction mode, and The final predictor for the target block is obtained by weighting and averaging the intra predictor and inter predictor of the target block. Includes, Determining the first intra-prediction mode of the target block using the TIMD method is: If the target block has a size less than or equal to a threshold, the TIMD method is used to determine the first intra prediction mode of the target block, and If the target block has a size exceeding a threshold, the first intra prediction mode of the target block is determined to be the planar mode. Methods that further include the above.
Citation Information
Patent Citations
Combined inter and intra prediction
WO2020142279A1