System and method for improving hybrid inter and intra prediction

By simplifying CIIP through disabling BDOF, converting double to single prediction, and harmonizing intra mode usage, the method addresses computational and memory challenges, enhancing decoding throughput and coding efficiency in video coding technologies like VVC.

JP2025111753AActive Publication Date: 2025-07-30BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025075975
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-01-09
Filing Date
2025-05-01
Publication Date
2025-07-30
Estimated Expiration
2040-01-09

AI Technical Summary

Technical Problem

Existing video coding technologies, such as VVC, face challenges in efficiently combining inter and intra prediction methods due to high computational complexity, memory bandwidth requirements, and non-unified design of merge-related modes, which hinder real-time decoding and coding efficiency.

Method used

The proposed method simplifies the composite inter and intra prediction (CIIP) process by disabling bidirectional optical flow (BDOF) for inter prediction, converting double to single prediction for memory efficiency, and harmonizing intra mode usage in MPM candidate lists, enabling CIIP only when CUs are singly predicted, and optimizing flag signaling.

Benefits of technology

This approach reduces computational complexity and memory bandwidth, improving decoding throughput and coding efficiency by enhancing the accuracy of prediction residuals, thus facilitating real-time decoding and better video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111753000001_ABST
    Figure 2025111753000001_ABST
Patent Text Reader

Abstract

To provide a method and an apparatus for hybrid inter- and intra-prediction (CIIP) for video coding.SOLUTION: A method includes obtaining a first reference image and a second reference image associated with a current prediction block, generating a first prediction L0 on the basis of a first motion vector MV0 from the current prediction block to a reference block in the first reference image, generating a second prediction L1 on the basis of a second motion vector MV1 from the current prediction block to a reference block in the second reference image, determining whether a bidirectional optical flow (BDOF) operation is applied, and calculating a dual prediction of the current prediction block on the basis of the first prediction L0 and the second prediction L1 and the first gradient value and the second gradient value.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority based on Provisional Application No. 62 / 790,421, filed on Jan. 9, 2019, and incorporates all of its content herein by reference. All of its content is incorporated herein by reference.

[0002] This application relates to video coding and compression. More specifically, this application relates to methods and devices for composite inter and intra prediction (CIIP) methods for video decoding.

Background Art

[0003] To compress video data, various video coding techniques can be used. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (H.265 / HEVC), High Definition Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, and the like. Video coding generally utilizes prediction methods (e.g., inter prediction, intra prediction, etc.) that exploit the redundancy present in video images or sequences. An important goal of video coding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality.

Summary of the Invention

[0004] Examples of the present disclosure provide a method for improving the efficiency of syntax signaling of merge-related modes.

[0005] JPEG2025111753000002.jpg106161

[0006] According to a second aspect of the present disclosure, within a reference image list associated with a current prediction block obtaining a reference image, generating an inter prediction based on a first motion vector from the current image to the first reference image obtaining an intra prediction mode associated with the current prediction block, generating an intra prediction of the current prediction block based on the intra prediction averaging the inter prediction and the intra prediction to generate a final prediction of the current prediction block, and determining whether the current prediction block is treated as either an inter mode or an intra mode relative to a most probable mode (MPM)-based intra mode prediction a video coding method that compromises. averaging the inter prediction and the intra prediction to generate a final prediction of the current prediction block, and determining whether the current prediction block is treated as either an inter mode or an intra mode relative to a most probable mode (MPM)-based intra mode prediction block is treated as either an inter mode or an intra mode relative to a most probable mode (MPM)-based intra mode prediction to determine whether the current prediction block is treated as either an inter mode or an intra mode relative to a most probable mode (MPM)-based intra mode prediction a video coding method that compromises.

[0007] JPEG2025111753000003.jpg87161

[0008] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing instructions is provided. When executed by one or more processors, the instructions cause the one or more processors to obtain a reference image within a reference image list associated with a current prediction block, generate an inter prediction based on a first motion vector from the current image to the first reference image obtain an intra prediction mode associated with the current prediction block, generate an intra prediction of the current prediction block based on the intra prediction obtain a reference image within a reference image list associated with a current prediction block, generate an inter prediction based on a first motion vector from the current image to the first reference image generate an inter prediction based on a first motion vector from the current image to the first reference image, obtain an intra prediction mode associated with the current prediction block, generate an intra prediction of the current prediction block based on the intra prediction obtain an intra prediction mode associated with the current prediction block, generate an intra prediction of the current prediction block based on the intra prediction, average the inter prediction and the intra prediction to generate a final prediction of the current prediction block generate an intra prediction of the current prediction block based on the intra prediction, average the inter prediction and the intra prediction to generate a final prediction of the current prediction block average the inter prediction and the intra prediction to generate a final prediction of the current prediction block Generating a final prediction and determining whether the current prediction block is to be treated as either an inter-mode or an intra-mode for the most probable mode (MP M)-based intra-mode prediction, and causing the computing device to perform operations including these. It should be understood that both the foregoing general description and the following detailed description are merely examples and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure. A block diagram of an encoder according to an example of the present disclosure. A block diagram of a decoder according to an example of the present disclosure.

[0010] A flowchart showing a method for generating composite inter and intra prediction (CIIP) according to an example of the present disclosure. A flowchart showing a method for generating CIIP according to an example of the present disclosure.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 5E

Figure 6A

Figure 6B

Figure 6C

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0011] Here, examples of the present disclosure are referred to in detail, and the examples are shown in the accompanying drawings. The following description is otherwise Unless otherwise noted, the same reference numerals in different drawings refer to the same or similar elements. The attached drawings are referred to. The embodiments described in the following description of examples of the present disclosure do not represent all embodiments that are consistent with the present disclosure. Instead, they are merely examples of devices and methods that are consistent with aspects related to the present disclosure as described in the scope of the appended claims. Reference is made to the figures. The embodiments described in the following description of examples of the present disclosure do not represent all embodiments that are consistent with the present disclosure. Instead, they are merely examples of devices and methods that are consistent with aspects related to the present disclosure as described in the scope of the appended claims. Reference is made to the figures. The embodiments described in the following description of examples of the present disclosure do not represent all embodiments that are consistent with the present disclosure. Instead, they are merely examples of devices and methods that are consistent with aspects related to the present disclosure as described in the scope of the appended claims.

[0012] The terms used in the present disclosure are for the sole purpose of describing particular embodiments and are not intended to limit the present disclosure. As used in the present disclosure and the scope of the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein means any one or all possible combinations of one or more of the related listed items and is also understood to be intended to include them. The terms used in the present disclosure are for the sole purpose of describing particular embodiments and are not intended to limit the present disclosure. As used in the present disclosure and the scope of the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein means any one or all possible combinations of one or more of the related listed items and is also understood to be intended to include them. Here, various information can be described using terms such as "first," "second," "third," etc., but it should be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information can be referred to as the second information, and similarly, the second information can also be referred to as the first information. Here, when the term "if" is used, it can be understood to mean "when" or "upon" or "depending on the judgment" depending on the context.

[0013] Here, various information can be described using terms such as "first," "second," "third," etc., but it should be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information can be referred to as the second information, and similarly, the second information can also be referred to as the first information. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information can be referred to as the second information, and similarly, the second information can also be referred to as the first information. Here, when the term "if" is used, it can be understood to mean "when" or "upon" or "depending on the judgment" depending on the context. Here, when the term "if" is used, it can be understood to mean "when" or "upon" or "depending on the judgment" depending on the context. Here, when the term "if" is used, it can be understood to mean "when" or "upon" or "depending on the judgment" depending on the context.

[0014] The first version of the HEVC standard was completed in October 2013, which is the previous generation of video... Compared with the decoding standard H.264 / MPEG AVC, it provides about 50% bit rate savings or equivalent perceptual quality. The HEVC standard provides significant coding improvements over its predecessors, but there is evidence that by adding coding tools to HEVC, excellent coding efficiency can be achieved. Based on this, both VCEG and MPEG started an investigation of new coding technologies for future video coding standardization. Advanced technologies that enable significant improvements in coding efficiency were the subject of important research. In October 2015, one Joint Video Exploration Team (JVET) was formed by ITU-T VCEG and ISO / IEC MPEG. A reference software called the Joint Exploration Model (JEM) was maintained by JVET by integrating some additional coding tools on top of the HEVC test model (HM).

[0015] In October 2017, a Call for Proposals (CfP) regarding video compression with features beyond HEVC was issued by ITU-T and ISO / IEC. In April 2018, at the 10th JVET meeting, 23 CfP responses were received and evaluated, demonstrating a compression efficiency gain of approximately 40% over HEVC. Based on such evaluation results, JVET launched a new project to develop a new generation of video coding standard called Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. ​

[0016] Like HEVC, VVC is a block-based hybrid video coding framework. Figure 1 (explained below) shows a typical block-based A block diagram of a hybrid video coding system is given. The input video signal is divided into blocks (codes). In VTM-1.0, the C U can be up to 128x128 pixels, but is based only on the quadtree. Unlike HEVC, which divides blocks based on quad / binary / ternary, VVC To accommodate various local characteristics based on the tree, one coding tree unit is used. A CTU is divided into CUs. The concept of unit types has been removed, i.e., CU, prediction unit (PU) and transform unit The separation of the CU (TU) no longer exists in VVC, instead each CU always has an additional party It is used as the basic unit for both prediction and transformation without any partitioning. Multi-type tree structure In this example, a CTU is first partitioned by a quad tree structure. Then, each quad Tree leaf nodes can be further partitioned into binary and ternary tree structures. As shown in Figures 5A, 5B, 5C, 5D, 5E (described below), ,respectively, quaternary partitioning, horizontal binary partitioning, and vertical binary partitioning. partitioning, horizontal three-way partitioning, and vertical three-way partitioning. There is a split type.

[0017] In FIG. 1 (described below), spatial and / or temporal prediction can be performed. Spatial prediction (or "intra prediction") uses samples of already-coded adjacent blocks (referred to as reference samples) in the same video image / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as "inter prediction" or "motion-compensated prediction") uses the reconstructed pixels from already-coded video images to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a particular CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, when multiple reference images are supported, one reference image index is additionally transmitted. This is used to identify from which reference image in the reference image store the temporal prediction signal comes. After spatial prediction and / or temporal prediction, the mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate-distortion optimization method. Next, the prediction block is subtracted from the current video block, and the prediction residue is decorrelated using transformation and quantization. The quantized residual coefficients are inverse quantized and inverse transformed to form the reconstructed residue, which is then added to the prediction block to form the reconstructed signal of the CU. Further in-loop filtering, such as deblocking filter, sample adaptive offset (SAO), adaptive loop filter (ALF), can be applied to the reconstructed CU before being placed in the reference image store for use in coding future video blocks.

[0018] <- ​​​​​​​​​​​​​​​​​ To form a video bit stream, the coding mode (intra or inter), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit, further compressed and packed to form a bit stream.

[0019] FIG. 2 (described below) shows a general block diagram of a block-based video decoder. The video bit stream is first entropy decoded by an entropy decoding unit. The coding mode and prediction information are sent to either a spatial prediction unit (when intra coding) or a temporal prediction unit (when inter coding) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct a residual block. Next, the prediction block and the residual block are added together. The reconstructed block can further pass through in-loop filtering before being stored in a reference image store. Next, the reconstructed video in the reference image store is sent to drive a display device and is also used to predict future video blocks.

[0020] FIG. 1 shows a typical encoder 100. The encoder 100 includes a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, image buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory ​​​​​​​​​​​​​​​124, in-loop filter 122, entropy coding 138, and bit stream 144. have.

[0021] Figure 2 shows a block diagram of a typical decoder 200. The decoder 200 includes a bit stream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 2 18, intra / inter mode selection 220, intra prediction 222, memory 230, in loop filter 228, motion compensation 224, picture buffer 226, prediction related information 234, and video output 232.

[0022] Figure 3 shows an exemplary method 300 for generating composite inter and intra prediction (CIIP) according to the present disclosure. shown.

[0023] In step 310, a first reference image and a second reference image associated with the current prediction block are obtained. Here, the first reference image is before the current image in the display order, and the second reference image is after the current image in the display order.

[0024] In step 312, a first prediction L0 is obtained based on a first motion vector MV0 from the current prediction block to a reference block in the first reference image.

[0025] In step 314, a second prediction L1 is obtained based on a second motion vector MV1 from the current prediction block to a reference block in the second reference image.

[0026] JPEG2025111753000004.jpg55161

[0027] JPEG2025111753000005.jpg32161 ​​​

[0028] Figure 4 shows an exemplary method for generating a CIIP according to the present disclosure. For example, the method includes single prediction-based inter-prediction and MPM-based intra-prediction to generate a CIIP.

[0029] In step 410, obtain reference images in the reference image list associated with the current prediction block.

[0030] In step 412, generate inter-prediction based on the first motion vector from the current image to the first reference image.

[0031] In step 414, obtain the intra-prediction mode associated with the current prediction block.

[0032] In step 416, generate the intra-prediction of the current prediction block based on the intra-prediction.

[0033] In step 418, generate the final prediction of the current prediction block by averaging the inter-prediction and the intra-prediction.

[0034] In step 420, determine whether the current prediction block is to be treated as either an inter-mode or an intra-mode for the most probable mode (MPM) -based intra-mode prediction.

[0035] Figure 5A shows a diagram illustrating a block quad-partition in a multi-type tree structure according to an example of the present disclosure.

[0036] Figure 5B shows a diagram illustrating a block vertical binary partition in a multi-type tree structure according to an example of the present disclosure. Shows a diagram illustrating a partition.

[0037] Figure 5C shows a diagram illustrating a block horizontal binary partition in a multi-type tree structure according to an example of the present disclosure. Shows a diagram illustrating a partition.

[0038] Figure 5D shows a diagram illustrating a block vertical ternary partition in a multi-type tree structure according to an example of the present disclosure. Shows a diagram illustrating a partition.

[0039] Figure 5E shows a diagram illustrating a block horizontal ternary partition in a multi-type tree structure according to an example of the present disclosure. Shows a diagram illustrating a partition.

[0040] Compound inter and intra prediction As shown in FIGS. 1 and 2, the inter and intra prediction methods are used in a hybrid video coding scheme. Here, each PU is allowed to select either inter prediction or intra prediction in either the temporal domain or the spatial domain to utilize the correlation, but not both. However, as pointed out in the prior art, the residual signals generated by the inter prediction blocks and the intra prediction blocks may have very different characteristics from each other. Therefore, if the two types of predictions can be efficiently combined, another accurate prediction can be expected to reduce the energy of the prediction residual and improve the coding efficiency. Furthermore, in natural video content, the motion of moving objects may become complex. For example, there may be regions containing both old content (e.g., objects included in previously coded images) and new content (e.g., objects excluded in previously coded images). Here, each PU is allowed to select either inter prediction or intra prediction in either the temporal domain or the spatial domain to utilize the correlation, but not both. Here, each PU is allowed to select either inter prediction or intra prediction in either the temporal domain or the spatial domain to utilize the correlation, but not both. However, as pointed out in the prior art, the residual signals generated by the inter prediction blocks and the intra prediction blocks may have very different characteristics from each other. However, as pointed out in the prior art, the residual signals generated by the inter prediction blocks and the intra prediction blocks may have very different characteristics from each other. Therefore, if the two types of predictions can be efficiently combined, another accurate prediction can be expected to reduce the energy of the prediction residual and improve the coding efficiency. Therefore, if the two types of predictions can be efficiently combined, another accurate prediction can be expected to reduce the energy of the prediction residual and improve the coding efficiency. Furthermore, in natural video content, the motion of moving objects may become complex. For example, there may be regions containing both old content (e.g., objects included in previously coded images) and new content (e.g., objects excluded in previously coded images). For example, there may be regions containing both old content (e.g., objects included in previously coded images) and new content (e.g., objects excluded in previously coded images). For example, there may be regions containing both old content (e.g., objects included in previously coded images) and new content (e.g., objects excluded in previously coded images). In such a scenario, neither inter prediction nor intra prediction can provide an accurate prediction of one of the current blocks.

[0041] To further improve prediction efficiency, the VVC standard adopts combined inter and intra prediction (CIIP) that combines the intra prediction and inter prediction of one CU coded by the merge mode. Specifically, for each merge CU, one additional flag is signaled to indicate whether CIIP is enabled for the current CU. For the luma component, CIIP supports four frequently used intra modes including the planar mode, DC mode, horizontal mode, and vertical mode. For the chroma component, DM (i.e., the chroma reuses the same intra mode as the luma component) is always applied without additional signaling. Further, in the existing CIIP design, weighted averaging is applied and the inter prediction samples and intra prediction samples of one CIIP CU are combined. Specifically, equal weights (i.e., 0.5) are applied when the planar mode or DC mode is selected. Otherwise (i.e., either the horizontal mode or the vertical mode is applied), the current CU is first divided into four equally sized regions horizontally (in the case of the horizontal mode) or vertically (in the case of the vertical mode).

[0042] JPEG2025111753000006.jpg60161

[0043] Further, in the current VVC operation specification, the intra mode of one CIIP CU is determined via the most probable mode (MPM) mechanism based on the intra modes of its adjacent CIIP CUs. ​​​​​​​​​​​​​​​ It can be used as a predictor for predicting the intra mode. Specifically, for each CI Regarding the IP CU, when its adjacent blocks are also CIIP CUs, if so the intra modes of those adjacent blocks are first rounded to the nearest mode within the planar mode, DC mode, horizontal mode, and vertical mode, and then added to the MPM candidate list of the current CU . However, when constructing the MPM list of each intra CU, if one of its adjacent blocks is coded in the CIIP mode, it is regarded as unavailable . That is, the intra mode of one CIIP CU is not allowed to predict the intra mode of its adjacent intra CU. FIGS. 7A and 7B (described below) compare the MPM list generation processes of the intra CU and the CIIP CU.

[0044] JPEG2025111753000007.jpg72162

[0045] JPEG2025111753000008.jpg84163

[0046] JPEG2025111753000009.jpg137163

[0047] JPEG2025111753000010.jpg61163

[0048] Here, shift and o offset are equal to 15-BD and 1 << (14-BD) + 2 · (1 << 13 ), respectively, and are the right shift value and offset value applied to combine the L0 and L1 prediction signals of the dual prediction.

[0049] FIG. 6A shows a diagram of composite inter and intra prediction in the horizontal mode according to an example of the present disclosure. shown

[0050] FIG. 6B shows a diagram of vertical mode composite inter and intra prediction according to an example of the present disclosure. shown

[0051] FIG. 6C shows a diagram of plane mode and DC mode composite inter and intra prediction according to an example of the present disclosure. shown

[0052] FIG. 7A shows a flowchart of an intra CUS MPM candidate list generation process according to an example of the present disclosure. shown

[0053] FIG. 7B shows a flowchart of a CIIP CU MPM candidate list generation process according to an example of the present disclosure. shown

[0054] Improvements to CIIP CIIP can improve the efficiency of conventional motion compensation prediction, but its design can be further improved. Specifically, the following problems in the existing CIIP design in VVC are identified in the present disclosure. First, as described in the "Composite Inter and Intra Prediction" section, since CIIP combines samples of inter and intra prediction, each CIIP CU needs to generate a prediction signal using its reconstructed adjacent samples. This means that the decoding of one CIIP CU depends on the complete reconstruction of its adjacent blocks. Due to such interdependence, in an actual hardware implementation, CIIP needs to be executed at the reconstruction stage where the reconstructed adjacent samples become available for intra prediction. The decoding of CUs at the reconstruction stage must be performed sequentially (i.e., one by one).

[0055] ​​​​​​​Therefore, the calculation operations involved in the CIIP process (e.g., multiplication, addition, bit shifting) ) should not be too high to ensure sufficient throughput for real-time decoding. It is not possible to do so.

[0056] As mentioned in the "Bidirectional Optical Flow" section, BDOF is the forward and backward Two reference blocks from both backward time directions are used to generate one inter-coded This is enabled so that prediction quality improves when a given CU is predicted. As shown in the figure, in the current VVC, BDOF is also used for the inter-prediction sample in CIIP mode. Considering the additional complexity of BDOF, Such a design allows for hardware codec encoding / decoding when CIIP is enabled. Decoding throughput can be significantly reduced.

[0057] Second, in the current CIIP design, one CIIP CU is double-predicted. When referring to a merge candidate, generate motion compensated prediction signals for both lists L0 and L1. If one or more MVs are not integer precision, partial sampling is required. To interpolate the samples at the row positions, an additional interpolation process must be invoked. Such a process not only increases computational complexity but also requires more access from external memory. It also increases memory bandwidth when reference samples need to be accessed.

[0058] Then, as discussed in the "Combined Inter- and Intra-Prediction" section, the current CI In IP design, the intra mode of CIIP CU and the intra mode of intra CU are , they are treated differently when constructing the MPM lists of their adjacent blocks. Specifically, , when one current CU is coded in CIIP mode, its adjacent CIIP CUs are considered as intra, that is, the intra modes of the adjacent CIIP CUs can be added to the MPM candidate list. However, when the current CU is coded in intra mode , its adjacent CIIP CUs are considered as inter, that is, the intra modes of the adjacent CIIP CUs are excluded from the MPM candidate list. Such a non-unified design may not be optimal for the final version of the VVC standard.

[0059] Fig. 8 shows a diagram of the workflow of the existing CIIP design in VVC according to an example of the present disclosure.

[0060] Simplification of CIIP In the present disclosure, a method for simplifying the existing CIIP design is provided to facilitate hardware codec implementation. Generally, the main aspects of the technology proposed in the present disclosure are summarized as follows.

[0061] First, in order to improve the CIIP coding / decoding throughput, it is proposed to exclude BDOF from the generation of inter prediction samples in CIIP mode.

[0062] Next, in order to reduce the computational complexity and memory bandwidth consumption, when one CIIP CU is double predicted (i.e., has both L0 and L1 MVs), a method is proposed to convert the block from double prediction to single prediction in order to generate inter prediction samples. ​

[0063] Then, the two methods are proposed to harmonize the intra modes of the adjacent blocks in forming the MPM candidates for the intra CU and the CIIP.

[0064] CIIP without BDOF As pointed out in the "Problem Statement" section, BDOF is always enabled in order to generate an inter prediction sample for the CIIP mode when the current CU is double predicted. Due to the further complexity of BDOF, the existing CIIP design may significantly reduce the encoding / decoding throughput, especially when real-time decoding becomes difficult for the VVC decoder. On the other hand, for the CIIP CU, its final prediction sample is generated by averaging the inter prediction sample and the intra prediction sample. In other words, the improved prediction sample by BDOF is not directly used as the prediction signal for the CIIP CU. Therefore, compared with the conventional double predicted CU (where BDOF is directly applied to generate the prediction sample), the corresponding improvement obtained from BDOF is less efficient in the CIIP CU. Thus, based on the above circumstances, it is proposed to disable BDOF when generating the inter prediction sample for the CIIP mode. FIG. 9 (described below) shows the corresponding workflow of the proposed CIIP process after removing BDOF. Therefore, based on the above circumstances, it is proposed to disable BDOF when generating the inter prediction sample for the CIIP mode. FIG. 9 shows the corresponding workflow of the proposed CIIP process after removing BDOF. FIG. 9 shows a diagram illustrating the workflow of the proposed CIIP method by removing BDOF according to an example of the present disclosure.

[0065] FIG. 9 shows a diagram illustrating the workflow of the proposed CIIP method by removing BDOF according to an example of the present disclosure.

[0066] ​​​CIIP Based on Single Prediction As described above, the merge candidates referred to by one CIIP CU are double predicted Sometimes, both L0 and L1 prediction signals are generated to predict samples within the CU. Memory To reduce the memory bandwidth and interpolation complexity, in one embodiment of the present disclosure, (even if the current CU is double predicted), the inter prediction samples generated using single prediction are used to combine with the intra prediction samples in the CIIP mode. Specifically when the current CIIP CU is in single prediction, the inter prediction samples are directly combined with the intra prediction samples. Otherwise (i.e., when the current CU is double predicted ), the inter prediction samples used by CIIP are generated based on a single prediction from one prediction list (L0 or L1). Various methods can be applied to select the prediction list. In the first method, for any CIIP block predicted by two reference images , it is proposed to always select the first prediction (i.e., list L0).

[0067] In the second method, it is proposed to always select the second prediction (i.e., list L1) for any CIIP block predicted by two reference images . In the third method , one adaptive method is applied when the prediction list associated with one reference image with a small picture order count (POC) distance from the current image is selected. FIG. 10 (described below ) shows the workflow of single prediction-based CIIP that selects the prediction list based on the POC distance.

[0068] ​​​​Finally, in the last approach, the CIIP mode is proposed to be enabled only when the current CU is singly predicted. Further, to reduce the overhead, the signaling of the CIIP enable / disable flag depends on the prediction direction of the current CIIP CU. When the current CU is singly predicted, the CIIP flag is signaled in the bitstream to indicate whether CIIP is enabled or disabled. Otherwise (i.e., when the current CU is doubly predicted), the signaling of the CIIP flag is skipped and it is always assumed to be false, i.e., CIIP is always disabled. Figure 10 shows a diagram illustrating the workflow of a single-prediction-based CIIP that selects a prediction list based on the POC distance according to an example of the present disclosure. Harmonization of the intra mode of the intra CU and CIIP for MPM candidate list construction

[0069] As described above, the current CIIP design is not unified with respect to how to form the MPM candidate lists of their adjacent blocks using the intra modes of the intra CU and the CIIP CU. Specifically, in both the intra modes of the intra CU and the CIIP CU, the intra mode of an adjacent block coded in the CIIP mode can be predicted. However, only in the intra mode of the intra CU can the intra mode of the intra CU be predicted. To achieve another unified design, two methods are proposed in this section to harmonize the usage of the intra modes of the intra CU and the CIIP for MPM list construction.

[0070]

[0071] ​​​​​​​​​In the first method, the CIIP mode is treated as an inter-mode for MPM list configuration. It has been proposed to do so. Specifically, when generating the MPM list of either one CIIP CU or one intra-CU, if the adjacent block is coded in the CIIP mode, the intra-mode of the adjacent block is marked as unavailable. In such a method, the intra-mode of the CIIP block cannot be used to configure the MPM list. Conversely, in the second method, it has been proposed to treat the CIIP mode as an intra-mode for MPM list configuration. Specifically, in this method, in the intra-mode of the CIIP CU, the intra-modes of both the adjacent CIIP block and the intra-block can be predicted. FIGS. 11A and 11B (described below) show the MPM candidate list generation process when the above two methods are applied. In the intra-mode of the CIIP CU, the intra-modes of both the adjacent CIIP block and the intra-block can be predicted. FIGS. 11A and 11B (described below) show the MPM candidate list generation process when the above two methods are applied. show the MPM candidate list generation process when the above two methods are applied.

[0072] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and implementation of the disclosure presented herein. This application is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles thereof and including departures from the present disclosure within the scope of knowledge or customary practice in the art. The true scope and spirit of the disclosure are indicated by the following patent claims, and the specification and examples are to be regarded only as illustrative examples. claims, and the specification and examples are to be regarded only as illustrative examples. claims, and the specification and examples are to be regarded only as illustrative examples. intended.

[0073] It should be understood that the present disclosure is not limited to the specific examples described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is intended to be limited only by the appended patent claims. intended. It is.

[0074] Figure 11A shows a flowchart of a method for enabling a CIIP block for MPM candidate list generation according to an example of the present disclosure. indicating.

[0075] Figure 11B shows a flowchart of a method for disabling a CIIP block for MPM candidate list generation according to an example of the present disclosure. indicating.

[0076] Figure 12 shows a computing environment 12 10 coupled to a user interface 1260. The computing environment 1210 can be part of a data processing server. The computing environment 1210 includes a processor 1220, a memory 1240, and an I / O interface 1250.

[0077] The processor 1220 generally controls the overall operation of the computing environment 1210, such as operations related to display, data acquisition, data communication, and image processing. The processor 1220 can include one or more processors for executing instructions to perform all or some of the steps of the above methods. Further, the processor 1220 can include one or more circuits that facilitate the interaction between the processor 1 220 and other components. The processor can be a central processing unit (CPU), a microprocessor, a single-chip machine , a GPU, etc. The memory 1240 is configured to store various types of data to support the operation of the computing environment 1210. Examples of such data are computing

[0078] types of data to support the operation of the computing environment 1210. Examples of such data are computing types of data to support the operation of the computing environment 1210. Examples of such data are computing Instructions, videos, etc. for use in any application or method operating in the wing environment 1210 Includes data, image data, etc. The memory 1240 is of any type of volatile or non-volatile Memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM) , erasable programmable read-only memory (EPROM), programmable read Only memory (PROM), read-only memory (ROM), magnetic memory, flash memory , can be realized using magnetic disks or optical disks.

[0079] The I / O interface 1250 provides an interface between the processor 1220 and peripheral interface modules such as keyboards, click hoes , buttons, etc. Buttons include, but are not limited to, home buttons, scan start buttons, and scan stop buttons. The I / O interface 1250 can be coupled to an encoder And a decoder.

[0080] In one embodiment, a non-transitory computer-readable storage medium including a plurality of programs such as those included in the memory 1240, executable by the processor 1220 within the computing environment 1210, is also provided for performing the above-described method. For example, the non-transitory Computer-readable storage medium can be ROM, RAM, CD-ROM, magnetic tape, floppy Disk, optical data storage device, etc.

[0081]

[0081] The non-transitory computer-readable storage medium has one or more processors It stores a plurality of programs for execution by a computing device therein , when the plurality of programs are executed by one or more processors, the computing ting device executes a method for predicting the above-described operation.

[0082] In one embodiment, the computing environment 1210 is configured to execute the above-described method using one or more application specific integrated circuits (ASICs), digital signal processors (DS Ps), digital signal processing devices (DSPDs), programmable logic devices (PL Ds), field programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors ors, or other electronic components.

Claims

1.

2. Determining whether the BDOF operation is not applied further includes determining that the BDOF operation is not applied on the condition that the CIIP is applied to generate the final prediction of the current prediction block. The method according to claim 1. Determining that the BDOF operation is applied further includes determining that the BDOF operation is applied when the CIIP is not applied to generate the final prediction of the current prediction block. The method according to claim 1. The dual prediction of the current block is calculated based on averaging the first prediction L0 and the second prediction L1. The method according to claim 2.

3.

4.

5.

6.

7.

8.

9.

10.

11. Obtaining a reference image in a reference image list associated with the current prediction block, Generating an inter prediction based on a first motion vector from the current image to a first reference image, Obtaining an intra prediction mode associated with the current prediction block, Generating an intra prediction of the current prediction block based on the intra prediction, Generating a final prediction of the current prediction block by averaging the inter prediction and the intra prediction, 、 Determining whether the current prediction block is treated as either an inter mode or an intra mode for the most probable mode (MPM)-based intra mode prediction, A video coding method for compromising. When the current prediction block is predicted from one reference image in the reference image list L0, the reference image list is L0. The method according to claim 6. When the current prediction block is predicted from one reference image in the reference image list L1, the reference image list is L1. The method according to claim 6. When the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, the reference image list is L0. The method according to claim 6. When the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, the reference image list is L1. The method according to claim 6.

12.

13.

14.

15.

16.

17.

18.

19.

20.

21.

22.

23.

24.

25.

26. The reference image list is such that when the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, it is associated with one reference image having a smaller picture order count (POC) distance to the current image. The method according to claim 6. When the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, it is associated with one reference image having a smaller picture order count (POC) distance to the current image. The method according to claim 6. When the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, it is associated with one reference image having a smaller picture order count (POC) distance to the current image. The method according to claim 6. The method according to claim 6, wherein the reference image list is associated with one reference image having a smaller picture order count (POC) distance to the current image when the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1. Claim 12 The current prediction block is treated as an inter mode, and the intra prediction mode of the current prediction block is not used for MPM-based intra mode prediction. The method according to claim 6. The current prediction block is treated as an intra mode, and the intra prediction mode of the current prediction block is not used for MPM-based intra mode prediction. The method according to claim 6. The method according to claim 6, wherein the current prediction block is treated as an intra mode, and the intra prediction mode of the current prediction block is not used for MPM-based intra mode prediction. Claim 13 The current prediction block is treated as an intra mode, and the intra prediction mode of the current prediction block is used for MPM-based intra mode prediction. The method according to claim 6. The current prediction block is treated as an intra mode, and the intra prediction mode of the current prediction block is used for MPM-based intra mode prediction. The method according to claim 6. The method according to claim 6, wherein the current prediction block is treated as an intra mode, and the intra prediction mode of the current prediction block is used for MPM-based intra mode prediction. Claim 14 Claim 15 Determining whether the BDOF operation is not applied further includes determining that the BDOF operation is not applied under the condition that composite inter and intra prediction (CIIP) is applied to generate the final prediction of the current prediction block. The non-transitory computer-readable storage medium according to claim 14. Determining whether the BDOF operation is not applied further includes determining that the BDOF operation is not applied under the condition that composite inter and intra prediction (CIIP) is applied to generate the final prediction of the current prediction block. The non-transitory computer-readable storage medium according to claim 14. Determining whether the BDOF operation is not applied further includes determining that the BDOF operation is not applied under the condition that composite inter and intra prediction (CIIP) is applied to generate the final prediction of the current prediction block. The non-transitory computer-readable storage medium according to claim 14. The non-transitory computer-readable storage medium according to claim 14, wherein determining whether the BDOF operation is not applied further includes determining that the BDOF operation is not applied under the condition that composite inter and intra prediction (CIIP) is applied to generate the final prediction of the current prediction block. Claim 16 Determining that the BDOF operation is applied further includes determining that the BDOF operation is applied when CIIP is not applied to generate the final prediction of the current prediction block. The non-transitory computer-readable storage medium according to claim 14. Determining that the BDOF operation is applied further includes determining that the BDOF operation is applied when CIIP is not applied to generate the final prediction of the current prediction block. The non-transitory computer-readable storage medium according to claim 14. Determining that the BDOF operation is applied further includes determining that the BDOF operation is applied when CIIP is not applied to generate the final prediction of the current prediction block. The non-transitory computer-readable storage medium according to claim 14. Claim 17 Calculating the double prediction of the current prediction block is calculated based on the first prediction L0 and the second prediction L1. The non-transitory computer-readable storage medium according to claim 15. Calculating the double prediction of the current prediction block is calculated based on the first prediction L0 and the second prediction L1. The non-transitory computer-readable storage medium according to claim 15. The non-transitory computer-readable storage medium according to claim 15, wherein calculating the double prediction of the current prediction block is calculated based on the first prediction L0 and the second prediction L1. Claim 18 Claim 19 A non-transitory computer-readable storage medium storing a plurality of programs executed by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, obtain reference images in a reference image list associated with a current prediction block, generate an inter prediction based on a first motion vector from a current image to a first reference image, generate an inter prediction based on a first motion vector from a current image to a first reference image, obtain an intra prediction mode associated with the current prediction block, generate an intra prediction of the current prediction block based on the intra prediction, generate an intra prediction of the current prediction block based on the intra prediction, and By averaging the inter prediction and the intra prediction, a final prediction of the current prediction block is generated; identifying whether the current prediction block is treated as either an inter mode or an intra mode with respect to a most probable mode (MPM)-based intra mode prediction; and causing the computing device to perform operations including the above, a non-transitory computer readable storage medium. **Claim 20** The non-transitory computer readable storage medium according to claim 19, wherein when the current prediction block is predicted from one reference image in the reference image list L0, the reference image list is L0. **Claim 21** The non-transitory computer readable storage medium according to claim 19, wherein when the current prediction block is predicted from one reference image in the reference image list L1, the reference image list is L1. **Claim 22** The non-transitory computer readable storage medium according to claim 19, wherein when the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, the reference image list is L0. **Claim 23** The non-transitory computer readable storage medium according to claim 19, wherein when the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1, the reference image list is L1. **Claim 24** The non-transitory computer readable storage medium according to claim 19, wherein the reference image list is associated with one reference image having a smaller picture order count (POC) distance to the current image when the current prediction block is predicted from one first reference image in the reference image list L0 and one second reference image in the reference image list L1. **Claim 25** **Claim 26** The non-transitory computer readable storage medium according to claim 19, wherein the current prediction block is treated as an inter mode, and the intra prediction mode of the current prediction block is not used for MPM-based intra mode prediction. **Claim 27** The non-transitory computer readable storage medium according to claim 19, wherein the current prediction block is treated as an intra mode, and the intra prediction mode of the current prediction block is used for MPM-based intra mode prediction.

Citation Information

Patent Citations

  • Method for processing image based on joint inter-intra prediction mode and apparatus therefor

    US20180249156A1

  • Image decoding device, image decoding method, and image encoding device

    WO2013047811A1