Constrained and adjusted application of combined inter and intra prediction modes
By constraining and adjusting the application of CIIP mode, the problems of computational complexity and throughput in existing technologies are solved, achieving more efficient video encoding and decoding and improving the performance of hardware codecs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2020-03-12
- Publication Date
- 2026-05-12
AI Technical Summary
In existing video coding and decoding technologies, the combined inter-frame and intra-frame prediction mode (CIIP) has problems in terms of computational complexity and encoding/decoding throughput, especially in hardware implementation. Furthermore, the inconsistency in the processing of intra-frame and inter-frame modes affects coding and decoding efficiency.
By constraining and adjusting the application of CIIP mode, inter-frame prediction techniques such as BDOF and DMVR can be disabled or selectively bypassed. A unified criterion is adopted to form an MPM candidate list, reducing computational operations and memory accesses, and improving hardware encoding and decoding efficiency.
It reduces the computational complexity and memory bandwidth requirements of hardware codecs, improves encoding/decoding throughput, and achieves more efficient video encoding and decoding performance.
Smart Images

Figure CN119653082B_ABST
Abstract
Description
[0001] This application is a divisional application of application number "202080020524.0", filed on "March 12, 2020", entitled "Constrained and Modulated Application of Combined Inter-Frame and Intra-Frame Prediction Modes".
[0002] Cross-reference to related applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 817,503, filed March 12, 2019. The entire disclosure of the above application is incorporated herein by reference. Technical Field
[0004] This disclosure generally relates to video encoding / decoding and compression. More specifically, this disclosure relates to systems and methods for performing video encoding / decoding using constraints and adjustments for applications combining inter-frame and intra-frame prediction (CIIP) modes. Background Technology
[0005] This section provides background information relevant to this disclosure. The information contained in this section should not necessarily be construed as prior art.
[0006] Any of a variety of video codec technologies can be used to compress video data. Video codecs can be performed according to one or more video codec standards. Some illustrative video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM) codec, High Efficiency Video Codec (H.265 / HEVC), High-Advanced Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) codec.
[0007] Video encoding and decoding typically employ prediction methods that utilize the inherent redundancy in video images or sequences (such as inter-frame prediction, intra-frame prediction, etc.). One goal of video encoding and decoding technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing video quality degradation.
[0008] Prediction methods used in video coding and decoding typically involve performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data, and are often associated with block-based video coding and decoding.
[0009] Public content
[0010] This section provides a general overview of this disclosure and is not a full disclosure of its entire scope or all its features.
[0011] According to a first aspect of this disclosure, a video encoding / decoding method is executed at a computing device having one or more processors and a memory storing multiple programs to be executed by the one or more processors. The method includes: dividing each image in a video stream into multiple blocks or CUs. The method also includes: when a CU is bidirectionally predicted, during the application of a CIIP mode to the CU, operations that bypass one or more inter-frame prediction processes while generating inter-frame prediction samples. The one or more bypassed inter-frame prediction processes include decoder-side motion vector refinement (DMVR) and bidirectional optical flow (BDOF).
[0012] According to a second aspect of this disclosure, a video encoding / decoding method is executed at a computing device having one or more processors and a memory storing multiple programs to be executed by the one or more processors. The method includes: dividing each image in a video stream into multiple blocks or CUs. The method further includes: identifying CUs as candidate CUs for application in a CIIP mode. The method further includes: determining whether the candidate CUs identified as candidate CUs for application in a CIIP mode are predicted bidirectionally or unidirectionally. The method further includes: constraining the application of the CIIP mode to the CUs based on the determination.
[0013] According to a third aspect of this disclosure, a video encoding / decoding method is executed at a computing device having one or more processors and a memory storing multiple programs to be executed by the one or more processors. The method includes: dividing each frame in a video stream into multiple blocks or coding units (CUs). The method further includes: deriving an MPM candidate list for each CU. The method further includes: determining whether each of the neighboring CUs of the CU (“current CU”) is a block that is respectively encoded / decoded by CIIP. The method further includes: for each neighboring CU that is encoded / decoded by CIIP, in forming the MPM candidate list for the current CU using the intra-frame modes of the neighboring CUs, employing a unified criterion independent of determining whether the current CU is intra-frame encoded or CIIP encoded.
[0014] According to a fourth aspect of this application, a computing device includes one or more processors, a memory, and a plurality of programs stored in the memory. When executed by the one or more processors, these programs cause the computing device to perform the operations described above.
[0015] According to a fifth aspect of this application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors. When executed by the one or more processors, these programs cause the computing device to perform the operations described above. Attached Figure Description
[0016] In the following description, a collection of illustrative, non-limiting embodiments of the present disclosure will be described in conjunction with the accompanying drawings. Those skilled in the art can implement variations in structure, method, or function based on the examples presented herein, and such variations are all included within the scope of this disclosure. Where there is no conflict, the teachings of the different embodiments may, but are not required, be combined with each other.
[0017] Figure 1 It is an illustrative block diagram of a block-based hybrid video encoder that can be used in conjunction with many video codec standards.
[0018] Figure 2 It is a block diagram illustrating a video decoder that can be used in conjunction with many video codec standards.
[0019] Figures 3A-3E An exemplary splitting type according to an exemplary embodiment is shown, namely, a quadrilateral split ( Figure 3A ), horizontal binary segmentation ( Figure 3B Vertical binary segmentation ( Figure 3C ), horizontal ternary segmentation ( Figure 3D ) and vertical ternary segmentation ( Figure 3E ).
[0020] Figures 4A-4C This is a diagram illustrating the combination of inter-frame prediction and intra-frame prediction in CIIP mode.
[0021] Figures 5A-5B These are a pair of flowcharts illustrating the process of generating the MPM candidate list in the current VVC.
[0022] Figure 6 This is a flowchart illustrating the workflow of CIIP design using bidirectional optical flow (BDOF) in current VVC.
[0023] Figure 7 This is a flowchart illustrating an illustrative workflow for selectively bypassing DMVR and BDOF operations when calculating bidirectional predictions for the current prediction block.
[0024] Figure 8 This is a flowchart illustrating the illustrative workflow of the CIIP design presented in this disclosure.
[0025] Figure 9 This is a flowchart illustrating the illustrative workflow of the second CIIP design presented in this disclosure.
[0026] Figures 10A-10B This is a pair of flowcharts illustrating the two proposed methods for processing CIIP codec blocks during the generation of the MPM candidate list in this disclosure. Detailed Implementation
[0027] The terminology used in this disclosure is intended to illustrate specific examples and not to limit the disclosure. The singular forms “a,” “an,” and “the,” as used in this disclosure and the appended claims, also refer to the plural forms unless otherwise explicitly implied in the context. It should be understood that the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed items.
[0028] It should be understood that while the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish one type of information from another. For example, without departing from the scope of this disclosure, first information may be referred to as second information; and similarly, second information may be referred to as first information. As used herein, the term "if" may be understood to mean "when," "once," or "in response to," depending on the context.
[0029] Throughout this specification, references to "one embodiment," "an embodiment," "another embodiment," etc., in singular or plural form refer to one or more specific features, structures, or characteristics described in connection with an embodiment that are included in at least one embodiment of this disclosure. Therefore, phrases such as "in one embodiment," "in another embodiment," etc., appearing throughout this specification in singular or plural form, do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable manner.
[0030] Conceptually, many video codec standards are similar, including those mentioned earlier in the background section. For example, almost all video codec standards use block-based processing and share similar video codec diagrams for video compression.
[0031] In block-based video coding and decoding, the input video signal is processed block by block. For each block (also called a codec unit (CU)), spatial prediction and / or temporal prediction can be performed.
[0032] Spatial prediction (also known as "intra-frame prediction") uses pixels from samples (called reference samples) of neighboring blocks that have already been encoded and decoded within the same video frame / strip to predict the current block. Spatial prediction reduces the spatial redundancy inherent in the video signal.
[0033] Timing prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded and decoded video frames to predict the current block. Timing prediction reduces the inherent temporal redundancy in the video signal. The timing prediction signal for a given CU is typically transmitted via one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its timing reference. Additionally, when multiple reference frames are supported, a reference frame index is sent to identify which reference frame in the reference frame library the timing prediction signal originates from.
[0034] Following spatial and / or temporal prediction, the mode decision block in the encoder selects the optimal prediction mode, for example, using a rate-distortion optimization method. The prediction block is then subtracted from the current block, and the prediction residuals are decorrelated using transform and quantization. The quantized residual coefficients are inversely quantized and inversely transformed to form the reconstruction residuals, which are then added back to the prediction block to form the reconstructed signal for that block.
[0035] Following spatial and / or temporal prediction, and before the reconstructed CU is placed into the reference image library and used for encoding and decoding future video blocks, further loop filtering, such as deblocking filters, sample adaptive offset (SAO), and adaptive loop filters (ALF), can be applied to the reconstructed CU. To form the output video bitstream, the encoding / decoding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy encoding / decoding unit for further compression and packing to form the bitstream.
[0036] During the decoding process, the video bitstream is first entropy-decoded in the entropy decoding unit. The encoding / decoding mode and prediction information are sent to the spatial prediction unit (in intra-frame encoding / decoding) or the temporal prediction unit (in inter-frame encoding / decoding) to form prediction blocks. Residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. The prediction blocks and residual blocks are then added together. The reconstructed blocks may undergo further loop filtering before being stored in the reference picture library. The reconstructed video from the reference picture library is then sent to drive the display device and is used to predict future video blocks.
[0037] In many hybrid video codec schemes, each block can employ either inter-frame or intra-frame prediction methods, rather than both. However, the residual signals generated by inter-frame and intra-frame prediction blocks can exhibit very different characteristics. Therefore, if the two prediction methods can be combined efficiently, a more accurate prediction can be expected to reduce the energy of the prediction residuals, thereby improving codec efficiency. Furthermore, in some video content, the motion of moving objects can be complex. For example, there may be regions containing both old content (e.g., objects included in previously codec images) and newly revealed content (e.g., objects not included in previously codec images). In such cases, neither inter-frame nor intra-frame prediction can provide an accurate prediction for the current block.
[0038] To further improve prediction efficiency, the current version of VVC introduces a Combined Inter-Frame and Intra-Frame Prediction (CIIP) mode, which combines intra-frame and inter-frame predictions from a single CU encoded and decoded using a merged mode. The application of CIIP mode generally improves encoding and decoding efficiency.
[0039] However, the application of CIIP mode involves more operations than those involved in inter-frame or intra-frame modes, which often increases computational complexity and reduces encoding / decoding throughput.
[0040] Furthermore, in the context of constructing the most likely mode (MPM) candidate list for adjacent blocks of intra-mode blocks and CIIP mode blocks, the current version of VVC is not entirely consistent in its treatment of intra-mode blocks and CIIP mode blocks.
[0041] Figure 1 An illustrative block diagram of a block-based hybrid video encoder 100 is shown, which can be used in conjunction with many video codec standards. In encoder 100, video frames are divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method. In inter-frame prediction, one or more prediction values are formed based on pixels from previously reconstructed frames through motion estimation and motion compensation. In intra-frame prediction, prediction values are formed based on pixels reconstructed in the current frame. Through mode decision, the best prediction value can be selected to predict the current block.
[0042] The prediction residual, representing the difference between the current video block and its predicted value, is sent to the transform circuit 102. The transform coefficients are then sent from the transform circuit 102 to the quantization circuit 104 for entropy reduction. The quantization coefficients are then fed to the entropy encoding / decoding circuit 106 to generate a compressed video bitstream. Figure 1As shown, prediction-related information 110 from the inter-frame prediction circuit and / or intra-frame prediction circuit 112, such as video block segmentation information, motion vectors, reference picture indexes, and intra-frame prediction modes, is also fed through the entropy encoding / decoding circuit 106 and stored in the compressed video bitstream 114.
[0043] In encoder 100, circuitry associated with the decoder is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed via inverse quantization circuit 116 and inverse transform circuit 118. This reconstructed prediction residual is combined with block prediction values 120 to generate unfiltered reconstructed pixels for the current video block.
[0044] To improve encoding / decoding efficiency and visual quality, loop filters are commonly used. For example, deblocking filters are available in current versions of AVC, HEVC, and VVC. In HEVC, an additional loop filter called SAO (Sample Adaptive Offset) is defined to further improve encoding / decoding efficiency. In the current version of the VVC standard, another loop filter called ALF (Adaptive Loop Filter) is under active research and is very likely to be included in the final standard.
[0045] These loop filter operations are optional. Performing these operations helps improve encoding / decoding efficiency and visual quality. They can also be turned off by the encoder 100 to save computational complexity.
[0046] It should be noted that intra-frame prediction is typically based on unfiltered reconstructed pixels, while inter-frame prediction is based on filtered reconstructed pixels if these filter options are enabled by encoder 100.
[0047] Figure 2 This is a block diagram illustrating a video decoder 200 that can be used in conjunction with many video codec standards. The decoder 200 is similar to a device residing in... Figure 1 The reconstruction-related part in encoder 100. In decoder 200 ( Figure 2In this process, the incoming video bitstream 201 is first decoded via entropy decoding 202 to derive quantization coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed via inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residual. The block prediction mechanism implemented in the intra / inter-frame mode selector 208 is configured to perform intra-frame prediction 210 or motion compensation 212 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by adding the reconstructed prediction residual from inverse transform 206 to the prediction output generated by the block prediction mechanism using adder 214. With the loop filter enabled, filtering is performed on these reconstructed pixels to derive the final reconstructed video. The reconstructed video from the reference image library is then sent to drive the display device and used to predict future video blocks.
[0048] In video codec standards such as HEVC, blocks can be segmented based on quadtrees. Newer video codec standards (such as the current VVC) employ more segmentation methods and can split a codec tree unit (CTU) into multiple CUs to accommodate different local characteristics based on quadtrees, binary trees, or ternary trees. In the current VVC, most codec modes do not separate CUs, prediction units (PUs), and transform units (TUs), and each CU always serves as the basic unit for both prediction and transform without further segmentation. However, in some specific codec modes (such as intra-frame sub-segmentation codecs), each CU can still contain multiple TUs. In multi-type tree structures, a CTU is first segmented by a quadtree structure. Then, each quadtree leaf node can be further segmented using binary and ternary tree structures.
[0049] Figures 3A-3E Five example splitting types are shown, namely quadruple splitting ( Figure 3A ), horizontal binary segmentation ( Figure 3B Vertical binary segmentation ( Figure 3C ), horizontal ternary segmentation ( Figure 3D ) and vertical ternary segmentation ( Figure 3E In video codec standards such as HEVC and the current VVC, intra-frame prediction or inter-frame prediction can be performed for each CU.
[0050] The current VVC also introduces a combined intra-frame and inter-frame prediction (CIIP) coding and decoding mode, in which the intra-frame and inter-frame predictions of a CU encoded and decoded by the merge mode are combined.
[0051] In CIIP mode, for each merged CU, an additional flag is signaled to indicate whether CIIP is enabled for the current CU. For the luma component of the current CU, CIIP supports four frequently used intra-frame modes: planar mode, DC mode, horizontal mode, and vertical mode. For the chroma component of the current CU, DM mode is always applied without additional signaling; DM mode means reusing the same intra-frame mode for the luma component.
[0052] Furthermore, in CIIP mode, for each merged CU, a weighted average is applied to combine the inter-frame prediction samples and intra-frame prediction samples of the current CU. Specifically, when choosing planar mode or DC mode, equal weights (i.e., 0.5) are applied. Otherwise (i.e., applying horizontal mode or vertical mode), the current CU is first horizontally split (for horizontal mode) or vertically split (for vertical mode) into four regions of equal size. This is represented as (w_intra i ,w_inter i Four weight sets will be applied to combine inter-frame prediction samples and intra-frame prediction samples from different regions, where i = 0 and i = 3 represent the nearest and farthest regions from the neighboring samples used for intra-frame prediction reconstruction. In the current CIIP design, the predefined set of values for the weight sets are: (w_intra0, w_inter0) = (0.75, 0.25), (w_intra1, w_inter1) = (0.625, 0.375), (w_intra2, w_inter2) = (0.375, 0.625), and (w_intra3, w_inter3) = (0.25, 0.75).
[0053] Figures 4A-4C This is a diagram illustrating the combination of inter-frame prediction and intra-frame prediction in CIIP mode. Figure 4A The diagram illustrates the operation of combining weights for CU 401 in horizontal mode. The current CU is horizontally divided into four equal-sized regions 402, 403, 404, and 405, represented by four different grayscale shades. For each region, a weight set is applied to combine inter-frame prediction samples and intra-frame prediction samples. For example, for the leftmost region 402, represented by the darkest grayscale shade, the weight set (w_intra0, w_inter0) = (0.75, 0.25) is applied, meaning that a CIIP prediction is obtained as the sum of (i) 0.75 or 3 / 4 times the intra-frame prediction and (ii) 0.25 or 1 / 4 times the inter-frame prediction. This is also illustrated by formula 406, which is attached to the same region 402 with an arrow. To illustrate. The weight sets used for the other three regions, 403, 404, and 405, are similarly shown. Figure 4B In the middle, with Figure 4AA similar approach was used to illustrate the operation of combined weights for the CU 407 in vertical mode. Figure 4C The section describes the operation of combining weights for CU 408 in either planar or DC mode, where only one set of weights—a set of equal weights—is applied to the entire CU 408. This is also illustrated by formula 409, which is attached to the entire CU 408 with an arrow. To illustrate.
[0054] The current VVC specification also stipulates that the intra-frame mode of a CIIP CU is used as the prediction value to predict the intra-frame mode of its neighboring CIIP CUs through the Most Probable Mode (MPM) mechanism. However, it does not stipulate that the same intra-frame mode of the same CIIP CU is used as the prediction value to predict the intra-frame mode of its neighboring intra-encoded CUs through the MPM mechanism.
[0055] Specifically, for each CIIP CU, if some of its neighboring blocks are also CIIP CUs, the intra-modes of these neighbors are first rounded to the closest among planar, DC, horizontal, and vertical modes, and then added to the current CU's MPM candidate list. This allows an intra-mode of a CIIP CU to predict the intra-modes of its neighboring CIIP CUs. However, for each intra-CU, if some of its neighboring blocks are encoded or decoded using a CIIP mode, these neighboring blocks are considered unavailable; that is, an intra-mode of a CIIP CU is not allowed to predict the intra-modes of its neighboring intra-CUs.
[0056] Figures 5A-5B These are a pair of flowcharts illustrating the process of generating the MPM candidate list in the current VVC. Figure 5A The diagram illustrates the generation process where the current CU is an intra-block. In this case, a determination is made as to whether the adjacent CU is an intra-block (501), and the intra-block mode of the adjacent CU is added to the MPM candidate list if and only if it is an intra-block (502). Figure 5B The diagram illustrates the generation process when the current CU is an intra-block. In this case, a determination is made as to whether the adjacent CU is an intra-block (503), and when it is an intra-block, the intra-block mode of the adjacent CU is rounded (504) and then added to the MPM candidate list (505). On the other hand, when it is not an intra-block, a subsequent determination is made as to whether the adjacent CU is a CIIP block (506), and when it is a CIIP block, the intra-block mode of the adjacent CU is added to the MPM candidate list (507).
[0057] Figures 5A-5BThe illustration shows the difference between the use of CIIP blocks and intra-frame blocks in the current VVC MPF candidate list generation process. Specifically, the method of using adjacent CIIP blocks in the generation process of the MPM candidate list depends on whether the current block is an intra-frame block or a CIIP block, and the two cases are different.
[0058] Newer video codec standards (such as the current VVC) have introduced new inter-mode codec tools, and two examples of these new inter-mode codec tools are: Bidirectional Optical Flow (BDOF) and Decoder-Side Motion Vector Refinement (DMVR).
[0059] Conventional bidirectional prediction in video encoding and decoding is a simple combination of two temporal prediction blocks obtained from a reconstructed reference image. However, due to the limitations of block-based motion compensation, there may be observable residual minute motion between samples of the two prediction blocks, thus reducing the efficiency of motion-compensated prediction. To address this issue, BDOF is applied in current VVC to reduce the impact of such motion on each sample within a block.
[0060] BDOF is a per-sample motion refinement performed on top of block-based motion-compensated predictions when using bidirectional prediction. The motion refinement for each 4×4 sub-block is calculated by minimizing the difference between the predicted samples from reference image list 0 (L0) and reference image list 1 (L1) after applying BDOF within a 6×6 window around that sub-block. Based on this derived motion refinement, the final bidirectional prediction samples for the CU are calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on an optical flow model.
[0061] DMVR is a bidirectional prediction technique for merging blocks, featuring two initial motion vectors (MVs) signaled by a signal. These MVs can be further refined using bilateral matching prediction. Bilateral matching derives the motion information of the current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference images. The cost function used in the matching process is the sum of absolute differences (SAD) of row subsamples. After the matching process, the refined MVs are used for motion compensation in the prediction phase, boundary strength calculation in the deblocking filter, temporal motion vector prediction for subsequent images, and cross-CTU spatial motion vector prediction for subsequent CUs. Under the assumption of continuous motion trajectories, the motion vectors MV0 and MV1 pointing to the two reference blocks should be proportional to the temporal distance (TD0 and TD1) between the current image and the two reference images. As a special case, when the current image is temporally located between the two reference images and the temporal distance from the current image to the two reference images is the same, bilateral matching becomes a mirror-based bidirectional MV.
[0062] In the current VVC, both BDOF and DMVR can be used in conjunction with CIIP mode.
[0063] Figure 6 This is a flowchart illustrating the illustrative workflow of a CIIP design using bidirectional optical flow (BDOF) in current VVC. In this workflow, L0 motion compensation and L1 motion compensation (601 and 602) are processed by BDOF (603), and the output from BDOF is then processed together with intra-frame prediction (604) to form a weighted average (605) in the CIIP mode.
[0064] The current VVC also offers similar applications that combine DMVR and CIIP modes.
[0065] In current VVC, the CIIP mode can improve the efficiency of prediction after motion compensation. However, this disclosure has identified three problems in the current CIIP design included in current VVC.
[0066] First, because CIIP combines samples from inter-frame prediction and intra-frame prediction, each CIIP CU needs to use its reconstructed neighboring samples to generate the prediction signal. This means that decoding a CIIP CU depends on the complete reconstruction of its neighboring blocks. Due to this interdependence, for practical hardware implementations, CIIP needs to be performed during the reconstruction phase, where adjacent reconstructed samples become available for intra-frame prediction. Because decoding of CUs during the reconstruction phase must be performed sequentially (i.e., one by one), the number of computational operations (e.g., multiplication, addition, and shifting) involved in the CIIP process cannot be too high to guarantee sufficient real-time decoding throughput. Furthermore, current CIIP designs in VVC involve new inter-modal codec tools (such as BDOF and DMVR) to generate inter-frame prediction samples for CIIP modes. Given the additional complexity introduced by these new inter-modal codec tools, this design can significantly reduce the encoding / decoding throughput of the hardware codec when CIIP is enabled.
[0067] Second, in the current CIIP design within VVC, when a CIIP CU references a bidirectionally predicted merge candidate, both motion-compensated prediction signals in lists L0 and L1 need to be generated. When one or more MVs are not of integer precision, an additional interpolation process must be invoked to interpolate the samples at fractional sample locations. This process not only increases computational complexity but also increases memory bandwidth due to the need to access more reference samples from external memory, and therefore can potentially severely degrade the encoding / decoding throughput of the hardware codec when CIIP is enabled.
[0068] Third, in the current CIIP design of VVC, the intra-frame modes of CIIP CUs and the intra-frame modes of intra-frame CUs are treated differently when constructing their neighboring block MPM lists. Specifically, when a current CU is encoded / decoded in CIIP mode, its neighboring CIIP CUs are considered intra-frame, meaning their intra-frame modes can be added to the MPM candidate list. However, when a current CU is encoded / decoded in intra-frame mode, its neighboring CIIP CUs are considered inter-frame, meaning their intra-frame modes are excluded from the MPM candidate list. This non-uniform design may not be optimal for the final version of the VVC standard.
[0069] This disclosure proposes a constrained and regulated application of the CIIP model to address the three issues mentioned above.
[0070] According to this disclosure, after identifying CUs as candidates for applying the CIIP pattern, a determination will be made as to whether the CUs identified as candidates for applying the CIIP pattern are bidirectional or unidirectional predictions, and then based on this determination, the application of the CIIP pattern to the CUs will be constrained.
[0071] According to one embodiment of this disclosure, constraining the application of CIIP mode to a CU based on determination includes: when the CU is bidirectionally predicted, disabling the operation of one or more predefined inter-frame prediction techniques during the generation of inter-frame prediction samples while applying CIIP mode.
[0072] In one example, one or more predefined inter-frame prediction techniques include BDOF and DMVR.
[0073] Figure 7This is a flowchart illustrating an explanatory workflow for selectively bypassing DMVR and BDOF operations when calculating bidirectional predictions for the current prediction block. When processing the current prediction block, a first reference image and a second reference image associated with the current prediction block are obtained (702), wherein the first reference image precedes the current image and the second reference image follows the current image in display order. Subsequently, a first prediction L0 is obtained based on a first motion vector MV0 from the current prediction block to the reference block in the first reference image (703), and a second prediction L1 is obtained based on a second motion vector MV1 from the current prediction block to the reference block in the second reference image (704). Then, a determination is made regarding whether to apply the DMVR operation (705). If applied, the DMVR operation adjusts the first motion vector MV0 and the second motion vector MV1 based on the first prediction L0 and the second prediction L1, generating an updated first prediction L0' and an updated second prediction L1' (706). Then, a second determination is made regarding whether to apply the BDOF operation (707). If applied, the BDOF operation will compute first horizontal and vertical gradient values for the predicted sample associated with the updated first prediction L0', and second horizontal and vertical gradient values associated with the updated second prediction L1' (708). Finally, based on the first prediction L0, the second prediction L1, the optional updated first prediction L0', the optional updated second prediction L1', the optional first horizontal and vertical gradient values, and the optional second horizontal and vertical gradient values, the bidirectional prediction of the current prediction block is computed (709).
[0074] In another example, one or more predefined inter-frame prediction techniques include BDOF. Figure 8 This is a flowchart illustrating the illustrative workflow of the CIIP design in this example of the present disclosure. In this workflow, L0 motion compensation and L1 motion compensation (801 and 802) are averaged (803) instead of being processed by BDOF, and the resulting average is then processed together with intra-frame prediction (804) to form a weighted average (805) in the CIIP mode.
[0075] In the third example, one or more predefined inter-frame prediction techniques include DMVR.
[0076] According to another embodiment of this disclosure, constraining the application of CIIP mode to a CU based on determination includes: selecting a plurality of unidirectional prediction samples from all available unidirectional prediction samples using predefined criteria, to be used in combination with intra-frame prediction samples for the CU during the application of CIIP mode to the CU when the CU is bidirectionally predicted.
[0077] In one example, the predefined criteria include: selecting all one-way prediction samples based on the images in the reference image list 0 (L0).
[0078] In another example, the predefined criteria include: selecting all one-way prediction samples based on the images in reference image list 1 (L1).
[0079] In the third example, the predefined criteria include: selecting all one-way prediction samples based on a reference image that has the smallest picture order count (POC) distance to the image where the CU is located ("current image"). Figure 9 This is a flowchart illustrating the illustrative workflow of the CIIP design proposed in this example of the present disclosure. In this workflow, a determination is made regarding whether the L0 reference image or the L1 reference image is closer to the current image in terms of POC distance (901), and motion compensation from a reference image having the smallest POC distance to the current image is selected (902 and 903), and then processed together with intra-frame prediction (904) to form a weighted average in the CIIP mode (905).
[0080] According to another embodiment of this disclosure, constraining the application of CIIP mode to the CU based on determination includes: disabling the application of CIIP mode to the CU when the CU is bidirectionally predicted, and disabling signaling for the CIIP flag for the CU. Specifically, to reduce overhead, the signaling for the CIIP enable / disable flag depends on the prediction direction of the current CIIP CU. If the current CU is unidirectionally predicted, the CIIP flag will be signaled in the bitstream to indicate whether CIIP is enabled or disabled. Otherwise (i.e., the current CU is bidirectionally predicted), the signaling for the CIIP flag will be skipped and always inferred as false, i.e., CIIP is always disabled.
[0081] Furthermore, according to this disclosure, when forming an MPM candidate list for a CU (“current CU”), a determination is made regarding whether each of the neighboring CUs of the current CU is individually encoded or decoded by CIIP, and then for each neighboring CU that is encoded or decoded by CIIP, a uniform criterion is adopted in the process of forming an MPM candidate list for the current CU using the intra-frame mode of the neighboring CUs, which is independent of the determination of whether the current CU is intra-frame encoded or decoded by CIIP.
[0082] According to one embodiment of this aspect of the disclosure, the unification criteria include: when adjacent CUs are encoded and decoded by CIIP, the intra-frame modes of adjacent CUs are considered unusable for forming an MPM candidate list.
[0083] According to another embodiment of this aspect of the present disclosure, the unification criteria include: when adjacent CUs are encoded and decoded by CIIP, the CIIP modes of adjacent CUs are regarded as equivalent to the intra-frame modes in the formation of the MPM candidate list.
[0084] Figures 10A-10BThese are a pair of flowcharts illustrating the illustrative workflow of two embodiments of this aspect of the present disclosure. Figure 10A The flowchart of a first embodiment of this aspect of the present disclosure is illustrated. In this flowchart, a determination of whether a neighboring CU is an intra-block or a CIIP block is first made (1001), and its intra-block mode is added to the MPM candidate list of the current CU if and only if it is an intra-block or a CIIP block (1002), regardless of whether the current CU is an intra-block or a CIIP block. Figure 10B The flowchart of a second embodiment of this aspect of the present disclosure is illustrated. In this flowchart, a determination of whether a neighboring CU is an intra-block is first made (1003), and its intra-block mode is added to the MPM candidate list of the current CU if and only if it is an intra-block, regardless of whether the current CU is an intra-block or a CIIP block.
[0085] In one or more examples, the described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions can be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium or a communication medium, where a computer-readable storage medium corresponds to a tangible medium such as a data storage medium, and a communication medium includes any medium that facilitates the transfer of a computer program (e.g., according to a communication protocol) from one place to another. In this manner, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the embodiments described herein. Computer program products may include computer-readable media.
[0086] Furthermore, the methods described above can be implemented using a device comprising one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The device can perform the methods described above using circuits combined with other hardware or software components. Each module, submodule, unit, or subunit disclosed above can be implemented at least partially using one or more circuits.
[0087] Other embodiments of the invention will be apparent to those skilled in the art in light of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention following its general principles, and includes deviations from this disclosure in what is known or customary practice in the art. It is intended that the specification and examples be considered exemplary only, and that the true scope and spirit of the invention are indicated by the following claims.
[0088] It will be understood that the invention is not limited to the exact examples described above and illustrated in the accompanying drawings, and that various modifications and changes can be made without departing from its scope. It is intended that the scope of the invention be defined only by the appended claims.
Claims
1. A video decoding method, comprising: Obtain a first reference image and a second reference image associated with the current block, wherein, in display order, the first reference image precedes the current image and the second reference image follows the current image; In response to the combined inter-frame and intra-frame prediction flags for the current block indicating that the combined inter-frame and intra-frame predictions should not be applied to the current block, it is determined that a decoder-side motion vector refinement operation is to be applied to the current block. In response to the combined inter-frame and intra-frame prediction flags for the current block indicating that the combined inter-frame and intra-frame predictions should not be applied to the current block, it is determined that a bidirectional optical flow operation should be applied to the current block. Based on the first reference image, the second reference image, and the determination made by applying the decoder-side motion vector thinning operation and the bidirectional optical flow operation, the bidirectional prediction of the current block is calculated. The calculation of the bidirectional prediction for the current block includes: The first motion vector and the second motion vector are adjusted to generate a first prediction and a second prediction, wherein the first motion vector is from the current block to a reference block in the first reference image preceding the current image in display order, and the second motion vector is from the current block to a reference block in the second reference image following the current image in display order. Calculate the first horizontal gradient value and the first vertical gradient value associated with the first prediction; and Calculate the second horizontal gradient value and the second vertical gradient value associated with the second prediction.
2. The method according to claim 1, wherein calculating the bidirectional prediction of the current block based on the first reference image, the second reference image, and the determination of the decoder-side motion vector thinning operation and the determination of the bidirectional optical flow operation comprises: Based on the first prediction, the second prediction, the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, and the second vertical gradient value, calculate the bidirectional prediction of the current block.
3. A video encoding method, comprising: Obtain a first reference image and a second reference image associated with the current block, wherein, in display order, the first reference image precedes the current image and the second reference image follows the current image; When the combined inter-frame and intra-frame prediction flags for the current block indicate that the combined inter-frame and intra-frame predictions should not be applied to the current block, a decoder-side motion vector refinement operation is applied to the current block. When the combined inter-frame and intra-frame prediction flags for the current block indicate that the combined inter-frame and intra-frame predictions should not be applied to the current block, a bidirectional optical flow operation is applied to the current block. Based on the first reference image, the second reference image, and the determination made by applying the decoder-side motion vector thinning operation and the bidirectional optical flow operation, the bidirectional prediction of the current block is calculated. The calculation of the bidirectional prediction for the current block includes: The first motion vector and the second motion vector are adjusted to generate a first prediction and a second prediction, wherein the first motion vector is from the current block to a reference block in the first reference image preceding the current image in display order, and the second motion vector is from the current block to a reference block in the second reference image following the current image in display order. Calculate the first horizontal gradient value and the first vertical gradient value associated with the first prediction; and Calculate the second horizontal gradient value and the second vertical gradient value associated with the second prediction.
4. The method according to claim 3, wherein calculating the bidirectional prediction of the current block based on the first reference image, the second reference image, and the determination of the decoder-side motion vector thinning operation and the determination of the bidirectional optical flow operation includes: Based on the first prediction, the second prediction, the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, and the second vertical gradient value, calculate the bidirectional prediction of the current block.
5. A computing device, comprising: Storage medium; as well as One or more processors coupled to the storage medium, wherein the one or more processors are configured to perform the method of any one of claims 1-4.
6. A method for storing a bit stream, comprising: Perform the video encoding method according to any one of claims 3-4 to generate a bitstream; as well as Store the bit stream.
7. A non-transitory computer-readable storage medium storing instructions and a bit stream formed by the instructions, which, when executed by a computing device having one or more processors, cause the one or more processors to perform the video coding method of any one of claims 3-4.