Method and device for selectively applying bidirectional optical flow and decoder-side motion vector correction for video encoding - Patents.com

By classifying blocks into DMVR or BDOF classes and applying either tool but not both, the method optimizes the balance between coding efficiency and complexity/latency in video encoding, addressing inefficiencies in current VVC designs.

JP7776572B2Active Publication Date: 2025-11-26BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024076928
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-02-08
Filing Date
2024-05-10
Publication Date
2025-11-26
Estimated Expiration
2040-02-08

AI Technical Summary

Technical Problem

Current VVC designs face challenges in achieving an optimal balance between coding efficiency and complexity/latency due to constraints on enabling Bi-Directional Optical Flow (BDOF) and Decoder-side Motion Vector Refinement (DMVR) tools, leading to unnecessary increases in complexity or missed coding gains.

Method used

Classify current blocks into DMVR or BDOF classes based on predefined criteria, applying either DMVR or BDOF but not both, to optimize tool usage and reduce latency and complexity.

Benefits of technology

Improves coding efficiency by optimizing the application of BDOF and DMVR, reducing latency and complexity in video encoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776572000001
    Figure 0007776572000001
  • Figure 0007776572000002
    Figure 0007776572000002
  • Figure 0007776572000003
    Figure 0007776572000003
Patent Text Reader

Abstract

To provide a system and a method for performing video encoding using selective application of bidirectional optical flow and decoder-side motion vector refinement for an inter mode coded block.SOLUTION: A video coding method executed on a computing device that selectively applies bidirectional optical flow (BDOF) and decoder-side motion vector refinement (DMVR) to an inter-mode coded block used in a video coding standard such as the current versatile video coding (VVC) design includes determining whether a current block is suitable for application of both DMVR and BDOF on the basis of a plurality of predefined conditions, classifying the current block using a predefined criterion when the current block is suitable for both, and using the classification when applying either BDOF or DMVR, but not both, on the block.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 803,417, filed February 8, 2019. The entire disclosure of the above application is incorporated herein by reference in its entirety.

[0002] This disclosure relates generally to video encoding and compression. More particularly, this disclosure relates to systems and methods for performing video encoding using selective application of bidirectional optical flow and decoder-side motion vector correction to inter-mode coded blocks. [Background technology]

[0003]

[0003] This section provides background information related to the present disclosure. Information contained in this section should not necessarily be construed as prior art.

[0004]

[0004] To compress video data, any of a variety of video encoding techniques can be used. The video encoding can be performed according to one or more video encoding standards. Some exemplary video encoding standards include Versatile Video Coding (VVC), Joint Exploration Model (JEM) coding, High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), and Moving Picture Experts Group (MPEG) coding.

[0005]

[0005] Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in a video image or sequence. One goal of video coding techniques is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation of video quality.

[0006]

[0006] Prediction methods utilized in video coding typically involve performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data, and are typically associated with block-based video coding.

[0007] In block-based video coding, the input video signal is processed block by block, and for each block (also known as a coding unit (CU)), spatial and / or temporal prediction can be performed.

[0008]

[0008] Spatial prediction (also known as "intra prediction") predicts the current block using pixels from samples of already coded neighboring blocks (called reference samples) within the same video picture / slice. Spatial prediction reduces the spatial redundancy inherent in video signals.

[0009]

[0009] During the decoding process, the video bitstream is first entropy decoded in an entropy decoding unit. The coding mode and prediction information are sent to a spatial prediction unit (when intra-coded) or a temporal prediction unit (when inter-coded) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then combined together. The reconstructed block may further undergo in-loop filtering before being stored in a reference picture store. The reconstructed video in the reference picture store is then sent out to drive a display device and used to predict future video blocks.

[0010]

[0010] Newer video coding standards, such as the current VVC design, have introduced new inter-mode coding tools, such as Bi-Directional Optical Flow (BDOF) and Decoder-side Motion Vector Refinement (DMVR). Such new inter-mode coding tools generally help increase the efficiency of motion-compensated prediction and thus improve coding gain. However, such improvements can come at the cost of increased complexity and latency.

[0011]

[0011] To achieve a good balance between the improvement and cost associated with new inter-mode tools, the current VVC design imposes constraints on when new inter-mode coding tools such as DMVR and BDOF are enabled for the inter-mode coding block.

[0012] However, such constraints present in current VVC designs do not necessarily achieve the best balance between improvement and cost. On the one hand, the constraints present in current VVC designs allow both DMVR and BDOF to be applied to the same inter-mode coding block, but increase latency due to the dependency between the operations of the two tools. On the other hand, the constraints present in current VVC designs are too lenient in certain cases for DMVR, resulting in unnecessary increases in complexity and latency, and are too strict in certain cases for BDOF, resulting in missed opportunities for further coding gain. Summary of the Invention

[0013]

[0013] This section provides an overview of the disclosure and is not a complete scope or comprehensive disclosure of all features of the disclosure.

[0014] According to a first aspect of the present disclosure, a video encoding method is performed on a computing device having one or more processors and a memory storing a plurality of programs to be executed by the one or more processors. The method includes classifying a current block suitable for application of both DMVR and BDOF into one of two predefined classes, a DMVR class and a BDOF class, based on a plurality of predefined conditions, using predefined criteria based on mode information of the current block. The method further includes using the classification of the current block when applying either DMVR or BDOF, but not both, to the current block.

[0015] According to a second aspect of the present disclosure, a video encoding method is executed on a computing device having one or more processors and a memory storing a plurality of programs to be executed by the one or more processors. The method includes determining whether weighted prediction is possible for a current block suitable for DMVR encoding based on a plurality of predefined conditions. The method further includes determining whether different weights are used when averaging List 0 predictor samples and List 1 predictor samples for the current block. The method further includes determining whether to disable application of DMVR to the current block based on the two determinations.

[0016] According to a third aspect of the present disclosure, a video encoding method is performed on a computing device having one or more processors and a memory storing a plurality of programs to be executed by the one or more processors, the method including enabling application of BDOF to a current block when the current block is encoded as a sub-block merging mode.

[0017] According to a fourth aspect of the present application, a computing device includes one or more processors, a memory, and a plurality of programs stored in the memory that, when executed by the one or more processors, cause the computing device to perform the operations set forth above in the first three aspects of the present application.

[0018] According to a fifth aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, the programs, when executed by the one or more processors, causing the computing device to perform the operations set forth above in the first three aspects of the present application.

[0019]

[0019] Hereinafter, several sets of exemplary, non-limiting embodiments of the present disclosure will be described in conjunction with the accompanying drawings. Those skilled in the art in the relevant art may implement modifications of structure, method, or function based on the examples presented herein, and all such modifications are encompassed within the scope of the present disclosure. Where no contradiction exists, the teachings of different embodiments may be combined with each other, although not necessarily. [Brief explanation of the drawings]

[0020] [Figure 1]

[0020] FIG. 1 is a block diagram illustrating an exemplary block-based hybrid video encoder that can be used with many video coding standards. [Figure 2]

[0021] FIG. 1 is a block diagram illustrating an exemplary video decoder that can be used with many video encoding standards. [Figure 3]

[0022] 1 is a diagram of block partitions in several types of tree structures that can be used with many video coding standards. [Figure 4]

[0023] FIG. 1 is a diagram of a Bi-Directional Optical Flow (BDOF) process. [Figure 5]

[0024] FIG. 1 is a diagram of mutual matching used in decoder-side motion vector refinement (DMVR). [Figure 6]

[0025] 1 is a flow chart illustrating the operation of a first aspect of the present disclosure. [Figure 7]

[0026] 10 is a flow chart illustrating the operation of a second aspect of the present disclosure. [Figure 8]

[0027] 10 is a flow chart illustrating the operation of a third aspect of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0021]

[0028] The terms used in this disclosure are intended to illustrate specific examples, not to limit the disclosure. Unless otherwise clearly implied by the context, the singular forms "a," "an," and "the" used in this disclosure and the appended claims also refer to the plural. As used herein, the term "and / or" should be understood to refer to all possible combinations of one or more of the associated listed items.

[0022]

[0029] As used herein, terms such as "first," "second," and "third" may be used to describe various pieces of information; however, it should be understood that the information is not limited by these terms. These terms are used solely to distinguish one category of information from another. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information, without departing from the scope of this disclosure. As used herein, the term "if" will be understood to mean "when," "upon," or "in response to," depending on the context.

[0023]

[0030] Throughout this specification, references to "one embodiment," "an embodiment," "another embodiment," etc., mean that one or more particular features, structures, or characteristics described in connection with one embodiment are included in at least one embodiment of the present disclosure. Thus, the appearances of phrases such as "one embodiment," "in an embodiment," or "another embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, particular features, structures, or characteristics in one or more embodiments may be combined in any suitable manner.

[0024]

[0031] Conceptually, many video coding standards, including those already mentioned in the background section, are similar: for example, virtually all video coding standards use block-based processing to achieve video compression and share similar video coding block diagrams.

[0025]

[0032] 1 shows a block diagram of an exemplary block-based hybrid video encoder 100 that can be used with many video coding standards. In encoder 100, a video frame is divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either inter-prediction or intra-prediction techniques. In inter-prediction, one or more predictors are formed based on pixels from a previously reconstructed frame using motion estimation and motion compensation. In intra-prediction, a predictor is formed based on reconstructed pixels in the current frame. A mode decision allows the best predictor to be selected for predicting the current block.

[0026]

[0033] A prediction residual, which represents the difference between the current video block and its predictor, is sent to transform circuitry 102. The transform coefficients are then sent from transform circuitry 102 to quantization circuitry 104 for entropy reduction. The quantized coefficients are then sent to entropy coding circuitry 106 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 110 from inter-prediction and / or intra-prediction circuitry 112, such as video block partition information, motion vectors, reference picture indexes, and intra-prediction modes, is also sent through entropy coding circuitry 106 and stored in compressed video bitstream 114.

[0027]

[0034] Decoder-related circuitry is also required in encoder 100 to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed by inverse quantization circuit 116 and inverse transform circuit 118. This reconstructed prediction residual is combined with block predictor 120 to generate unfiltered reconstructed pixels for the current video block.

[0028]

[0035] Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") predicts a current video block using reconstructed pixels from an already coded video picture. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal references. In addition, if multiple reference pictures are supported, a reference picture index is also sent and used to identify which reference picture in the reference picture store the temporal prediction signal comes from.

[0029]

[0036] After spatial and / or temporal prediction is performed, an intra / inter mode decision circuit 121 in encoder 100 chooses the best prediction mode, for example, based on a rate-distortion optimization method. A block predictor 120 is then subtracted from the current video block, and the resulting prediction residual is de-correlated using transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are inverse quantized by inverse quantization circuit 116 and inverse transformed by inverse transform circuit 118 to form a reconstructed residual, which is then added back to the predictive block to form a reconstructed signal for the CU. Further in-loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), can be applied to the reconstructed CU, which is then placed in reference picture storage in picture buffer 117 and used to encode future video blocks. To form the output video bitstream 114, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 106 for further compression and packing to form the bitstream.

[0030]

[0037] For example, deblocking filters are available in AVC, HEVC, and the current version of VVC. HEVC defines an additional in-loop filter called SAO (Sample Adaptive Offset) to further improve coding efficiency. In the current version of the VVC standard, yet another in-loop filter called ALF (Adaptive Loop Filter) is being actively investigated and has a good chance of being included in the final standard.

[0031]

[0038] These in-loop filter operations are optional. Performing these operations helps improve coding efficiency and visual quality. These operations can also be turned off as a decision made by encoder 100 to reduce computational complexity.

[0032]

[0039] Note that when these filter options are turned on by the encoder 100, intra prediction is typically based on unfiltered reconstructed pixels, whereas inter prediction is based on filtered reconstructed pixels.

[0033]

[0040] Figure 2 is a block diagram illustrating an exemplary video decoder 200 that can be used with many video coding standards. This decoder 200 is similar to the reconstruction-related portions residing in the encoder 100 of Figure 1. In the decoder 200 (Figure 2), an incoming video bitstream 201 is first decoded by entropy decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by inverse quantization 204 and inverse transform 206 to obtain a reconstructed prediction residual. A block predictor mechanism implemented in an intra / inter mode selector 212 is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by adding the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block predictor mechanism using an adder 214.

[0034]

[0041] The reconstructed blocks may further undergo an in-loop filter 209 before being stored in a picture buffer 213, which acts as a reference picture store. The reconstructed video in the picture buffer 213 may then be sent out to drive a display device as well as used to predict future video blocks. In situations where the in-loop filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0035]

[0042] In video coding standards such as HEVC, blocks can be divided based on a quadtree. Newer video coding standards, such as the current VVC, use more division methods, and one coding tree unit (CTU) can be divided into multiple CUs to adapt to varying local characteristics based on a quadtree, binary tree, or ternary tree. The separation of CUs, prediction units (PUs), and transform units (TUs) does not exist in most coding modes of the current VVC, and each CU is always used as a basic unit for both prediction and transformation without further division. However, in some specific coding modes, such as an intra-subdivision coding mode, each CU can further include multiple TUs. In multiple types of tree structures, first, one CTU is divided by a quadtree structure. Then, the leaf nodes of each quadtree can be further divided by binary and ternary tree structures.

[0036]

[0043] 3 shows five partition types used in current VVC: 4-partition 301, horizontal 2-partition 302, vertical 2-partition 303, horizontal 3-partition 304, and vertical 3-partition 305. In a situation where multiple types of tree structures are used, one CTU is first partitioned by a quadtree structure. Then, the leaf nodes of each quadtree can be further partitioned by a binary tree and a ternary tree structure.

[0037]

[0044] 3 can be used to perform spatial prediction and / or temporal prediction using the configuration shown in FIG. 1. Spatial prediction (or "intra-prediction") predicts a current video block using pixels from samples (called reference samples) of already coded neighboring blocks within the same video picture / slice. Spatial prediction reduces spatial redundancy inherent in video signals.

[0038]

[0045] Newer video coding standards, such as the current VVC, have introduced new inter-mode coding tools; two examples of new inter-mode coding tools are bidirectional optical flow (BDOF) and decoder-side motion vector compensation (DMVR).

[0039]

[0046] Traditional bidirectional prediction in video coding is a simple combination of two temporal prediction blocks obtained from an already reconstructed reference picture. However, due to the limitations of block-based motion compensation, slight motion may remain observable between the samples of the two prediction blocks, thus reducing the efficiency of motion-compensated prediction. To solve this problem, current VVC designs apply BDOF to reduce the effect of such motion on all samples within a block.

[0040]

[0047] 4 is a diagram of the BDOF process. BDOF is a sample-by-sample motion correction performed in addition to block-based motion compensation prediction when bidirectional prediction is used. The motion correction for each 4x4 subblock 401 is calculated by minimizing the difference between the reference picture list 0 (L0) predicted samples 402 and the reference picture list 1 (L1) predicted samples 403 after BDOF is applied within a 6x6 window around the subblock. Based on the motion correction thus derived, the final bidirectional predicted samples for the CU are calculated by interpolating the L0 / L1 predicted samples along the motion trajectory based on the optical flow model.

[0041]

[0048] DMVR is a bidirectional prediction technique for merging blocks with two initially signaled MVs that can be further corrected by using mutually matched prediction.

[0042]

[0049] Figure 5 is a diagram of mutual matching used in DMVR. Mutual matching is used to derive motion information for a current CU 501 by finding the closest match between two blocks 503 and 504 along the current CU's 501 motion trajectory 502 in two different reference pictures 505 and 506. The cost function used in the matching process is row-subsampled sum of absolute differences (SAD). After the matching process is performed, the corrected MVs 507 and 508 are used for motion compensation in the prediction stage, boundary strength calculation for the deblocking filter, temporal motion vector prediction for subsequent pictures, and cross-CTU spatial motion vector prediction for subsequent CUs. Under the assumption that the motion trajectories are continuous, the motion vectors MV0 507 and MV1 508 pointing to the two reference blocks 503 and 504 are proportional to the temporal distances, i.e., TD0 509 and TD1 510, between the current picture 511 and the two reference pictures 505 and 506. As a special case, when the current picture 511 is temporally between two reference pictures 505 and 506 and the temporal distances from the current picture 511 to the two reference pictures 505 and 506 are the same, the mutual alignment becomes a mirror-based bidirectional MV.

[0043]

[0050] To strike an appropriate balance between, on the one hand, the increased coding efficiency that newer inter mode coding tools such as BDOF and DMVR can bring, and, on the other hand, the increased complexity and latency that comes with the newer inter mode tools, current VVC designs apply constraints on when BDOF or DMVR can be enabled for the current block.

[0044]

[0051] In current VVC designs, BDOF is only possible when all of the following predefined BDOF conditions, listed in the box immediately below this paragraph, are true:

[0045] 1. The current block uses bidirectional prediction, with one MV pointing to a reference picture that is before the current picture in display order and another MV pointing to a reference picture that is after the current picture in display order. 2. Weighted prediction is not possible. 3. The height of the current block is not equal to 4. 4. The size of the current block is not equal to 4x8 (i.e., 4 width and 8 height). 5. The current block is not coded as a symmetric motion vector differential (MVD) mode, which is a special MVD coding mode in VVC. 6. The current block is not coded as affine mode. 7. The current block is not coded as sub-block merging mode. 8. The current block does not use different heights when averaging predictor samples from list 0 and list 1 (e.g., bidirectional prediction with weighted averaging (BWA) with unequal weights).

[0046]

[0052] In the current VVC design, DMVR is enabled only when all of the following predefined DMVR conditions, listed in the box immediately below this paragraph, are true:

[0047] 1. The current block uses bidirectional prediction, one MV points to a reference picture that is before the current picture in display order, and another MV points to a reference picture that is after the current picture in display order, and the distance between the current picture and the forward reference picture and the distance between the current picture and the backward reference picture are the same. 2. The current block is coded as a merge mode, and the selected merge candidate is a normal merge candidate (e.g., a normal spatial merge candidate that is not a sub-block, or a temporal merge candidate, etc.). 3. The height of the current block is 8 or more. 4. The area of ​​the current block is 64 or more. 5. The current block is not coded as affine mode. 6. The current block is not coded as sub-block merging mode. 7. The current block is not coded as merge mode with motion vector differential (MMVD) mode.

[0048]

[0053] While the above constraints present in current VVC designs go a long way toward achieving a desired balance between coding efficiency on the one hand and complexity and latency on the other, they do not completely solve the problem.

[0049]

[0054] One remaining problem with current VVC designs is that, although some constraints are already applied to enable BDOF and DMVR, in some cases, two decoder-side inter-prediction tools, BDOF and DMVR, can both be enabled when encoding a block. In current VVC designs, when both decoder-side inter-prediction tools are enabled, BDOF depends on the final motion compensation sample of DMVR, creating latency issues for hardware designs.

[0050]

[0055] A second remaining problem with the current VVC design is that although some constraints are already applied to enable DMVR, these constraints are still collectively too lenient with respect to enabling DMVR, because the current VVC design enables DMVR in scenarios where disabling DMVR and then reducing complexity and latency would provide a better balance between coding efficiency on the one hand and complexity and latency on the other.

[0051]

[0056] A third remaining problem with current VVC designs is that the constraints already applied to enable BDOF are generally too strict with respect to enabling BDOF, because there are scenarios where enabling BDOF and subsequently increasing coding gain would provide a better balance between coding efficiency on the one hand and complexity and latency on the other, and the current VVC design does not enable BDOF in these scenarios.

[0052]

[0057] According to a first aspect of the present disclosure, when a current block is suitable for application of both DMVR and BDOF based on multiple predefined conditions, the current block is classified into one of two predefined classes, namely, a DMVR class and a BDOF class, using predefined criteria based on the mode information of the current block. The current block's classification is then used when applying either DMVR or BDOF, but not both, to the current block. This method can be combined with the current VVC in addition to the above predefined conditions, or can be implemented independently.

[0053]

[0058] FIG. 6 is a flow diagram illustrating the operation of a first aspect of the present disclosure. When processing a current block (601), the operation of this aspect of the present disclosure may apply multiple predefined conditions to the current block (602) and determine whether the current block is suitable for application of both DMVR and BDOF based on the multiple predefined conditions (603). If the current block is determined to be unsuitable for application of both DMVR and BDOF based on the multiple predefined conditions, the operation of this aspect of the present disclosure may continue the process existing in the current VVC design (604). On the other hand, if the current block is determined to be suitable for application of both DMVR and BDOF based on the multiple predefined conditions, the operation of this aspect of the present disclosure may classify the current block into one of two predefined classes, i.e., a DMVR class and a BDOF class, using predefined criteria based on the mode information of the current block (605). Next, the operation of this aspect of the present disclosure may apply either DMVR or BDOF to the current block, but not both, using the classification result of the current block (606). Applying either DMVR or BDOF, but not both, to the current block using the classification result of the current block (606) may include, if the current block is classified into the DMVR class, continuing the DMVR class (applying DMVR instead of BDOF) (607), and, if the current block is classified into the BDOF class, continuing the BDOF class (applying BDOF instead of DMVR) (608).

[0054]

[0059] The multiple predefined conditions under which the current block becomes eligible for both DMVR and BDOF application can be, but are not necessarily, the multiple predefined BDOF conditions and predefined DMVR conditions listed in the box above.

[0055]

[0060] Mode information used as a basis for the predefined criteria includes, but is not limited to, prediction mode, such as whether or not to use merge mode, merge mode index, motion vector, block shape, block size, and predictor sample value.

[0056]

[0061] According to one or more embodiments of the present disclosure, using the classification of the current block when applying either DMVR or BDOF, but not both, to the current block includes optionally signaling a flag to indicate the classification of the current block; i.e., in some examples of this embodiment, one flag is signaled to specify whether BDOF or DMVR has been applied to the block, and in some other examples of this embodiment, no such flag is signaled.

[0057]

[0062] According to one or more embodiments of the present disclosure, using the classification of the current block when applying either DMVR or BDOF, but not both, to the current block further includes applying DMVR, but not BDOF, to the current block when the current block is classified into a DMVR class, either by an existing merge candidate list or by a separately generated merge candidate list. That is, in some examples of this embodiment, when the current block is classified into a DMVR class and DMVR is applied to the current block, a separate merge candidate list is exclusively generated and used. To indicate this DMVR merge mode, syntax is signaled, and if the DMVR merge candidate list size is greater than 1, a merge index is also signaled. In some other examples of this embodiment, when the current block is classified into a DMVR class and DMVR is applied to the current block, such a separate merge candidate is not generated, and application of DMVR to the current block uses the existing merge candidate list without further syntax or signaling.

[0058]

[0063] According to one or more embodiments of the present disclosure, using the classification of the current block when applying either DMVR or BDOF, but not both, to the current block further includes applying BDOF, rather than DMVR, to the current block when the current block is classified into a BDOF class, either by an existing merge candidate list or by a separately generated merge candidate list. That is, in some examples of this embodiment, when the current block is classified into a BDOF class and BDOF is applied to the current block, a separate merge candidate list is exclusively generated and used. To indicate this BDOF merge mode, syntax is signaled, and if the BDOF merge candidate list size is greater than 1, a merge index is also signaled. In some other examples of this embodiment, when the current block is classified into a BDOF class and BDOF is applied to the current block, such separate merge candidates are not generated, and application of BDOF to the current block uses the existing merge candidate list without further syntax or signaling.

[0059]

[0064] According to another embodiment of the first aspect of the present disclosure, classifying the current block into one of two predefined classes, namely, a DMVR class and a BDOF class, using predefined criteria based on mode information of the current block includes classifying the current block into the DMVR class when the predefined criteria are satisfied, and classifying the current block into the BDOF class when the predefined criteria are not satisfied.

[0060]

[0065] According to another embodiment of the first aspect of the present disclosure, classifying the current block into one of two predefined classes, namely, a DMVR class and a BDOF class, using predefined criteria based on mode information of the current block includes classifying the current block into the BDOF class when the predefined criteria are satisfied, and classifying the current block into the DMVR class when the predefined criteria are not satisfied.

[0061]

[0066] According to another embodiment of the first aspect of the present disclosure, the predefined criteria include whether the normal mode is selected for the current block.

[0062]

[0067] According to another embodiment of the first aspect of the present disclosure, the predefined criteria include whether the encoded merge index of the current block possesses a predefined mathematical property.

[0063]

[0068] In another example, the predefined mathematical property includes the property of being greater than or equal to a predefined threshold number.

[0064]

[0069] According to another embodiment of the first aspect of the present disclosure, the predefined criteria include whether the motion vector of the current block satisfies a predefined test.

[0065]

[0070] In one example, the predefined test includes whether the sum of the magnitudes of all motion vector components is greater than a predefined threshold number.

[0066]

[0071] According to another embodiment of the first aspect of the present disclosure, the predefined criteria include whether the current block is of a predefined shape.

[0067]

[0072] In one example, the predefined shape is a square shape.

[0068]

[0073] According to another embodiment of the first aspect of the present disclosure, the predefined criteria include whether the block size of the current block possesses a predefined mathematical property.

[0069]

[0074] In one example, the predefined mathematical property includes the property of being greater than or equal to a predefined threshold number.

[0070]

[0075] According to another embodiment of the first aspect of the present disclosure, the predefined criteria include whether the sum of absolute differences or sum of squared differences (SAD or SSD) between the List 0 predictor samples and the List 1 predictor samples of the current block possesses predefined mathematical properties.

[0071]

[0076] In one example, the predefined mathematical property includes being greater than a predefined threshold number.

[0072]

[0077] According to a second aspect of the present disclosure, when a current block is suitable for DMVR coding based on a plurality of predefined conditions, a first determination is made as to whether weighted prediction is possible for the current block, a second determination is made as to whether different weights are used when averaging list 0 predictor samples and list 1 predictor samples for the current block, and a determination is made as to whether to disable application of DMVR to the current block based on the first and second determinations. This method can be combined with or implemented independently of the current VVC in addition to the above predefined conditions.

[0073]

[0078] 7 is a flow chart illustrating an exemplary method of a second aspect of the present disclosure. When processing a current block (701), the operation of this aspect of the present disclosure may apply multiple predefined conditions to the current block (702) and determine whether the current block is suitable for DMVR encoding based on the multiple predefined conditions (703). If the current block is determined to be unsuitable for DMVR encoding based on the multiple predefined conditions, the operation of this aspect of the present disclosure may continue the process that exists in the current VVC design (704). On the other hand, if the current block is determined to be suitable for DMVR encoding based on the multiple predefined conditions, the operation of this aspect of the present disclosure may determine whether weighted prediction is possible for the current block (705) and may also determine whether different weights are used when averaging list 0 predictor samples and list 1 predictor samples for the current block (706), and then determine whether to disable application of DMVR to the current block based on the two determinations (707).

[0074]

[0079] The predefined conditions under which the current block is suitable for DMVR encoding may be, but are not necessarily, the predefined DMVR conditions listed in the box above.

[0075]

[0080] According to one embodiment of the second aspect of the present disclosure, determining whether to disable application of DMVR to the current block based on the two determinations includes disabling application of DMVR to the current block when it is determined that weighted prediction is possible for the current block.

[0076]

[0081] According to another embodiment of the second aspect of the present disclosure, determining whether to disable application of DMVR to the current block based on the two determinations includes disabling application of DMVR to the current block when it is determined that different weights are used when averaging the List 0 predictor samples and the List 1 predictor samples for the current block.

[0077]

[0082] According to another embodiment of the second aspect of the present disclosure, determining whether to disable application of DMVR to the current block based on the two determinations includes disabling application of DMVR to the current block when it is determined that weighted prediction is possible for the current block and at the same time it is determined that different weights are used when averaging the List 0 predictor samples and the List 1 predictor samples for the current block.

[0078]

[0083] According to a third aspect of the present disclosure, when the current block is coded using sub-block merging mode, it is possible to enable the application of BDOF to the current block. This method can be combined with the current VVC in addition to the above predefined conditions, or can be implemented independently.

[0079]

[0084] 8 is a flow diagram illustrating the operation of a third aspect of the present disclosure. When processing a current block (801), the operation of this aspect of the present disclosure may determine whether the current block is coded in sub-block merging mode (802). If it is determined that the current block is not coded in sub-block merging mode, the operation of this aspect of the present disclosure may continue the process that exists in the current VVC design (803). On the other hand, if it is determined that the current block is coded in sub-block merging mode, the operation of this aspect of the present disclosure may enable the application of BDOF to the current block (804).

[0080]

[0085] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communications protocol. In this manner, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media or (2) communication media, such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the implementations described herein. A computer program product may include computer-readable media.

[0081]

[0086] Furthermore, the above methods can be implemented using an apparatus including one or more circuits, such as an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. The apparatus can use these circuits in combination with other hardware or software components to perform the above-described methods. Each module, sub-module, unit, or sub-unit disclosed above can be implemented at least in part using one or more circuits.

[0082]

[0087] Other embodiments of the invention will be apparent to those skilled in the art from consideration of the specification and practice of the invention as disclosed above. This application is intended to cover any modifications, uses, or applications of the invention which comply with the general principles of the invention, including departures from the present disclosure which come within known or customary practice in the art. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.

[0083]

[0088] It will be understood that the present invention is not limited to the exact examples described above and illustrated in the accompanying drawings, and that various modifications and changes may be made thereto without departing from the scope of the present invention, which is intended to be limited only by the appended claims.

Claims

1. 1. A method for video decoding, comprising: determining whether weighted prediction is possible for a current block in response to the current block satisfying a plurality of predefined conditions, and determining whether different weights are used when averaging list 0 predicted samples and list 1 predicted samples for the current block; determining whether to disable application of decoder-side motion vector correction (DMVR) to the current block based on the two determinations; obtaining a corrected motion vector based on DMVR in response to determining that the application of DMVR to the current block is not invalid; A method comprising:

2. 2. The method of claim 1, wherein determining whether to disable application of DMVR to the current block based on the two determinations comprises disabling application of DMVR to the current block when it is determined that the weighted prediction is possible for the current block.

3. 2. The method of claim 1 , wherein determining whether to disable application of DMVR to the current block based on the two determinations comprises disabling application of DMVR to the current block when it is determined that different weights are used when averaging the list 0 predicted samples and the list 1 predicted samples for the current block.

4. 2. The method of claim 1 , wherein determining whether to disable application of DMVR to the current block based on the two determinations comprises disabling application of DMVR to the current block when it is determined that the weighted prediction is possible for the current block and at the same time it is determined that different weights are used when averaging the list 0 predicted samples and the list 1 predicted samples for the current block.

5. The plurality of predefined conditions include: the current block is not coded in affine mode, The current block is not coded as a sub-block merge mode 10. The method of claim 1, comprising:

6. a computing device one or more processors; a non-transitory storage device coupled to the one or more processors; a plurality of programs stored in the non-transitory storage device; Equipped with The plurality of programs, when executed by the one or more processors, cause the computing device to: determining whether weighted prediction is possible for a current block in response to the current block satisfying a plurality of predefined conditions, and determining whether different weights are used when averaging list 0 predicted samples and list 1 predicted samples for the current block; determining whether to disable application of decoder-side motion vector correction (DMVR) to the current block based on the two determinations; responsive to determining that the application of DMVR to the current block is not invalid, obtaining a corrected motion vector based on DMVR; A computing device that causes the device to perform operations including:

7. 7. The computing device of claim 6, wherein determining whether to disable application of DMVR to the current block based on the two determinations comprises disabling application of DMVR to the current block when it is determined that the weighted prediction is possible for the current block.

8. 7. The computing device of claim 6, wherein determining whether to disable application of DMVR to the current block based on the two determinations comprises disabling application of DMVR to the current block when it is determined that different weights are used when averaging the list 0 predicted samples and the list 1 predicted samples for the current block.

9. 7. The computing device of claim 6, wherein determining whether to disable application of DMVR to the current block based on the two determinations comprises disabling application of DMVR to the current block when it is determined that the weighted prediction is possible for the current block and at the same time it is determined that different weights are used when averaging the list 0 predicted samples and the list 1 predicted samples for the current block.

10. The plurality of predefined conditions include: the current block is not coded in affine mode, The current block is not coded as a sub-block merge mode The computing device of claim 6 , comprising:

11. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors of a computing device, cause the computing device to store a bitstream on the non-transitory computer-readable storage medium and to perform a method according to any one of claims 1 to 5 on the stored bitstream.

12. A computer program comprising instructions that, when executed on a computing device, cause the computing device to perform the method of any one of claims 1 to 5.

13. A method for storing a bitstream, comprising: performing an encoding method to generate a bitstream; storing the generated bitstream; Equipped with The encoding method comprises: determining whether weighted prediction is possible for the current block when the current block satisfies a plurality of predefined conditions, and determining whether different weights are used when averaging list 0 predicted samples and list 1 predicted samples for the current block; determining whether to disable application of decoder-side motion vector correction (DMVR) to the current block based on the two determinations; responsive to determining that the application of DMVR to the current block is not invalid, obtaining a corrected motion vector based on DMVR; A method comprising:

Citation Information

Patent Citations

  • Moving image encoder, moving image encoding method, moving image encoding computer program, moving image decoder, moving image decoding method and moving image decoding computer program

    JP2018107580A

  • Side motion refinement in video encoding / decoding systems.

    JP2022515875A

  • JPP7232345B

  • JPP7339395B

  • JPP7546721B