Method and device for selectively applying bidirectional optical flow and decoder-side motion vector correction for video encoding

By classifying blocks and selectively applying DMVR or BDOF in the VVC design, the method addresses the suboptimal balance in the current VVC design, enhancing video encoding efficiency and reducing complexity and latency.

JP7696412B2Active Publication Date: 2025-06-20BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Application Number
JP2023194974
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-02-08
Filing Date
2023-11-16
Publication Date
2025-06-20
Estimated Expiration
2040-02-08

AI Technical Summary

Technical Problem

The current VVC design imposes constraints on the use of inter-mode coding tools like BDOF and DMVR, which do not achieve the optimal balance between coding efficiency and complexity/latency, leading to unnecessary increases in complexity and latency, and potential losses in coding gain.

Method used

A method for classifying current blocks into DMVR or BDOF classes based on predefined conditions and mode information, allowing for the selective application of either DMVR or BDOF, but not both, to achieve a better balance between coding efficiency and complexity/latency.

Benefits of technology

This approach enables a more balanced application of BDOF and DMVR, reducing unnecessary complexity and latency while maintaining coding efficiency, thereby improving the overall performance of video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696412000001
    Figure 0007696412000001
  • Figure 0007696412000002
    Figure 0007696412000002
  • Figure 0007696412000003
    Figure 0007696412000003
Patent Text Reader

Abstract

To provide video coding methods, computing devices and storage media for achieving a desired balance between the coding efficiency and the complexity and latency.SOLUTION: The invention provides methods to be implemented in a computing device for selectively applying bidirectional optical flow (BDOF) and decoder-side motion vector refinement (DMVR) to inter mode coded blocks employed in video coding standards, such as the current versatile video coding (VVC) design. When a current block is eligible for both applications of DMVR and BDOF based on a plurality of predefined conditions, the computing device uses a pre-defined criterion to classify the current block and then uses the classification in applying not both but one of BDOF and DMVR to the block.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 803,417, filed Feb. 8, 2019. The entire disclosure of the above application is incorporated herein by reference in its entirety.

[0002]

[0002] This disclosure generally relates to video encoding and compression. More particularly, this disclosure relates to systems and methods for performing video encoding using selective application of bidirectional optical flow and decoder - side motion vector correction for inter - mode encoded blocks.

Background Art

[0003]

[0003] This chapter provides background information related to the present disclosure. The information included in this chapter is not necessarily to be construed as prior art.

[0004]

[0004] To compress video data, any of a variety of video encoding techniques can be used. Video encoding can be performed according to one or more video encoding standards. Some exemplary video encoding standards include Versatile Video Coding (VVC), Joint Exploration Model (JEM) encoding, High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), and Moving Picture Experts Group (MPEG) encoding.

[0005]

[0005] Video encoding generally utilizes prediction methods (e.g., inter - prediction, intra - prediction, etc.) that exploit redundancies inherent in video images or sequences. One goal of video encoding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality.

[0006]

[0006] The prediction methods used in video encoding typically involve performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove the redundancy inherent in video data, and are typically associated with block-based video encoding.

[0007]

[0007] In block-based video encoding, the input video signal is processed block by block. For each block (also known as a coding unit (CU)), spatial prediction and / or temporal prediction can be performed.

[0008]

[0008] Spatial prediction (also known as "intra prediction") uses pixels from samples of already-encoded adjacent blocks (referred to as reference samples) within the same video picture / slice to predict the current block. Spatial prediction reduces the spatial redundancy inherent in the video signal.

[0009]

[0009] During the decoding process, the video bitstream is first entropy decoded by an entropy decoding unit. The coding mode and prediction information are sent to a spatial prediction unit (when intra-encoded) or a temporal prediction unit (when inter-encoded) to form a predicted block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct the residual block. Then, the predicted block and the residual block are integrated together. The reconstructed block can further undergo in-loop filtering before being stored in the reference picture memory. Then, the reconstructed video in the reference picture memory is sent out to drive a display device and is also used to predict future video blocks.

[0010]

[0010] In more recent video coding standards such as the current VVC design, new inter-mode coding tools such as Bi-Directional Optical Flow (BDOF) and Decoder-side Motion Vector Refinement (DMVR) have been introduced. Such new inter-mode coding tools generally help increase the efficiency of motion compensation prediction and thus improve coding gain. However, such improvements can be achieved at the expense of increased complexity and latency.

[0011]

[0011] To achieve a suitable balance between the improvements and costs associated with the new inter-mode tools, the current VVC design imposes constraints on when to enable new inter-mode coding tools such as DMVR and BDOF for inter-mode coding blocks.

[0012]

[0012] However, such constraints existing in the current VVC design do not necessarily achieve the best balance between improvements and costs. On the one hand, the constraints existing in the current VVC design allow both DMVR and BDOF to be applied to the same inter-mode coding block, but due to the dependency between the operations of the two tools, the latency increases. On the other hand, the constraints existing in the current VVC design are too lenient in certain cases regarding DMVR, resulting in an unnecessary increase in complexity and latency, and too strict in certain cases regarding BDOF, resulting in the loss of opportunities for further coding gain.

Summary of the Invention

[0013]

[0013] This chapter provides an overview of the present disclosure and is not a complete disclosure of the full scope of the present disclosure or all features of the present disclosure.

[0014] According to a first aspect of the present disclosure, a video encoding method is executed on a computing device having one or more processors and a memory storing a plurality of programs to be executed by the one or more processors. The method includes classifying a current block suitable for the application of both DMVR and BDOF based on a plurality of predefined conditions using a predefined criterion based on the mode information of the current block into one of two predefined classes, namely, a DMVR class and a BDOF class. The method further includes using the classification of the current block when applying either one but not both of DMVR and BDOF to the current block.

[0015] According to a second aspect of the present disclosure, a video encoding method is executed on a computing device having one or more processors and a memory storing a plurality of programs to be executed by the one or more processors. The method includes determining whether weighted prediction is possible for a current block suitable for DMVR encoding based on a plurality of predefined conditions. The method further includes determining whether different weights are used when averaging list 0 predictor samples and list 1 predictor samples for the current block. The method further includes determining whether to disable the application of DMVR to the current block based on the two determinations.

[0016] According to a third aspect of the present disclosure, a video encoding method is executed on a computing device having one or more processors and a memory storing a plurality of programs to be executed by the one or more processors. The method includes enabling the application of BDOF to a current block when the current block is encoded in sub-block merge mode.

[0017]

[0017] According to a fourth aspect of the present application, a computing device includes one or more processors, a memory, and a plurality of programs stored in the memory. When these programs are executed by one or more processors, the computing device is caused to perform the operations described above in the first three aspects of the present application.

[0018]

[0018] According to a fifth aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors. When these programs are executed by one or more processors, the computing device is caused to perform the operations described above in the first three aspects of the present application.

[0019]

[0019] Hereinafter, exemplary non-limiting embodiments of several sets of the present disclosure will be described together with the accompanying drawings. Those skilled in the relevant art can implement variations of the structure, method, or function based on the examples presented herein, and all such variations are included within the scope of the present disclosure. When there is no contradiction, the teachings of different embodiments, although not essential, can also be combined with each other.

Brief Description of the Drawings

[0020]

Figure 1

[0020] FIG. is a block diagram showing an exemplary hybrid video encoder based on blocks that can be used with many video coding standards.

Figure 2

[0021] FIG. is a block diagram showing an exemplary video decoder that can be used with many video coding standards.

Figure 3

[0022] FIG. is a diagram of block partitioning in a plurality of types of tree structures that can be used with many video coding standards.

Figure 4

[0023] It is a diagram of a bi - directional optical flow (BDOF) process.

Figure 5

[0024] It is a diagram of the mutual consistency used in decoder - side motion vector refinement (DMVR).

Figure 6

[0025] It is a flowchart showing the operation of the first aspect of the present disclosure.

Figure 7

[0026] It is a flowchart showing the operation of the second aspect of the present disclosure.

Figure 8

[0027] It is a flowchart showing the operation of the third aspect of the present disclosure.

Mode for Carrying Out the Invention

[0021]

[0028] The terms used in the present disclosure are not intended to limit the present disclosure but are intended to illustrate specific examples. Unless otherwise clearly incorporated in the context, the singular forms "a", "an", and "the" used in the present disclosure and the appended claims also refer to the plural forms. In this specification, it should be understood that the term "and / or" refers to any possible combination of one or more of the recited related items.

[0022]

[0029] In this specification, terms such as "first", "second", "third", etc. may be used to describe various information, but it should be understood that this information is not limited by these terms. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of the present disclosure, the first information can be called the second information, and similarly, the second information can be called the first information. In this specification, it should be understood that the term "if" means "when" or "upon" or "in response to" depending on the context.

[0023]

[0030] Throughout this specification, references to one or more "an embodiment", "embodiment", "another embodiment", etc. mean that one or more specific features, structures, or characteristics described in connection with an embodiment are included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases "in one embodiment", "in an embodiment", or "in another embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Further, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable manner.

[0024]

[0031] Conceptually, many video coding standards are similar, including what has already been mentioned in the background chapter. For example, virtually all video coding standards use block-based processing to achieve video compression and share similar video coding block diagrams.

[0025]

[0032] FIG. 1 shows a block diagram of an exemplary hybrid video encoder 100 based on blocks that can be used with many video coding standards. In encoder 100, a video frame is divided into a plurality of video blocks for processing. For each given video block, a prediction is formed based on an inter-prediction technique or an intra-prediction technique. In inter-prediction, one or more predictors are formed by motion estimation and motion compensation based on pixels from a previously reconstructed frame. In intra-prediction, a predictor is formed based on reconstructed pixels within the current frame. By mode decision, the best predictor can be selected to predict the current block.

[0026]

[0033] The prediction residual, which represents the difference between the current video block and its predictor, is sent to the conversion circuit 102. Next, the conversion coefficients are sent from the conversion circuit 102 to the quantization circuit 104 for entropy reduction. Next, the quantized coefficients are sent to the entropy encoding circuit 106 to generate a compressed video bit stream. As shown in FIG. 1, prediction relation information 110 from the inter prediction circuit and / or the intra prediction circuit 112, such as video block partition information, motion vectors, reference picture indices, and intra prediction modes, is also sent through the entropy encoding circuit 106 and stored into the compressed video bit stream 114.

[0027]

[0034] In the encoder 100, for prediction purposes, a decoder related circuit is also required to reconstruct pixels. First, the prediction residual is reconstructed by the inverse quantization circuit 116 and the inverse conversion circuit 118. Combining this reconstructed prediction residual with the block predictor 120 generates unfiltered reconstructed pixels for the current video block.

[0028]

[0035] Temporal prediction (also called "inter prediction" or "motion compensated prediction") predicts the current video block using reconstructed pixels from previously encoded video pictures. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and the temporal reference of the current CU. Also, when corresponding to multiple reference pictures, one reference picture index is further sent and used to identify from which reference picture in the reference picture memory the temporal prediction signal comes.

[0029]

[0036] After spatial and / or temporal prediction is performed, the intra / inter mode determination circuit 121 within the encoder 100 selects the best prediction mode, for example, based on a rate distortion optimization method. Next, a block predictor 120 is subtracted from the current video block, and the resulting prediction residual correlation is removed using the conversion circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are inverse quantized by the inverse quantization circuit 116 and inverse transformed by the inverse transform circuit 118 to form a reconstructed residual, which is then added to the prediction block again to form a reconstructed signal for the CU. Further in-loop filtering 115, such as deblocking filtering, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), can be applied to the reconstructed CU. Thereafter, the reconstructed CU is placed within the reference picture memory of the picture buffer 117 and used for encoding future video blocks. To form the output video bitstream 114, all of the encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy encoding unit 106, further compressed and packed to form a bitstream.

[0030]

[0037] For example, in AVC, HEVC, and the current version of VVC, deblocking filtering is available. In HEVC, an additional in-loop filter called SAO (sample adaptive offset) is defined to further improve the encoding efficiency. In the current version of the VVC standard, yet another in-loop filter called ALF (adaptive loop filter) is being actively investigated and has a good probability of being included in the final standard.

[0031]

[0038] These in-loop filter operations are optional. Performing these operations helps improve the encoding efficiency and visual quality. These operations can also be turned off as a decision made by the encoder 100 to save computational complexity.

[0032]

[0039] When these filter options are turned on by the encoder 100, it should be noted that intra prediction is usually based on non-filtered reconstructed pixels, while inter prediction is based on filtered reconstructed pixels.

[0033]

[0040] FIG. 2 is a block diagram showing an exemplary video decoder 200 that can be used with many video coding standards. This decoder 200 is similar to the reconstruction-related part resident in the encoder 100 of FIG. 1. In the decoder 200 (FIG. 2), first, the incoming video bitstream 201 is decoded by the entropy decoder 202 to derive quantized coefficient levels and prediction relationship information. Then, the quantized coefficient levels are processed by the inverse quantization 204 and the inverse transform 206 to obtain a reconstructed prediction residual. The block predictor mechanism implemented within the intra / inter mode selector 212 is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. By using the adder 214 to add the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block predictor mechanism, a set of non-filtered reconstructed pixels is obtained.

[0034]

[0041] The reconstructed block can further undergo a loop filter 209 and is then stored in a picture buffer 213 that functions as a reference picture memory unit. Then, the reconstructed video in the picture buffer 213 is sent out to drive a display device, and this can be used to predict future video blocks. In a situation where the loop filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0035]

[0042] In video coding standards such as HEVC, blocks can be divided based on a quadtree. In newer video coding standards such as the current VVC, more division methods are used, and one coding tree unit (CTU) can be divided into multiple CUs to adapt to local features that vary based on a quadtree, a binary tree, or a ternary tree. The separation of CUs, prediction units (PUs), and transform units (TUs) does not exist in most of the current VVC coding modes, and each CU is always used as the basic unit for both prediction and transform without further division. However, in some specific coding modes such as the intra subdivision coding mode, each CU can further contain multiple TUs. In multiple types of tree structures, first, one CTU is divided by a quadtree structure. Then, the leaf nodes of each quadtree can be further divided by binary and ternary tree structures.

[0036]

[0043] FIG. 3 shows five division types used in the current VVC, namely, four - way division 301, horizontal two - way division 302, vertical two - way division 303, horizontal three - way division 304, and vertical three - way division 305. In a situation where multiple types of tree structures are utilized, first, one CTU is divided by a quadtree structure. Then, the leaf nodes of each quadtree can be further divided by binary and ternary tree structures.

[0037]

[0044] Using one or more of the exemplary block divisions 301, 302, 303, 304, or 305 shown in FIG. 3 and using the configuration shown in FIG. 1, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra prediction") predicts the current video block using samples (referred to as reference samples) of already - encoded adjacent blocks within the same video picture / slice. Spatial prediction reduces the spatial redundancy inherent in the video signal.

[0038]

[0045] In more recent video coding standards such as the current VVC, new inter-mode coding tools have been introduced. Two examples of the new inter-mode coding tools are bi-directional optical flow (BDOF) and decoder-side motion vector refinement (DMVR).

[0039]

[0046] Conventional bi-directional prediction in video coding is a simple combination of two temporal prediction blocks obtained from already reconstructed reference pictures. However, due to the limitations of block-based motion compensation, there may remain slight motion that can be observed between samples of the two prediction blocks, thus reducing the efficiency of motion compensation prediction. To solve this problem, in the current VVC design, BDOF is applied to reduce the impact of such motion on all samples within one block.

[0040]

[0047] FIG. 4 is a diagram of the BDOF process. BDOF is a per-sample motion correction that is performed in addition to block-based motion compensation prediction when bi-directional prediction is used. The motion correction for each 4×4 sub-block 401 is calculated by minimizing the difference between the reference picture list 0 (L0) prediction sample 402 and the reference picture list 1 (L1) prediction sample 403 after BDOF is applied within one 6×6 window around the sub-block. Based on the motion correction derived in such a way, the final bi-directional prediction sample of the CU is calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical flow model.

[0041]

[0048] DMVR is a bi-directional prediction technique for merge blocks having two initially signaled MVs that can be further corrected by using mutually consistent prediction.

[0042]

[0049] Figure 5 is a diagram of mutual consistency used in DMVR. Mutual consistency is used to derive the motion information of the current CU501 by finding the closest match between two blocks 503 and 504 along the trajectory 502 of the motion of the current CU501 within two different reference pictures 505 and 506. The cost function used in the matching process is the sum of absolute differences (SAD) sub-sampled row by row. After the matching process, the corrected MVs 507 and 508 are used for motion compensation in the prediction stage, boundary strength calculation of the non-block filter, temporal motion vector prediction for subsequent pictures, and cross-CTU spatial motion vector prediction for subsequent CUs. Under the assumption that the motion trajectories are continuous, the motion vectors MV0 507 and MV1 508 pointing to the two reference blocks 503 and 504 are assumed to be proportional to the temporal distances, i.e., TD0 509 and TD1 510, between the current picture 511 and the two reference pictures 505 and 506. As a special case, when the current picture 511 is temporally between the two reference pictures 505 and 506 and the temporal distances from the current picture 511 to the two reference pictures 505 and 506 are the same, the mutual consistency becomes a mirror-based bidirectional MV.

[0043]

[0050] On the one hand, in order to strike an appropriate balance between the increased coding efficiency that can be brought about by newer inter-mode coding tools such as BDOF and DMVR, and on the other hand, the increased complexity and latency associated with the newer inter-mode tools, the current VVC design applies constraints regarding when BDOF or DMVR can be enabled for the current block.

[0044]

[0051] In the current VVC design, BDOF is only possible when all of the following predefined BDOF conditions listed within the frame immediately following this paragraph are met.

[0045] 1. The current block uses bidirectional prediction, one MV points to a reference picture before the current picture in display order, and the other MV points to a reference picture after the current picture in display order. 2. Weighted prediction is not possible. 3. The height of the current block is not equal to 4. 4. The size of the current block is not equal to 4×8 (i.e., width 4 and height 8). 5. The current block is not encoded as a symmetric motion vector difference (MVD) mode, which is a special MVD encoding mode in VVC. 6. The current block is not encoded as an affine mode. 7. The current block is not encoded as a sub-block merge mode. 8. When averaging predictor samples from list 0 and list 1, the current block does not use different heights (e.g., bidirectional prediction (BWA) with weighted averaging by unequal weights).

[0046]

[0052] In the current VVC design, DMVR is only possible when all of the following predefined DMVR conditions listed in the frame immediately below this paragraph are met.

[0047] 1. The current block uses bidirectional prediction, one MV points to a reference picture before the current picture in display order, another MV points to a reference picture after the current picture in display order, and further, the distance between the current picture and the forward reference picture, and the distance between the current picture and the backward reference picture are the same. 2. The current block is encoded as a merge mode, and the selected merge candidate is a normal merge candidate (e.g., a normal spatial merge candidate that is not a sub-block, or a temporal merge candidate, etc.). 3. The height of the current block is 8 or more. 4. The area of the current block is 64 or more. 5. The current block is not encoded as an affine mode. 6. The current block is not encoded as a sub-block merge mode. 7. The current block is not encoded as a merge mode with motion vector difference (MMVD) mode.

[0048]

[0053] The above constraints existing in the current VVC design are very helpful in achieving the desired balance between one encoding efficiency and the other complexity and latency, but do not completely solve this problem.

[0049]

[0054] One remaining problem with the current VVC design is that although some constraints have already been applied to enable BDOF and DMVR, in some cases, when encoding a block, it is possible to enable both of the two decoder-side inter-prediction correction tools BDOF and DMVR. In the current VVC design, when both decoder-side inter-prediction correction tools are enabled, BDOF depends on the final motion compensation samples of DMVR, causing a latency problem for the hardware design.

[0050]

[0055] The second remaining problem with the current VVC design is that although some constraints have already been applied to enable DMVR, these constraints are still too lenient overall with respect to enabling DMVR. Because there is a scenario where a better balance should be taken between one encoding efficiency and the other complexity and latency by disabling DMVR and then reducing complexity and latency, but the current VVC design enables DMVR in these scenarios.

[0051]

[0056] The third remaining problem with the current VVC design is that the constraints already applied to enable BDOF are too strict overall with respect to enabling BDOF. Because there is a scenario where a better balance should be taken between one encoding efficiency and the other complexity and latency by enabling BDOF and then increasing the encoding gain, but the current VVC design does not enable BDOF in these scenarios.

[0052]

[0057] According to a first aspect of the present disclosure, when a current block is suitable for the application of both DMVR and BDOF based on a plurality of predefined conditions, the current block is classified into one of two predefined classes, namely the DMVR class and the BDOF class, using a predefined criterion based on the mode information of the current block. Next, the classification of the current block is used when applying either one but not both of DMVR and BDOF to the current block. This method can be combined with the current VVC in addition to the above predefined conditions, or can be implemented independently.

[0053]

[0058] FIG. 6 is a flowchart showing the operation of the first aspect of the present disclosure. When processing the current block (601), this operation of this aspect of the present disclosure applies a plurality of predefined conditions to the current block (602), and can determine whether the current block is suitable for the application of both DMVR and BDOF based on the plurality of predefined conditions (603). If it is determined that the current block is not suitable for the application of both DMVR and BDOF based on the plurality of predefined conditions, this operation of this aspect of the present disclosure can continue the process existing in the current VVC design (604). On the other hand, if it is determined that the current block is suitable for the application of both DMVR and BDOF based on the plurality of predefined conditions, this operation of this aspect of the present disclosure can use a predefined criterion based on the mode information of the current block to classify the current block into one of two predefined classes, namely the DMVR class and the BDOF class (605). Next, this operation of this aspect of the present disclosure can apply either one but not both of DMVR and BDOF to the current block using the classification result of the current block (606). Applying either one but not both of DMVR and BDOF to the current block using the classification result of the current block (606) includes continuing with the DMVR class (applying DMVR instead of BDOF) if the current block is classified into the DMVR class (607), and can include continuing with the BDOF class (applying BDOF instead of DMVR) if the current block is classified into the BDOF class (608).

[0054]

[0059] The plurality of predefined conditions based on which the current block becomes suitable for the application of both DMVR and BDOF can be the plurality of predefined BDOF conditions and predefined DMVR conditions listed in the above framework, but this is not necessarily required.

[0055]

[0060] The mode information used as a basis for a predefined criterion is not limited thereto, and includes prediction modes such as whether to use a merge mode, a merge mode index, a motion vector, a block shape, a block size, and predictor sample values.

[0056]

[0061] According to one or more embodiments of the present disclosure, when applying either DMVR or BDOF (but not both) to the current block, using the classification of the current block includes optionally signaling a flag to indicate the classification of the current block. That is, in some examples of this embodiment, one flag is signaled to specify whether BDOF or DMVR has been applied to the block, and in some other examples of this embodiment, such a flag is not signaled.

[0057]

[0062] According to one or more embodiments of the present disclosure, when applying either DMVR or BDOF (but not both) to the current block, using the classification of the current block further includes applying DMVR rather than BDOF to the current block by an existing merge candidate list or by a separately generated merge candidate list when the current block is classified into the DMVR class. That is, in some examples of this embodiment, when the current block is classified into the DMVR class and DMVR is applied to the current block, a separate merge candidate list is exclusively generated and used. To indicate this DMVR merge mode, syntax is signaled, and when the DMVR merge candidate list size is greater than 1, the merge index is also signaled. In some other examples of this embodiment, when the current block is classified into the DMVR class and DMVR is applied to the current block, such a separate merge candidate is not generated, and the application of DMVR to the current block uses the existing merge candidate list without further syntax or signaling.

[0058]

[0063] According to one or more embodiments of the present disclosure, when applying either DMVR or BDOF but not both to the current block, using the classification of the current block further includes applying BDOF rather than DMVR to the current block by an existing merge candidate list or by a separately generated merge candidate list when the current block is classified into the BDOF class. That is, in some examples of this embodiment, when the current block is classified into the BDOF class and BDOF is applied to the current block, a separate merge candidate list is exclusively generated and used. To indicate this BDOF merge mode, a syntax is signaled, and when the BDOF merge candidate list size is greater than 1, a merge index is also signaled. In some other examples of this embodiment, when the current block is classified into the BDOF class and BDOF is applied to the current block, such a separate merge candidate is not generated, and the application of BDOF to the current block uses the existing merge candidate list without further syntax or signaling.

[0059]

[0064] According to another embodiment of the first aspect of the present disclosure, classifying the current block into one of two predefined classes, namely the DMVR class and the BDOF class, using a predefined criterion based on the mode information of the current block includes classifying the current block into the DMVR class when the predefined criterion is satisfied, and classifying the current block into the BDOF class when the predefined criterion is not satisfied.

[0060]

[0065] According to another embodiment of the first aspect of the present disclosure, classifying the current block into one of two predefined classes, namely the DMVR class and the BDOF class, using a predefined criterion based on the mode information of the current block includes classifying the current block into the BDOF class when the predefined criterion is satisfied, and classifying the current block into the DMVR class when the predefined criterion is not satisfied.

[0061]

[0066] According to another embodiment of the first aspect of the present disclosure, the predefined criterion includes whether the normal mode is selected for the current block.

[0062]

[0067] According to another embodiment of the first aspect of the present disclosure, the predefined criterion includes whether the coding merge index of the current block has predefined mathematical properties.

[0063]

[0068] In another example, the predefined mathematical property includes the property of being equal to or greater than a predefined threshold number.

[0064]

[0069] According to another embodiment of the first aspect of the present disclosure, the predefined criterion includes whether the motion vector of the current block satisfies a predefined test.

[0065]

[0070] In one example, the predefined test includes whether the sum of the magnitudes of all motion vector components is greater than a predefined threshold number.

[0066]

[0071] According to another embodiment of the first aspect of the present disclosure, the predefined criterion includes whether the current block has a predefined shape.

[0067]

[0072] In one example, the predefined shape is a square shape.

[0068]

[0073] According to another embodiment of the first aspect of the present disclosure, the predefined criterion includes whether the block size of the current block has predefined mathematical properties.

[0069]

[0074] In one example, the predefined mathematical property includes the property of being equal to or greater than a predefined threshold number.

[0070]

[0075] According to another embodiment of the first aspect of the present disclosure, the predefined criterion includes whether the sum of absolute differences or the sum of squared differences (SAD or SSD) between the list 0 predictor sample and the list 1 predictor sample of the current block has a predefined mathematical property.

[0071]

[0076] In one example, the predefined mathematical property includes the property of being greater than a predefined number of thresholds.

[0072]

[0077] According to the second aspect of the present disclosure, when the current block is suitable for DMVR coding based on a plurality of predefined conditions, a first determination is made as to whether weighted prediction is possible for the current block, and a second determination is made as to whether different weights are used when averaging the list 0 predictor sample and the list 1 predictor sample for the current block. Based on the first and second determinations, it is possible to determine whether to disable the application of DMVR to the current block. This method can be combined with the current VVC in addition to the above predefined conditions, or can be implemented independently.

[0073]

[0078] FIG. 7 is a flowchart illustrating an exemplary method of a second aspect of the present disclosure. When processing the current block (701), this operation of this aspect of the present disclosure applies a plurality of predefined conditions to the current block (702), and can determine whether the current block is suitable for DMVR encoding based on the plurality of predefined conditions (703). If it is determined that the current block is not suitable for DMVR encoding based on the plurality of predefined conditions, this operation of this aspect of the present disclosure can continue the process existing in the current VVC design (704). On the other hand, if it is determined that the current block is suitable for DMVR encoding based on the plurality of predefined conditions, this operation of this aspect of the present disclosure can determine whether weighted prediction is possible for the current block (705), and can also determine whether different weights are used when averaging the list 0 predictor sample and the list 1 predictor sample for the current block (706), and then based on the two determinations, can determine whether to disable the application of DMVR to the current block (707).

[0074]

[0079] The plurality of predefined conditions based on which the current block becomes suitable for DMVR encoding can be the plurality of predefined DMVR conditions listed in the above framework, but this is not necessarily required.

[0075]

[0080] According to an embodiment of the second aspect of the present disclosure, determining whether to disable the application of DMVR to the current block based on the two determinations includes disabling the application of DMVR to the current block when it is determined that weighted prediction is possible for the current block.

[0076]

[0081] According to another embodiment of the second aspect of the present disclosure, determining whether to disable the application of DMVR to the current block based on two determinations includes disabling the application of DMVR to the current block when it is determined that different weights are used when averaging the list 0 predictor sample and the list 1 predictor sample for the current block.

[0077]

[0082] According to another embodiment of the second aspect of the present disclosure, determining whether to disable the application of DMVR to the current block based on two determinations includes disabling the application of DMVR to the current block when it is determined that weighted prediction is possible for the current block and at the same time it is determined that different weights are used when averaging the list 0 predictor sample and the list 1 predictor sample for the current block.

[0078]

[0083] According to the third aspect of the present disclosure, when the current block is encoded as a sub-block merge mode, the application of BDOF to the current block can be enabled. This method can be combined with the current VVC in addition to the above predefined conditions, or can be implemented independently.

[0079]

[0084] FIG. 8 is a flowchart showing the operation of the third aspect of the present disclosure. When processing the current block (801), this operation of this aspect of the present disclosure can determine whether the current block is encoded as a sub-block merge mode (802). If it is determined that the current block is not encoded as a sub-block merge mode, this operation of this aspect of the present disclosure can continue the process existing in the current VVC design (803). On the other hand, if it is determined that the current block is encoded as a sub-block merge mode, this operation of this aspect of the present disclosure can enable the application of BDOF to the current block (804).

[0080]

[0085] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a processing unit based on hardware. The computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transmission of a computer program from one location to another, for example, in accordance with a communication protocol. In this way, the computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the implementation examples described in this application. A computer program product can include a computer-readable medium.

[0081]

[0086] Furthermore, the above method can be implemented using an apparatus including one or more circuits, and such circuits include application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. This apparatus can use these circuits in combination with other hardware or software components to execute the above-described method. Each module, sub-module, unit, or sub-unit disclosed above can be implemented using at least partially one or more circuits.

[0082]

[0087] Other embodiments of the invention will be apparent to those of ordinary skill in the art from a consideration of the specification and practice of the invention as disclosed above. This application is intended to cover any variations, uses, or adaptations of the invention following, in general, the principles of the invention and including such departures from the present disclosure as come within known or customary practice in the art to which the invention pertains. The specification and examples are to be considered as illustrative only, and the true scope and spirit of the invention is intended to be indicated by the following claims.

[0083]

[0088] It is to be understood that the invention is not limited to the exact examples described and illustrated above, and various modifications and changes can be made without departing from the scope of the invention. The scope of the invention is intended to be limited only by the appended claims.

Claims

1. A method for video encoding, comprising the step of obtaining a plurality of blocks obtained by dividing a video picture, each of the plurality of blocks being encoded in an intra prediction mode or an inter prediction mode, for a current block among the plurality of blocks, determining whether weighted prediction is possible for the current block when the current block satisfies a plurality of predefined conditions, and determining whether different weights are used when averaging list 0 prediction samples and list 1 prediction samples for the current block; determining whether to disable decoder-side motion vector refinement (DMVR) for the current block based on the two determinations; and outputting prediction mode information of the current block by a video bitstream.

2. The step of determining whether to disable the application of DMVR to the current block based on the two determinations includes disabling the application of DMVR to the current block when it is determined that weighted prediction is possible for the current block. The method according to claim 1.

3. The step of determining whether to disable the application of DMVR to the current block based on the two determinations includes disabling the application of DMVR to the current block when it is determined that different weights are used when averaging the list 0 prediction samples and the list 1 prediction samples for the current block. The method according to claim 1.

4. The step of determining whether to invalidate the application of DMVR to the current block based on the two determinations includes invalidating the application of DMVR to the current block when it is determined that weighted prediction is possible for the current block and at the same time it is determined that different weights are used when averaging the list 0 prediction sample and the list 1 prediction sample for the current block. The method according to claim 1.

5. The plurality of predefined conditions include the current block is not encoded in affine mode, and the current block is not encoded in sub-block merge mode The method according to claim 1 including this.

6. A computing device for video encoding, comprising one or more processors, a non-transitory storage device coupled to the one or more processors, and a plurality of programs stored in the non-transitory storage device The plurality of programs, when executed by the one or more processors, cause the computing device to obtain a plurality of blocks obtained by dividing a video picture, each of the plurality of blocks being encoded in an intra prediction mode or an inter prediction mode, including the step of For the current block among the plurality of blocks, determining whether weighted prediction is possible for the current block when the current block satisfies a plurality of predefined conditions, and determining whether different weights are used when averaging the list 0 prediction sample and the list 1 prediction sample for the current block; ​Based on the two determinations, determining whether to disable the application of decoder-side motion vector refinement (DMVR) to the current block; A computing device that executes an operation including: outputting prediction mode information of the current block by a video bitstream.

7. The step of determining whether to disable the application of DMVR to the current block based on the two determinations includes disabling the application of DMVR to the current block when it is determined that weighted prediction is possible for the current block. The computing device according to claim 6.

8. The step of determining whether to disable the application of DMVR to the current block based on the two determinations includes disabling the application of DMVR to the current block when it is determined that different weights are used when averaging the list 0 prediction sample and the list 1 prediction sample for the current block. The computing device according to claim 6.

9. The step of determining whether to disable the application of DMVR to the current block based on the two determinations includes disabling the application of DMVR to the current block when it is determined that weighted prediction is possible for the current block and at the same time different weights are used when averaging the list 0 prediction sample and the list 1 prediction sample for the current block. The computing device according to claim 6.

10. The plurality of predefined conditions include The current block is not encoded as an affine mode, The current block is not encoded as a sub-block merge mode The computing device according to claim 6, including this.

11. A non-transitory computer-readable storage medium storing a plurality of instructions, when the plurality of instructions are executed by one or more processors of a computing device, cause the computing device to execute the method according to any one of claims 1 to 5 to generate a video bitstream, and store the generated video bitstream in the non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium.

12. A computer program including instructions, which, when executed on a computing device, cause the computing device to perform the method according to any one of claims 1 to 5 to generate a video bitstream. A computer program.

13. A method for storing a bitstream, performing the method according to any one of claims 1 to 5 to generate a bitstream; and storing the generated bitstream in a non-transitory computer-readable storage medium. A method.

Citation Information

Patent Citations

  • Moving image encoder, moving image encoding method, moving image encoding computer program, moving image decoder, moving image decoding method and moving image decoding computer program

    JP2018107580A

  • Side motion refinement in video encoding / decoding systems.

    JP2022515875A

  • JPP7339395B

  • Apparatus for video encoding, apparatus for video decoding, and non-transitory computer-readable storage medium

    US20180184116A1

  • Side motion refinement in video encoding / decoding systems

    US20210136401A1

Cited By

  • Method and device for selectively applying bidirectional optical flow and decoder-side motion vector refinement for video encoding

    JP2024099849A

  • Methods and devices for selectively applying bi-directional optical flow and decoder-side motion vector refinement for video coding

    US12407815B2

  • Methods and devices for selectively applying bi-directional optical flow and decoder-side motion vector refinement for video coding

    US12684112B2