Decoder-Side Motion Vector Refinement for Affine Motion Compensation

Decoder-side motion vector refinement for affine motion compensation addresses the limitation of exclusive DMVR application to non-affine blocks, enhancing coding efficiency and accuracy in affine-coded blocks.

JP2025521805APending Publication Date: 2025-07-10ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024577159
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-03
Filing Date
2023-07-05
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Current video coding technologies exclusively apply decoder-side motion vector refinement (DMVR) only to non-affine coded blocks, neglecting its potential application in affine motion compensation, leading to suboptimal motion vector accuracy in affine-coded blocks.

Method used

Implement decoder-side motion vector refinement for affine motion compensation by refining motion vectors of affine-coded blocks using a refined search method that includes bilateral matching and adaptive refinement techniques, allowing for improved motion vector accuracy.

Benefits of technology

Enhances the coding efficiency and accuracy of motion vectors in affine-coded blocks, thereby improving the overall video coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521805000001_ABST
    Figure 2025521805000001_ABST
Patent Text Reader

Abstract

To refine the motion vector accuracy and thereby improve the coding efficiency, a VVC standard encoder and a VVC standard decoder are provided that implement the application of DMVR in an affine merge mode coded block. The refined motion vector (MV) search is performed for the control point motion vector (CPMV) of an inter-coded coding block (CB), and the refined MV of the CB is output. The refined MV search includes deriving the MV of the sub-blocks of the CB based on the CPMV of the CB, performing sub-block MV refinement for the MV of the sub-blocks, and outputting the refined MV of the CB based on the refined MV of the sub-blocks. The refined MV search further includes deriving the affine model parameters based on multiple CPMVs of the CB, performing affine parameter offset search for the affine model parameters, and outputting the refined MV of the CB based on the optimal parameter offset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 358,257, filed Jul. 5, 2022, entitled "DECODER-SIDE MOTION VECTOR REFINEMENT FOR AFFINE MOTION COMPENSATION"; U.S. Provisional Patent Application No. 63 / 406,122, filed Sep. 13, 2022, entitled "DECODER-SIDE MOTION VECTOR REFINEMENT FOR AFFINE MOTION COMPENSATION"; U.S. Provisional Patent Application No. 63 / 433,748, filed Dec. 19, 2022, entitled "DECODER-SIDE MOTION VECTOR REFINEMENT FOR AFFINE MOTION COMPENSATION"; and U.S. Patent Application No. 18 / 346,766, filed Jul. 3, 2023, entitled "DECODER-SIDE MOTION VECTOR REFINEMENT FOR AFFINE MOTION COMPENSATION". All of the above applications are hereby expressly incorporated by reference in their entirety.

[0002] The present disclosure generally relates to video processing, and more particularly, to methods and systems for implementing decoder-side motion vector refinement for affine motion compensation.

Background Art

[0003] In 2020, the Joint Video Experts Team (JVET) of the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) published the final draft of the next-generation video codec specification, Versatile Video Coding (VVC). This specification further improves video coding performance compared to previous standards such as H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding). JVET continues to propose additional techniques beyond the scope of the VVC standard itself, which were collected under the name of the Enhanced Compression Model (ECM) in January 2021 and used as a new software base for developing tools beyond the VVC standard.

[0004] Inter-picture prediction is important for video encoding and is transmitted through motion vectors (MVs). Techniques for signaling MVs reduce the bitrate for transmission to the decoder but generate inaccurate MVs, leading to ambiguous predictions. To improve the results of these techniques, decoder-side refinement and / or compensation techniques can be used. Decoder-side Motion Vector Refinement (DMVR) includes multi-pass DMVR and adaptive DMVR and is a coding tool for refining MVs on the decoder side without additional signaling because MVs inherited from neighboring blocks may not perfectly match the current block. Also, affine motion compensation can be used to capture the affine motion between two different frames.

[0005] However, in the current design, these two coding tools are used exclusively with respect to each other. That is, DMVR is applied only in non-affine-coded blocks in order to refine the MV of translational motion. Also, when a block is coded in affine mode, DMVR is not used to refine the MV. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0006] Embodiments of the present disclosure are directed to decoder-side motion vector refinement for affine motion compensation.

[0007] In a first aspect, an embodiment of the present disclosure is a computer system, comprising one or more processors, and a computer-readable storage medium communicatively coupled to the one or more processors, the computer-readable storage medium storing computer-readable instructions executable by the one or more processors, the computer-readable instructions, when executed by the one or more processors, performing related operations, the computer-readable storage medium wherein the related operations include performing refined motion vector (MV) search for a control point motion vector (CPMV) of an inter-coded coding block (CB), and outputting the refined MV of the CB, and providing a computer system.

[0008] In a second aspect, an embodiment of the present disclosure is a method, comprising performing refined MV search for a CPMV of an inter-coded CB, and outputting the refined MV of the CB, and providing a method.

[0009] In a third aspect, one embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate the method according to the second aspect above.

[0010] In a fourth aspect, one embodiment of the present disclosure provides a computer program product including computer program instructions that enable a computer to execute the method according to the second aspect above.

[0011] In a fifth aspect, one embodiment of the present disclosure provides a computer program that enables a computer to execute the method according to the second aspect above.

[0012] The detailed description is set forth with reference to the accompanying drawings. In the drawings, the leftmost digit of a reference number identifies the figure in which the reference number first appears. The use of the same reference number in different figures indicates similar or equivalent items or features.

Brief Description of the Drawings

[0013]

Fig. 1A

Fig. 1B

Fig. 2

Fig. 3A

Fig. 3B

Fig. 3C

Fig. 4

Fig. 5A

Fig. 5B

Fig. 6

Fig. 7

Fig. 8A

Fig. 8B

Fig. 9

Fig. 10

Fig. 11

Fig. 12

Fig. 13

Fig. 14

Fig. 15

Fig. 16

Best Mode for Carrying Out the Invention

[0014] According to the VVC video coding standard (the "VVC standard") and the motion prediction described therein, computer-readable instructions stored on a computer-readable storage medium are executable by one or more processors of a computing system to configure the one or more processors to perform the operations of an encoder and a decoder as described by the VVC standard. Some of these encoder operations and decoder operations according to the VVC standard will be described in more detail later, but these subsequent descriptions should not be understood to cover all of the encoder operations and decoder operations according to the VVC standard. Later, "VVC standard encoder" and "VVC standard decoder" will be taken to represent the respective computer-readable instructions stored on a computer-readable storage medium that configure one or more processors to perform these respective operations (which may sometimes be referred to, for example, as the "reference implementation form" of the encoder or decoder).

[0015] Moreover, according to an exemplary embodiment of the present disclosure, a VVC standard encoder and a VVC standard decoder further include computer-readable instructions stored on a computer-readable storage medium that are executable by one or more processors of a computing system to configure the one or more processors to perform operations not specified by the VVC standard. The VVC standard encoder should not be understood as being limited to the operations of the reference implementation of the encoder, but should be understood as including further computer-readable instructions that configure one or more processors of a computing system to perform the further operations described herein. The VVC standard decoder should not be understood as being limited to the operations of the reference implementation of the decoder, but should be understood as including further computer-readable instructions that configure one or more processors of a computing system to perform the further operations described herein.

[0016] FIG. 1A and FIG. 1B respectively show exemplary block diagrams of an encoding process 100 and a decoding process 150 according to an exemplary embodiment of the present disclosure.

[0017] In the symbolization process 100, the VVC standard encoder configures one or more processors of the computing system to receive, as input, one or more input pictures from the image source 102. The input picture includes a number of pixels sampled by an image capture device such as a photosensor array and includes an uncompressed stream of multiple color channels (such as RGB color channels) that store color data at the original resolution of the picture, where each channel uses a number of bits to store the color data of each pixel of the picture. The VVC standard encoder configures one or more processors of the computing system to store this uncompressed color data in a compressed format, where the color data is stored at a resolution lower than the original resolution of the picture and is encoded as a luma (“Y”) channel and two chroma (“U” and “V”) channels at a resolution lower than the luma channel.

[0018] The VVC standard encoder encodes a picture (the picture being encoded, called the "current picture", which is distinguished from other pictures received from the image source 102) by configuring one or more processors of the computing system to divide the original picture into units and sub-units according to a partition structure. The VVC standard encoder configures one or more processors of the computing system to further divide the picture into macroblocks (MBs), each having dimensions of 16×16 pixels, and the MBs can be further divided into partitions. The VVC standard encoder configures one or more processors of the computing system to divide the picture into coding tree units (CTUs), and the luma and chroma components of the CTUs can be further divided into coding tree blocks (CTBs), and the CTBs are further divided into coding blocks (CBs). Alternatively, the VVC standard encoder configures one or more processors of the computing system to divide the picture into units of N×N pixels, and those units can then be further divided into sub-units. Each of these largest divided units of the picture is generally sometimes referred to as a "block" in the present disclosure.

[0019] The CB is coded using one block of luma samples and two corresponding blocks of chroma samples, where the picture is not monochrome and is coded using one coding tree.

[0020] The VVC standard encoder configures one or more processors of the computing system to divide the block into partitions having dimensions that are multiples of 4×4 pixels. For example, the partitions of the block can have dimensions of 8×4 pixels, 4×8 pixels, 8×8 pixels, 16×8 pixels, or 8×16 pixels.

[0021] Rather than the pixel color information of the original picture at full resolution, by encoding the color information of the picture blocks and the subdivision of the blocks, the VVC standard encoder configures one or more processors of the computing system to encode the color information of the picture at a resolution lower than that of the input picture and store the color information in fewer bits than the input picture.

[0022] Furthermore, the VVC standard encoder encodes the picture by configuring one or more processors of the computing system to perform motion prediction on the blocks of the current picture. Motion prediction coding refers to storing the image data of the blocks of the current picture (where the blocks of the original picture before coding are called "input blocks") using motion information and prediction units (PU) rather than pixel data, by means of intra prediction 104 or inter prediction 106.

[0023] Motion information refers to data that describes the motion of the block structure of a picture or a unit or its sub-units, such as motion vectors and references to blocks of the current picture or reference pictures. A PU may refer to a unit or multiple sub-units corresponding to a block structure among multiple block structures of a picture, such as an MB or CTU, where the block is divided based on the picture data and coded according to the VVC standard. The motion information corresponding to the PU may describe the motion prediction encoded by the VVC standard encoder described herein.

[0024] The VVC standard encoder configures one or more processors of the computing system to code the motion prediction information across each block of the picture in the coding order among the blocks, such as raster scan order where the first block to be decoded is the topmost and leftmost block of the picture. The block being coded is called the "current block" to be distinguished from other blocks of the same picture.

[0025] According to the intra prediction 104, one or more processors of the computing system are configured to encode a block based on motion information of one or more other blocks of the same picture and a reference to the PU. According to the intra prediction coding, one or more processors of the computing system execute the intra prediction 104 (also called spatial prediction) calculation by coding the motion information of the current block based on spatially neighboring samples from spatially neighboring blocks of the current block.

[0026] According to the inter prediction 106, one or more processors of the computing system are configured to encode a block based on motion information of one or more other pictures and a reference to the PU. One or more processors of the computing system are configured to store, in a reference picture buffer, one or more previously encoded and decoded pictures in inter prediction coding, and these stored pictures are called reference pictures.

[0027] One or more processors are configured to execute the inter prediction 106 (also called temporal prediction or motion compensation prediction) calculation by coding the motion information of the current block based on samples from one or more reference pictures. The inter prediction can be further calculated according to single prediction or dual prediction, that is, in single prediction, only one motion vector pointing to one reference picture is used to generate a prediction signal for the current block. In dual prediction, two motion vectors each pointing to a respective reference picture are used to generate a prediction signal for the current block.

[0028] The VVC standard encoder configures one or more processors of a computing system to code a coding block (CB) to include a reference index for identifying a prediction signal of a current block for reference by a VVC standard decoder. One or more processors of the computing system can code the CB to include an inter prediction indicator. The inter prediction indicator indicates a list 0 prediction regarding a first reference picture list called list 0, a list 1 prediction regarding a second reference picture list called list 1, or a dual prediction regarding both reference picture lists called list 0 and list 1, respectively.

[0029] When the inter prediction indicator indicates a list 0 prediction or a list 1 prediction, one or more processors of the computing system are each configured to code the CB to include a reference index pointing to a reference picture in a reference picture buffer referenced by list 0 or list 1. When the inter prediction indicator indicates a dual prediction, one or more processors of the computing system are configured to code the CB to include a first reference index pointing to a first reference picture in a reference picture buffer referenced by list 0 and a second reference index pointing to a second reference picture in a reference picture buffer referenced by list 1.

[0030] The VVC standard encoder configures one or more processors of a computing system to individually code each current block of a picture and output a prediction block for each. According to the VVC standard, a CTU can be the same size as 128×128 luma samples (and corresponding chroma samples depending on the chroma format). A CTU can be further divided into CUs according to a quadtree, a binary tree, or a ternary tree. One or more processors of the computing system are configured to finally record a set of coding parameters such as a coding mode (intra mode or inter mode), motion information (reference index, motion vector, etc.) for an inter-coded block, and quantized residual coefficients, in the syntax structure of the leaf nodes of the partition structure.

[0031] After the prediction block is output, the VVC standard encoder configures one or more processors of a computing system to send a set of coding parameters such as a coding mode (i.e., intra or inter prediction), the mode of intra prediction or the mode of inter prediction, and motion information, to an entropy encoder 124 (described later).

[0032] The VVC standard provides semantics for recording coding parameter sets for CUs. For example, regarding the coding parameter sets described above, the pred_mode_flag for a CU is set to 0 for an inter-coded block and 1 for an intra-coded block, the general_merge_flag for a CU is set to indicate whether the merge mode is used in the inter prediction of the CU, the inter_affine_flag and cu_affine_type_flag for a CU are set to indicate whether affine motion compensation is used in the inter prediction of the CU, the mvp_l0_flag and mvp_l1_flag are set to indicate the motion vector index in list 0 and list 1 respectively, and the ref_idx_l0 and ref_idx_l1 are set to indicate the reference picture index in list 0 and list 1 respectively. It should be understood that the VVC standard includes semantics for recording various other information, flags, and options that are outside the scope of this disclosure.

[0033] The VVC standard encoder further implements one or more mode decisions and encoder control settings 108, including rate control settings. One or more processors of a computing system are configured to perform mode decision by selecting an optimized prediction mode for the current block based on a rate distortion optimization method after intra or inter prediction.

[0034] The rate control setting configures one or more processors of a computing system to assign different quantization parameters (QP) to different pictures. The magnitude of the QP determines the scale at which picture information is quantized by one or more processors during encoding (as will be described later), and thus determines the degree to which the encoding process 100 discards picture information from the MBs of the sequence (by the information falling between the steps of the scale) during coding.

[0035] The VVC standard encoder further implements an adder 110. One or more processors of the computing system are configured to perform an addition operation by calculating the difference between the input block and the prediction block. Based on the optimized prediction mode, the prediction block is subtracted from the input block. The difference between the input block and the prediction block is called the prediction residual, or simply the "residual" for brevity.

[0036] Based on the prediction residual, the VVC standard encoder further implements a transform 112. One or more processors of the computing system are configured to perform a transform operation on the residual by matrix arithmetic operations to derive an array of coefficients (which may be referred to as "residual coefficients", "transform coefficients", etc.), thereby encoding the current block as a transform block (TB). The transform coefficients may refer to coefficients representing one of several spatial transforms, such as diagonal inversion, vertical inversion, or rotation, that can be applied to sub-blocks.

[0037] It should be understood that the coefficients can be stored as two components, an absolute value and a sign, as will be described in more detail later.

[0038] Sub-blocks of CB, such as PU and TB, can be arranged in any combination of sub-block dimensions as described above. The VVC standard encoder configures one or more processors of the computing system to further partition the CB into a residual quadtree (RQT), which is a hierarchical structure of TB. The RQT provides an order for motion prediction and residual coding that spans the sub-blocks at each level of the RQT and recursively descends through each level of the RQT.

[0039] The VVC standard encoder further implements quantization 114. One or more processors of the computing system are configured to perform a quantization operation on the residual coefficients by matrix arithmetic operations based on the quantization matrix and the QP assigned above. Residual coefficients within a certain interval are retained, and residual coefficients outside that interval step are discarded.

[0040] The VVC standard encoder further implements inverse quantization 116 and inverse transform 118. One or more processors of the computing system are configured to perform an inverse quantization operation and an inverse transform operation on the quantized residual coefficients by matrix arithmetic operations, which are the inverses of the quantization operation and the transform operation described above. The inverse quantization operation and the inverse transform operation result in a reconstructed residual.

[0041] The VVC standard encoder further implements an adder 120. One or more processors of the computing system are configured to perform an addition operation by adding the prediction block and the reconstructed residual and output the reconstructed block.

[0042] The VVC standard encoder further implements a loop filter 122. One or more processors of a computing system are configured to apply loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF) to a reconstructed block and output the filtered and reconstructed block.

[0043] The VVC standard encoder further configures one or more processors of a computing system to output the filtered and reconstructed block to a decoded picture buffer (DPB) 200. The DPB 200 stores reconstructed pictures that are used as reference pictures by one or more processors of the computing system when coding pictures other than the current picture, as described above with respect to inter prediction.

[0044] The VVC standard encoder further implements an entropy encoder 124. One or more processors of a computing system are configured to perform entropy coding, where symbols constituting the quantized residual coefficients are coded by mapping to a binary sequence (hereinafter referred to as "bin") according to a Context-Adaptive Binary Arithmetic Codec (CABAC), and this can be transmitted in an output bitstream at a compressed bitrate. The symbols of the quantized residual coefficients to be coded include the absolute values of the residual coefficients (these absolute values are hereinafter referred to as "residual coefficient levels").

[0045] The entropy coder configures one or more processors of a computing system to perform the following operations: coding the residual coefficient levels of a block; bypassing the coding of the residual coefficient signs and recording the residual coefficient signs together with the coded block; recording coding parameters sets such as coding mode, intra prediction mode or inter prediction mode, and motion information (picture parameter set (PPS) found in the picture header, and sequence parameter set (SPS) found in the sequence of multiple pictures, etc.) that are coded within the syntax structure of the coded block; and outputting the coded block.

[0046] The VVC standard encoder configures one or more processors of a computing system to output a coded picture composed of coded blocks from the entropy coder 124. The coded picture is output to a transmission buffer, where it is finally packed into a bitstream for the output from the VVC standard encoder.

[0047] In the decoding process 150, the VVC standard decoder configures one or more processors of a computing system to receive, as input, one or more coded pictures from the bitstream.

[0048] The VVC standard decoder implements an entropy decoder 152. One or more processors of the computing system are configured to perform entropy decoding, wherein, according to CABAC, by reversing the mapping from symbols to bins, the bins are decoded, whereby the entropy-coded quantized residual coefficients are restored. The entropy decoder 152 outputs the quantized residual coefficients, outputs the coding-bypassed residual coefficient symbols, and also outputs the syntax structure such as PPS and SPS.

[0049] The VVC standard decoder further implements an inverse quantization 154 and an inverse transform 156. One or more processors of the computing system are configured to perform an inverse quantization operation and an inverse transform operation on the decoded quantized residual coefficients by matrix arithmetic operations which are the inverses of the quantization operation and the transform operation described above. The inverse quantization operation and the inverse transform operation result in a reconstructed residual.

[0050] Furthermore, based on the coding parameter set recorded in the syntax structure such as PPS and SPS by the entropy coder 124 (alternatively, received by out-of-band transmission or coded in the decoder) and the coding mode included in the coding parameter set, the VVC standard decoder determines whether to apply intra prediction 158 (i.e., spatial prediction) or motion compensation prediction 160 (i.e., temporal prediction) to the reconstructed residual.

[0051] When the coding parameter set specifies intra prediction, the VVC standard decoder configures one or more processors of the computing system to perform intra prediction 158 using the prediction information specified in the coding parameter set. Intra prediction 158 thereby generates a prediction signal.

[0052] When the coding parameter set specifies inter prediction, the VVC standard decoder configures one or more processors of the computing system to perform motion compensation prediction 160 using reference pictures from the DPB200. The motion compensation prediction 160 thereby generates a prediction signal.

[0053] The VVC standard decoder further implements an adder 162. The adder 162 configures one or more processors of the computing system to perform an addition operation on the reconstructed residual and the prediction signal, thereby outputting a reconstructed block.

[0054] The VVC standard decoder further implements a loop filter 164. One or more processors of the computing system are configured to apply loop filters such as a deblocking filter, an SAO filter, and an ALF to the reconstructed block and output a filtered reconstructed block.

[0055] The VVC standard decoder further configures one or more processors of the computing system to output the filtered reconstructed block to the DPB200. As described above, the DPB200 stores reconstructed pictures that are used as reference pictures by one or more processors of the computing system when coding pictures other than the current picture, as described above for motion compensation prediction.

[0056] The VVC standard decoder further configures one or more processors of the computing system to output the reconstructed picture from the DPB to a user-viewable display of the computing system, such as a television display, a personal computing monitor, a smartphone display, or a tablet display.

[0057] Therefore, as shown by the encoding process 100 and the decoding process 150 described above, the VVC standard encoder and the VVC standard decoder each implement motion prediction coding according to the VVC specification. The VVC standard encoder and the VVC standard decoder each configure one or more processors of a computing system to generate a reconstructed picture based on the reconstructed picture before the DPB according to the motion compensation prediction described by the VVC standard, where the previous reconstructed picture serves as a reference picture in motion compensation prediction as described herein.

[0058] The VVC standard adopts decoder-side motion vector refinement (DMVR) based on bilateral matching (BM) in bi-prediction to improve the accuracy of the MV in the merge mode. In DMVR, the refined MV is searched near the initial MV, MV0, and MV1 in reference picture list 0 (L0) and reference picture list 1 (L1), where the refined MV is denoted as MV0' and MV1', respectively. The BM method calculates the distortion between two candidate blocks in reference pictures L0 and L1, respectively.

[0059] Figure 2 shows the motion prediction performed in the current picture 202 according to dual prediction, where the offset blocks of the reference pictures are used to calculate refined motion vectors, and the refined motion vectors are then used to generate a dual prediction signal. The current picture 202 includes a current block 202A. Two collocated reference pictures 204 and 206 are a reference picture from reference list 0 in the first temporal direction and a reference picture from reference list 1 in the second temporal direction, respectively, and are shown according to dual prediction. The motion information of the current block 202A refers to the collocated reference block 204A of the collocated reference picture 204 and the collocated reference block 206A of the collocated reference picture 206. The collocated reference picture 204 further includes an offset block 204B near the collocated reference block 204A, and the collocated reference picture 206 further includes an offset block 206B near the collocated reference block 206A.

[0060] As shown in FIG. 2, the sum of absolute differences ("SAD") between the reference block 204A and the offset block 204B, and the SAD between the reference block 206A and the offset block 206B are calculated. The MV candidate with the lowest SAD is set as the refined MV and used to generate a dual prediction signal.

[0061] Furthermore, according to the VVC standard, the application of DMVR is restricted and is only applicable to the CBs coded using the following modes and features, which are the CB-level merge mode using dual-prediction MVs, where the dual-prediction MVs point to respective reference pictures in different temporal directions for the current picture (i.e., one reference picture is from the past and the other reference picture is from the future), the distance from the two reference pictures to the current picture (i.e., the picture order count (POC) difference) is the same, both reference pictures are short-term reference pictures, the current CB has more than 64 luma samples, both the CB height and the CB width are 8 luma samples or more, the bidirectional prediction with coding unit weights (BCW) weight index indicates equal weights (in the context of the weighted-average dual-prediction formula where the weighted average of two prediction signals is calculated, "equal weights" should be understood as the weight parameters that cause the two prediction signals to be equally weighted in the formula), weighted bi-prediction (WP) is not enabled for the current block, and the combined inter-intra prediction (CIIP) mode is not used for the current block.

[0062] The refined MVs derived by DMVR are used to generate inter-prediction samples and are also used in the temporal motion vector prediction for future picture coding. The original MVs are used in deblocking and are also used in the spatial motion vector prediction for future CB coding.

[0063] Additional features of DMVR will be described later.

[0064] In DMVR, the refined MV search starts from the search center, encompasses the search range of the refined MVs that directly surround the initial MV, the span of the search range defines the search window, and the range of the searched refined MVs is offset according to the MV difference mirroring rule. In other words, any point searched by DMVR, indicated by the candidate MV pair (MV0, MV1), follows Equations 1 and 2 below respectively. MV0' = MV0 + MV offset MV1' = MV1 - MV offset where MV offset represents the MV refinement offset between the initial MV and the refined MV in one of the reference pictures. The refined MV search range (hereinafter also referred to as the "search step") is two integer distance luma samples from the initial MV. The refined MV search includes two stages, namely, integer sample offset search and fractional sample refinement.

[0065] For the purpose of understanding the exemplary embodiments of the present disclosure, all subsequent references to one or more "points" being searched should be understood to refer to the individual luma samples of blocks or sub-blocks separated by integer distances.

[0066] A full search of 25 points is applied for the integer sample offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR ends. Otherwise, the remaining 24 search points are searched in raster scan order, and the SAD of each search point is calculated. The search point with the minimum SAD is selected as the integer-distance refined MV, and it is output by the integer sample offset search. To reduce the penalty for the uncertainty of DMVR refinement, the original MV can be preferred during the DMVR process. The SAD between the reference blocks referred to by the initial MV candidates is reduced by 1 / 4 of the SAD value.

[0067] After integer sample offset search, fractional sample refinement may continue. To reduce computational complexity, fractional sample refinement is performed by solving a parametric error surface equation rather than further searching by SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample offset search. When the integer sample offset search ends in a state where the center has the minimum SAD in either the first iterative search or the second iterative search, fractional sample refinement is further applied. Otherwise, the integer distance refined MV may be output as the refined MV.

[0068] In parametric error surface-based sub-pixel offset estimation, the center position cost and the costs at four neighboring positions from the center are used to fit a 2-D parabolic error surface described by Equation 3 below. E(x,y)=A(x - x min ) 2 +B(y - y min ) 2 +C where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By using the cost values of five search points and solving the above equation, (x min ,y min ) is calculated according to Equations 4 and 5 below, respectively. x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) y min =(E(0,-1)-E(0,1)) / (2(E(0,-1)+E(0,1)-2E(0,0)))

[0069] x min and y minThe value is constrained to be between -8 and 8 by default because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pel offset at 1 / 16 pel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer-distance refined MV to obtain a subpixel-accurate refined delta MV. The subpixel-accurate refined delta MV can be output as the refined MV instead of the integer-distance refined MV.

[0070] In VVC, the resolution of the MV is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, since the refined MV search points are the points directly surrounding the initial fractional pel MV with integer sample offsets, the samples at those fractional positions need to be interpolated for DMVR refined MV search. To reduce the computational complexity, a bilinear interpolation filter is used to generate fractional samples for DMVR refined MV search. Moreover, by using the bilinear filter with a 2-sample search range, DMVR does not access more reference samples compared to the standard motion compensation process. After the refined MV is output by DMVR refined MV search, the standard 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the standard motion compensation process, samples required for the interpolation process based on the refined MV, rather than the original MV, are padded from those available samples.

[0071] When the width and / or height of the CB is greater than 16 luma samples, the CB will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size for DMVR refined MV search is limited to 16×16.

[0072] In the ECM, to further improve the coding efficiency, multi-pass decoder-side motion vector refinement is applied. In the first pass, BM is applied to the coding block. In the second pass, BM is applied to each 16×16 sub-block within the coding block. In the third pass, the MV in each 8×8 sub-block is refined by applying bi-directional optical flow ("BDOF"). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction.

[0073] In the first pass, the refined MVs are derived by applying BM to the coding block. Similar to DMVR, in bi-prediction, the refined MVs are searched near the two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived near the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.

[0074] The BM-based refinement performs a local search to derive the integer sample accuracy intDeltaMV. The local search applies a 3×3 square search pattern and loops through the search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size and the maximum values of sHor and sVer are 8 or other values. For example, in Figure 3A, point 0 is the position referenced by the initial MV and is set as the first search center. Therefore, points 1 to 8 around the initial point are searched first and the cost at each position is calculated.

[0075] In the first search iteration, it is discovered that point 7 has the minimum cost, point 7 is set as the second search center, and points 9, 10, and 11 are searched. In the next search iteration, since it is discovered that the cost of point 10 is smaller than the costs of points 7, 9, and 11, the third search center is set to point 10, and points 12, 13, and 14 are searched. In the next search iteration, since it is discovered that point 12 has the minimum cost among points 6 to 14, point 12 is set as the fourth search center. In the next search iteration, it is discovered that the costs of points 10, 11, 13, and 15 to 19 around point 12 are all greater than the cost of point 12. Then, point 12 is the optimal point, the refined MV search ends, and the refined MV corresponding to the optimal point is output. Therefore, FIG. 3A shows an exemplary diagram of a search pattern (e.g., 3×3 square search pattern) used in the first pass of multi-pass decoder-side motion vector refinement.

[0076] The bilateral matching cost can be calculated as bilCost = mvDistanceCost + sadCost, where sadCost is the SAD between the L0 predictor (i.e., the reference block from the reference picture L0) and the L1 predictor (i.e., the reference block from the reference picture L1) at the search point, and mvDistanceCost is based on intDeltaMV (i.e., the distance between the search point and the initial point). When the block size cbW (CB width, in pixels) * cbH (CB height, in pixels) is greater than 64, the mean-removed SAD (MRSAD) cost function is applied to remove the influence of the discrete cosine (DC) of the distortion between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search ends. Otherwise, the current minimum-cost search point is set as the new center point of the 3×3 search pattern, and the search for the minimum cost continues until the end of the search range is reached.

[0077] The existing fractional sample refinement is further applied to derive the fractional MV refinement fracDeltaMV, and the final deltaMV is derived as intDeltaMV + fracDeltaMV. Then, the refined MVs after the first pass are derived according to Equations 6 and 7 below, respectively. MV0_pass1 = MV0 + deltaMV MV1_pass1 = MV1 - deltaMV

[0078] In the second pass, the refined MVs are derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the refined MVs are searched near the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between two reference sub-blocks in L0 and L1.

[0079] For each sub-block, BM-based refinement performs an exhaustive search to derive the integer sample accuracy intDeltaMV(sbIdx2). The exhaustive search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8 or other values.

[0080] The bilateral matching cost can be calculated by applying a cost factor to the sum of absolute transformed differences ("SATD") cost between two reference sub-blocks, such as bilCost = satdCost * costFactor. The search area (2 * sHor + 1) * (2 * sVer + 1) is divided into up to five diamond-shaped search regions, as shown in FIG. 4. FIG. 4 shows a diagram of the bilateral matching cost (each matching cost corresponding to a diamond-shaped search region of a different shade) used in the second pass of multi-pass decoder side motion vector refinement. Each search region is assigned a costFactor, which is determined by the distance intDeltaMV(sbIdx2) between each search point and the starting MV, and each diamond-shaped region is processed in order starting from the center of the search area. In each region, the search points are processed in raster scan order starting from the upper left corner of the region and moving towards the lower right corner. When the minimum bilCost within the current search region is less than a threshold equal to sbW (sub-block width) * sbH (sub-block height), the integer pel full search ends; otherwise, the integer pel full search continues to the next search region until all search points have been examined. Additionally, if the difference between the previous minimum cost and the current minimum cost in an iteration is less than a threshold equal to the area of the block, the search ends.

[0081] Furthermore, the bilateral matching cost described above can also be calculated based on MRSAD instead of SAD, and can also be calculated based on the mean-removed sum of absolute transformed differences ("MRSATD") instead of SATD.

[0082] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MVs in the second pass are derived according to Equations 8 and 9 below, respectively. MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2) MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2)

[0083] In the third pass, the refined MVs are derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, the BDOF refinement starts from the refined MVs of the parent sub-block in the second pass and is applied to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample accuracy and clipped between -32 and 32.

[0084] The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) in the third pass are derived according to the following equations 10 and 11, respectively. MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) - bioMv

[0085] In ECM, adaptive decoder-side motion vector refinement is an extension of multi-pass DMVR that includes two new merge modes to refine the MVs only in one of the temporal directions of the bi-prediction for merge candidates that satisfy the DMVR condition, either in reference picture L0 or reference picture L1. The multi-pass DMVR process is applied for the selected merge candidates to refine the motion vectors, provided that either MVD0 or MVD1 is set to 0 in the first-pass (i.e., PU level) DMVR. Thus, a new merge candidate list is constructed for adaptive decoder-side motion vector refinement. The new merge mode for the new merge candidate list is called BM merge in ECM.

[0086] Merge candidates for the BM merge mode are derived from spatially neighboring coded blocks, a temporal based motion vector predictor (TMVP), non-adjacent blocks, a history-based motion vector predictor (HMVP), and pairwise candidates, similar to the normal merge mode. The difference is that only those merge candidates that satisfy the DMVR conditions are added to the merge candidate list. The same merge candidate list is used by the two new merge modes. The merge index is coded as in the case of the normal merge mode, except that if the list of BM candidates contains inherited BCW weights and the distortion calculation is done using MRSAD or MRSATD is used when the weights are not equal and biprediction is weighted with BCW weights, and the DMVR process remains unchanged.

[0087] In HEVC, only the translational motion model is applied for motion compensation prediction (MCP). However, in the real world, many types of motion occur, such as zoom in / out, rotation, viewpoint motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As shown in FIGS. 5A and 5B, the affine motion field of a block is described by the motion information of two control point motion vectors (4 parameters) (FIG. 5A) or three control point motion vectors (6 parameters) (FIG. 5B).

[0088] In affine motion compensation, for the 4-parameter affine motion model, the motion vector at the sample location (x, y) in the block is derived according to Equation 12 below.

[0089]

Equation

[0090] In the case of a 6-parameter affine motion model, the motion vector at the sample location (x, y) in the block is derived according to Equation 13 below.

[0091]

Equation

[0092] However, (mv 0x , mv 0y ) is the motion vector of the control point at the upper left corner, (mv 1x , mv 1y ) is the motion vector of the control point at the upper right corner, and (mv 2x , mv 2y ) is the motion vector of the control point at the lower left corner.

[0093] To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector for each 4×4 luma sub-block, as shown in Figure 6, the motion vector of the central sample of each sub-block is calculated according to the above equation and rounded to 1 / 16 fractional precision. Figure 6 shows a diagram of the affine motion vectors of the luma sub-blocks calculated for each sub-block central sample. According to ECM, the sub-block size is determined adaptively. If the motion vector difference between two neighboring luma sub-blocks is smaller than the threshold, the neighboring luma sub-blocks will be merged into a larger sub-block. If the motion vector difference between two neighboring larger sub-blocks is still smaller than the threshold, the larger sub-block will continue to be merged until the motion vector difference between the two neighboring sub-blocks is larger than the threshold or the sub-block becomes equal to the entire block.

[0094] After the motion vectors of the sub - blocks are derived, a motion - compensated interpolation filter is applied to generate predictors for each sub - block using the derived motion vectors. The sub - block size of the chroma component depends on the size of the luma sub - block. The MV of the chroma sub - block is calculated as the average of the MVs of the top - left luma sub - block and the bottom - right luma sub - block in the collocated luma region.

[0095] Similar to what is done for translational motion inter - prediction, there are two affine motion inter - prediction modes, namely, the affine merge mode and the advanced motion vector prediction (AMVP) mode.

[0096] The affine merge mode (AF_MERGE) can be applied to CBs with both width and height of 8 or more. In this mode, the control - point motion vector (CPMV) of the current CB is generated based on the motion information of spatially neighboring CBs. There can be up to 15 affine candidates, and an index is signaled to indicate which one will be used for the current CB. The following eight types of candidates are used to form the affine merge candidate list, namely, candidates inherited from adjacent neighbors, candidates inherited from non - adjacent neighbors, candidates constructed from adjacent neighbors, a second type of constructed affine candidate from non - adjacent neighbors, a first type of constructed affine candidate from non - adjacent neighbors, regression - based affine merge candidates, pairwise affine, and zero MV.

[0097] The inherited affine candidates are derived from the affine motion models of adjacent or non-adjacent blocks. When an adjacent affine CB or a non-adjacent affine CB is identified, its control point motion vector is used to derive a control point motion vector prediction (「CPMVP」) candidate in the affine merge candidate list of the current CB. As shown in FIG. 7 (which shows a diagram of control point motion vector inheritance from which candidates are used to form the affine merge candidate list), when the neighbor lower left block A is coded in the affine mode, the motion vectors v2, v3, and v4 of the upper left corner, upper right corner, and lower left corner of the CB containing block A are obtained. For block A coded using the 4-parameter affine model, two CPMVs of the current CB are calculated according to v2 and v3. For block A coded using the 6-parameter affine model, three CPMVs of the current CB are calculated according to v2, v3, and v4.

[0098] For candidates inherited from non-adjacent neighbors, non-adjacent spatial neighbors are checked based on their distances to the current block, i.e., from near to far. At a specific distance, only the first available neighbors (coded using the affine mode) from each side (e.g., the left side and the upper side) of the current block are included for inherited candidate derivation. As shown by the dashed arrows in FIG. 8A, the check order of neighbors on the left side and the upper side is from bottom to top and from right to left, respectively. FIGS. 8A and 8B show, respectively, a diagram of candidates inherited from non-adjacent neighbors for the affine merge candidate list, and a diagram of the first type of constructed candidates for the affine merge candidate list.

[0099] The constructed affine candidates from adjacent neighbors are candidates constructed by combining the neighbor translation motion information of each control point. The motion information for the control points is derived from the specified spatial neighbors and temporal neighbors, as shown in FIG. 9 (FIG. 9 shows a diagram of the locations of the constructed affine candidates from adjacent neighbors for the affine merge candidate list). CPMV k (k = 1, 2, 3, 4) represents the k-th control point. In the case of CPMV1, the B2->B3->A2 block is checked and the MV of the first available block is used. In the case of CPMV2, the B1->B0 block is checked, and in the case of CPMV3, the A1->A0 block is checked. TMVP is used as CPMV4 if available.

[0100] After the MVs of the four control points are obtained, the affine merge candidates are constructed based on that motion information. The following combinations of control point MVs are used to construct in order. {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3} Combinations of three CPMVs construct 6-parameter affine merge candidates, and combinations of two CPMVs construct 4-parameter affine merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the related combinations of control point MVs are discarded.

[0101] For the case of the first type of constructed candidates from non - adjacent neighbors, as shown in Figure 8B, the positions of one left and one upper non - adjacent spatial neighbor are first determined independently, and then the location of the upper - left neighbor can be determined accordingly, which can enclose a rectangular virtual block together with the left and upper non - adjacent neighbors. Then, as shown in Figure 10, the motion information of three non - adjacent neighbors is used to form the CPMVs at the upper - left (A), upper - right (B), and lower - left (C) of the virtual block, and they are finally projected onto the current CB to generate the corresponding constructed candidates. Figure 10 shows a diagram of the locations of the constructed affine candidates from non - adjacent neighbors for the affine merge candidate list.

[0102] For the case of the second type of constructed candidates, the affine model parameters are inherited from non - adjacent spatial neighbors. Specifically, the second - type affine - constructed candidates are generated from the combination of 1) the MVs of the adjacent 4×4 blocks of the neighbors and 2) the affine model parameters inherited from non - adjacent spatial neighbors defined in Figure 8A.

[0103] For the regression-based affine merge candidates, the sub-block motion field from the previously coded affine CB and the motion information from the neighboring sub-blocks of the current CB are used as inputs to the regression process to derive the affine candidates. The previously coded affine CB can be identified by scanning through non-adjacent positions and the affine HMVP table. The neighboring sub-block information of the current CB is fetched from the 4×4 sub-blocks represented by the gray zone as shown in FIG. 11. For each sub-block, when a reference list is given, the corresponding motion vector and center coordinates of the sub-block can be used. For each affine CB, up to two regression-based affine candidates can be derived, those with and without neighboring sub-block information. All of the linear regression-generated candidates are pruned and collected into one candidate subgroup, and when ARMC is enabled, the TM cost-based ARMC process is applied. Then, when N affine CBs are found, up to N linear regression-generated candidates are added to the affine merge candidate list.

[0104] After inserting all of the above merge candidates into the merge candidate list, if the list is still not full, zero MVs are inserted at the end of the list.

[0105] Sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity compared to pixel-based motion compensation at the cost of prediction accuracy penalty. To achieve a finer granularity of motion compensation, prediction refinement with optical flow (「PROF」) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luma prediction samples are refined by adding the differences derived by the optical flow formula. PROF is described as the following four steps.

[0106] First, sub-block-based affine motion compensation is performed to generate a sub-block prediction I(i,j).

[0107] Second, the spatial gradients g x (i,j) and g y (i,j) are calculated at each sample location using a 3-tap filter [-1,0,1] according to the following equations 14 and 15. The gradient calculation is the same as the gradient calculation in BDOF. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j)>>shift1) g y (i,j)=(I(i,j+1)>>shift1)-(I(i,j-1)>>shift1) shift1 is used to control the gradient accuracy. The sub-block (i.e., 4×4) prediction is extended by only one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extended boundaries are copied from the nearest integer pixel positions in the reference picture.

[0108] Third, the luma prediction refinement is calculated according to the following optical flow equation 16. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) However, Δv(i,j) is the difference between the sample MV calculated for the sample location (i,j) indicated by v(i,j) as shown in FIG. 12 and the sub-block MV of the sub-block to which the sample (i,j) belongs. Δv(i,j) is quantized in units of 1 / 32 luma sample accuracy. FIG. 12 shows the sub-block motion vector v at the location SBAlso shown is a diagram of the difference Δv(i,j) between the motion vector v(i,j) calculated at that location, and that difference is used as part of the prediction refinement by optical flow for sub-block-based affine motion compensation prediction.

[0109] Since the affine model parameters and the sample location relative to the sub-block center do not change for each sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CB. Let dx(i,j) and dy(i,j) be the horizontal offset and vertical offset from the sample location (i,j) to the center (x SB ,y SB ) of the sub-block, derived according to Equation 17 below. Then, Δv(x,y) can be derived according to Equation 18 below.

[0110]

Equation

[0111] To maintain accuracy, the center (x SB ,y SB ) of the sub-block is calculated as ((WSB - 1) / 2, (HSB - 1) / 2), where WSB and HSB are the width and height of the sub-block, respectively.

[0112] For the 4-parameter affine model, the parameters C, D, E, and F are derived according to Equation 19 below.

[0113]

Equation

[0114] For the 6-parameter affine model, the parameters C, D, E, and F are derived according to Equation 20 below.

[0115]

Equation

[0116] However, (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the top - left, top - right, and bottom - left control point motion vectors, and w and h are the width and height of the CB.

[0117] Fourthly, the loop filter prediction refinement ΔI(i, j) is added to the sub - block prediction I(i, j). The final prediction I' is generated according to Equation 21 below. I'(i, j)=I(i, j)+ΔI(i, j) PROF is not applied in two cases for affine - coded CBs, i.e., 1) when all control point MVs are the same, which indicates that the CB has only translational motion, 2) when the sub - block - based affine motion compensation is degraded to CB - based motion compensation to avoid large memory access bandwidth requirements, so the affine model parameters are larger than the specified limits.

[0118] According to the VVC standard and ECM proposal described above, DMVR is applied only to non - affine - coded blocks to refine the motion vectors of translational motion. For blocks coded using the affine mode, DMVR is not used to refine the motion vectors. However, similar to the motion vectors of translation - compensated - coded blocks, the motion vectors of affine - merge - mode - coded blocks are also inherited from previously coded blocks and may not exactly match the current block. Therefore, in the present disclosure, the VVC standard encoder and decoder configure one or more processors of the computing system to apply DMVR in affine - merge - mode - coded blocks to refine the motion vector accuracy and thereby improve the coding efficiency.

[0119] The "decoder side" does not mean that this method is exclusively implemented by the decoder. Rather, it should be understood that the steps of this method can be implemented equally or equivalently by the encoder and the decoder, as will be described later.

[0120] As described herein, in the affine model used to code a block in the affine merge mode, the motion vector at sample location (x, y) can be derived according to Equation 22 below.

[0121]

Equation

[0122] Here, (mv x , mv y ) is the derived motion vector at sample location (x, y), and (mv 0x , mv 0y ) represents the MV of the affine model, which is the motion vector at sample location (0, 0). a, b, c, and d are the parameters of the affine model. Affine motion can include, but is not limited to, translation, rotation, and zooming. The MV of the affine model represents the translational motion of the affine model. The affine model parameters a, b, c, and d represent the non-translational motion of the affine model, including rotation and zooming, as well as other non-translational motions. The affine model parameters can be derived based on the motion vectors at two sample locations in the case of a 4-parameter affine model and at three non-collinear sample locations in the case of a 6-parameter affine model. As a generalization of Equation 22 above, the MV of the affine model can be the motion vector at any sample location, not necessarily at location (0, 0). The motion vector at sample location (w, h) is ((mv wx , mv wy) taken as the MV of the affine model shown as), the motion vector at the sample location (x, y) can be formulated according to Equation 23 below.

[0123]

Number

[0124] For the 4-parameter affine model, b is equal to c and d is equal to a. Therefore, the 4-parameter affine model can be formulated according to Equation 24 below.

[0125]

Number

[0126] Theoretically, all of the parameters of the affine model, including a, b, c, d, mv wx , and mv wy can be refined at once in DMVR. However, in order to limit the computational complexity of the refinement, according to an exemplary embodiment of the present disclosure, while the affine model MV (mv wx , mv wy ) is refined, the affine model parameters a, b, c, and d are fixed, and while the affine model parameters a, b, c, and d are refined, the affine model MV is fixed. Later, the refined MV search where the affine model parameters are fixed is described, and the affine parameter offset search where the affine model MV is fixed is also described.

[0127] Similar to the standard DMVR implementation, the VVC standard encoder and decoder configure one or more processors of the computing system to apply DMVR-refined MV search for the initial MV as described above, except as shown below. In contrast to the standard DMVR implementation, during the refined MV search, when calculating the bilateral matching cost (regardless of whether it is based on SAD or SATD, or MRSAD or MRSATD) between two predictors of reference picture L0 and reference picture L1 (i.e., the reference block from reference picture L0 and the reference block from reference picture L1), motion compensation is performed at the sub-block level. Therefore, the refined MV search is performed for the initial MV over a number of iterations of integer sample offset search, followed by fractional sample refinement, and the refined MV search outputs the refined MV. The affine model is used to derive the MV for each sub-block of the current affinely coded block, and motion compensation is performed at the sub-block level to obtain the predictor of the current affinely coded block. The SAD or SATD (or MRSAD or MRSATD) of two predictors of the current affinely coded block (one from the L0 reference picture and the other from the L1 reference picture) is calculated to derive the bilateral matching cost for each current search point. The integer sample offset search ends after a number of search iterations, resulting in an optimal point that can be output as the refined MV, or fractional sample refinement can be further applied to it before outputting the refined MV.

[0128] In one example, the CPMV is refined. Starting from an initial set of CPMVs that reference the initial points, MV refinement offsets are added to each CPMV according to the following equations 25, 26, 27, 28, 29, and 30 to obtain the surrounding search points. CPMV0_l0' = CPMV0_l0 + MV_offset CPMV0_l1' = CPMV0_l1 - MV_offset CPMV1_l0' = CPMV1_l0 + MV_offset CPMV1_l1' = CPMV1_l1 - MV_offset CPMV2_l0' = CPMV2_l0 + MV_offset CPMV2_l1' = CPMV2_l1 - MV_offset Here, CPMVx_l0 is the x-th CPMV referring to the reference picture L0, and CPMVx_l1 is the x-th CPMV referring to the reference picture L1. MV_offset is the MV refinement offset for the search point, that is, the difference between the initial CPMV and the refined CPMV. An affine model is applied to calculate the MVs of each sub-block, and then sub-block level motion compensation is applied to obtain the predictor of the current block. During the refined MV search, the affine model parameters described above are fixed. The refined MV search includes multiple search iterations, but these iterations collectively constitute one search, and it should be understood that the refined MV search itself (since the MV refinement offset is not variable) does not need to be executed multiple times.

[0129] During each iteration in the refined MV search, multiple search points are searched around the search center, where each search point is an individual luma sample as described above, and a standard search method can be applied to search for multiple search points. For example, as shown in Figure 3A, a 3×3 square search around the search center can be performed in the integer sample offset search, and then fractional search as well as fractional error surface estimation method can be applied to derive the optimal MV offset. As another example, a cross search around the search center is used to reduce the points to be searched. As shown in Figure 3B, point 0 is the position referred to by the initial MV and is set as the first search center. Therefore, points 1 to 4 around the search center are searched first, and the cost of each search point is calculated.

[0130] In the first search iteration, it is discovered that point 4 has the minimum cost, point 4 is set as the second search center, and points 5, 6, and 7 are searched. In the next search iteration, since it is discovered that the cost of point 7 is smaller than the costs of points 4, 5, and 6, the third search center is set to point 7, and points 8 - 10 are searched. In the next search iteration, since it is discovered that point 9 has the minimum cost among points 7 - 10, point 9 is set as the fourth search center. In the next search iteration, it is discovered that the costs of points 6, 16, and 18 around point 9 are all greater than the cost of point 9, and then point 9 is the optimal point.

[0131] As another example, a 3×3 square search and a 3×3 cross search can be executed in cooperation. The square search is executed in the first k iterations, and then followed by the cross search executed to determine the optimal point, or the cross search is executed first, and then the square search is executed to determine the optimal point.

[0132] In another embodiment, a VVC standard encoder and decoder configure one or more processors of a computing system to apply an adaptive search step to accelerate refined MV search. In the first one or more of the k earlier search iterations, the search step is set to an initial step (e.g., 2), and in one or more subsequent search iterations following the earlier search iterations, the search step is changed to a smaller step (e.g., 1). As shown in Figure 3C, point 0 is the position referred to by the initial MV and is set as the first search center. Points 1 - 8 around the initial point are first searched, and since the search step is set to 2, the distance between points 1, 2, 3, 4, 5, 6, 7, 8 and point 0 is 2 pixels (the Manhattan distance is used here instead of the Euclidean distance).

[0133] Next, in the second search iteration, the cost at point 7 is the minimum cost, the search center is point 7, and the search step is still set to 2. The search points 9, 10, and 11 are all 2 pixels away from the search center point 7, and it is found that point 10 has the minimum cost.

[0134] Next, in the third search iteration, the search center is point 10 and the search step is changed to 1. The search points 12 to 14 are each 1 pixel away from point 10. Assume that point 13 is found to have the minimum cost.

[0135] Next, in the fourth search iteration, the search center is 13, the search step is 1, and the points 11, 15, 16, 17, and 18 are each 1 pixel away from the search center point 13. The adaptive search step can also be applied in a cross-search pattern or another search pattern.

[0136] For each search point, the sub-block MV is recalculated using the CPMV of the current search point, and the block predictor is derived by performing sub-block based affine motion compensation using the derived sub-block MV.

[0137] The cost of each search point can be calculated as cost = mvDistanceCost + sadCost, where sadCost is the SAD between the L0 predictor and the L1 predictor of the current block, and mvDistanceCost is based on the distance between the search point and the initial point (i.e., the difference between the refined CPMV and the initial CPMV).

[0138] To control the complexity of refinement, PROF may or may not be applied before SAD calculation during refined MV search. When PROF is applied, for each search point, after obtaining the predictor, PROF is applied to refine the predictor, and SAD is calculated between the two PROF-refined predictors. When PROF is not applied, for each search point, SAD is calculated directly after the L0 predictor and the L1 predictor are obtained.

[0139] Also, to further reduce the complexity of refined MV search, for each search point, the predictor can be generated using a 2-tap bilinear interpolation filter instead of an 8-tap or 12-tap interpolation filter.

[0140] In another example, a VVC standard encoder and decoder configure one or more processors of a computing system to perform sub-block MV refinement. The sub-block MV is derived using an initial CPMV, and then an MV refinement offset is added to each sub-block MV to refine the sub-block MV according to the following equations 31 and 32. MV(sbx)_l0' = MV(sbx)_l0 + MV_offset MV(sbx)_l1' = MV(sbx)_l1 - MV_offset Here, MV(sbx)_l0 and MV(sbx)_l1 are the MVs of sub-block X referring to reference picture L0 and reference picture L1, respectively. The MV_offset from the current search point is added to all of the sub-block MVs to obtain refined sub-block MVs. Then, motion compensation is performed at the sub-block level using each sub-block MV to obtain two predictors for the entire CB. SAD or SATD (or MRSAD or MRSATD) is calculated between the two predictors to obtain the bilateral matching cost of the current search. The search point with the minimum cost is treated as the optimal point, and the corresponding MV_offset is obtained as the optimal MV refinement offset. The refined sub-block MVs can be calculated according to equations 31 and 32. PROF may or may not be applied before SAD calculation during refined MV search. An interpolation filter with a reduced number of taps can be used instead of an 8-tap or 12-tap interpolation filter to generate the predictor during refined MV search. The search method and cost calculation method applied in CPMV refinement can also be applied in sub-block MV refinement.

[0141] To reduce search complexity, pre-interpolated samples of predictors at all search points within a search window can be stored in a buffer before refined MV search. Then, for each search point, the predictors of each sub-block can be fetched directly from the buffer without interpolation. As shown in FIG. 13, a coding block is divided into 16 4×4 sub-blocks for affine motion compensation. The initial CPMV is used to calculate the initial sub-block MV for each sub-block. For each sub-block, a reference sub-block (i.e., the predictor of the sub-block) can be located in the reference picture by the initial sub-block MV. Since refinement is applied to the sub-block MV, for each search point, all of the reference sub-blocks are shifted by the same MV refinement offset. Thus, for each sub-block, samples of the reference sub-blocks at all search points within each search window can be pre-interpolated as the shaded areas in FIG. 13. Then, for a search point, samples of the reference block at that search point can be fetched directly from the search window without interpolation, thereby saving the computation time and resources spent in further interpolation.

[0142] The VVC standard encoder and decoder can configure one or more processors of a computing system to apply pre - interpolation in integer sample offset search and fractional sample refinement (i.e., so that for each search point, all samples in the search window are interpolated before refined MV search such that the predicted sample can be directly fetched and interpolation does not need to be called, i.e., the predicted sample can be directly fetched and there is no need to call interpolation). In the case of integer sample offset search, the pre - interpolated samples are all 1 pixel apart from each other. In phase - dependent fractional sample refinement, the pre - interpolated samples can be stored separately. As shown in FIG. 14, squares indicate samples for integer search points that can all be pre - interpolated, the × symbol indicates samples at the 1 / 2 - pel position horizontally but at integer positions vertically, triangles indicate samples at the 1 / 2 - pel position vertically but at integer positions horizontally, and circles indicate samples at the 1 / 2 - pel position in both horizontal and vertical directions. In view of the 1 / 2 - pel search process, there can be three different phases of samples, and the distance between two neighboring fractional samples with the same phase is also 1 pixel. Thus, fractional samples of the same type can be pre - interpolated together, and samples of different phases can be stored separately.

[0143] After CB - level MV refinement, sub - CB - level MV refinement can also be applied. For example, in affine motion compensation, a CB is divided into a plurality of sub - CBs (e.g., 16×16 sub - CBs) that are larger than sub - blocks. For each sub - CB, the affine model MV is further refined, and sub - blocks within one sub - CB share the same MV refinement offset. The final MV for each respective sub - block can be formulated according to the following equations (33) and (34). MV(sbx)_l0' = MV(sbx)_l0+MV_offset+MV_offset(sbCUx) MV(sbx)_l1' = MV(sbx)_l1 - MV_offset - MV_offset(sbCUx) Here, MV(sbx)_l0 and MV(sbx)_l1 are the L0 and L1 MVs of sub-block X respectively. MV_offset is the MV refinement offset obtained in the CB-level MV refinement process, and MV_offset(sbCUx) is the MV refinement offset obtained in the sub-CB level MV refinement for sub-CB X where sub-block X is located. MV(sbx)_l0' and MV(sbx)_l1' are the final refined MVs for sub-block X. Motion compensation is performed at the sub-block level using the refined sub-block MVs.

[0144] The search range can be set according to the PU size or QP in order to reduce the calculation complexity at the cost of performance reduction, or to improve the performance at the cost of increased calculation complexity. As an example, since larger CBs may benefit from higher refinement, a larger search range can be set for larger CBs and a smaller search range can be set for smaller CBs. As another example, a larger QP may benefit from a larger search range due to larger distortion. Thus, a larger search range can be set for larger QPs and a smaller search range can be set for smaller QPs. Alternatively, in order to achieve encoding time reduction, the number of points to be searched can be significantly reduced by setting a smaller search range for larger QPs. Thus, in some examples, a smaller search range can be set for larger QPs and a larger search range can be set for smaller QPs.

[0145] In order to further reduce the calculation complexity, a PU size limit can be imposed. The PU size limit parameter specifies that the process of affine DMVR is skipped for some sizes of CBs. For example, affine DMVR is not applied to CBs smaller than 8×8 or 16×16, or affine DMVR is not applied to CBs larger than 64×64 or 128×128.

[0146] In some examples, to reduce the computational complexity, one or more processors of a computing system are configured such that a VVC standard encoder and decoder apply early termination to the DMVR refined MV search for affine blocks. As an example, if a search point is checked in the search of previous candidates in an affine merge candidate list, that search point is skipped. As another example, if the SAD at a search point is less than a predefined threshold, the refined MV search ends and that search point is used as the optimal position after refinement.

[0147] In some examples, to reduce the complexity for the encoder and decoder, one or more processors of a computing system are configured such that a VVC standard encoder and decoder apply a fast algorithm. As an example, if the difference between the current affine merge candidate (i.e., the current initial MV) and the previous checked affine merge candidate (i.e., the previous initial MV) is less than a threshold, the entire process of DMVR for the current affine merge candidate is skipped. The current affine merge candidate is used directly for motion compensation of the current block without refinement. The threshold may depend on the current block size such that larger blocks have a larger threshold than smaller blocks. As another example, DMVR is disabled for affine coded blocks having a size smaller / greater than a threshold. DMVR is not applied to blocks smaller than the threshold or to blocks greater than the threshold to refine the motion. The threshold may be fixed or signaled in the bitstream.

[0148] As described above, the parameters of the affine model, including a, b, c, d, and mv wx wx , mv wy wy can be refined in DMVR, and thus, in addition to the MV of the affine model, the parameters of the affine model can also be refined.

[0149] In some examples, the VVC standard encoder and decoder configure one or more processors of a computing system to perform four-parameter refinement in an affine model. Based on the affine model shown in Equation 23, offset_a, offset_b, offset_c, and offset_d are added to parameters a, b, c, and d, respectively, to refine these parameters. Both the encoder and decoder search for offset_a, offset_b, offset_c, and offset_d to minimize the bilateral matching cost (whether calculated based on SAD or SATD as described above, or based on MRSAD or MRSATD) between the L0 predictor (i.e., the reference block from reference picture L0) and the L1 predictor (i.e., the reference block from reference picture L1) of the affine-coded block (hereinafter referred to as "affine parameter offset search", where the parameter offset that minimizes the bilateral matching cost is hereinafter referred to as the "optimal parameter offset"). The refinement of each parameter proceeds according to the following Equations 35, 36, 37, 38, 39, 40, 41, and 42. a_l0' = a_l0 + offset_a a_l1' = a_l1 - offset_a b_l0' = b_l0 + offset_b b_l1' = b_l1 - offset_b c_l0' = c_l0 + offset_c c_l1' = c_l1 - offset_c d_l0' = d_l0 + offset_d d_l1' = d_l1 - offset_d Here, a_l0, b_l0, c_l0, and d_l0 are the parameters of the affine model for reference list 0, and a_l1, b_l1, c_l1, and d_l1 are the parameters of the affine model for reference list 1. Each of the four parameters is refined.

[0150] After the affine parameter refinement, the sub-block MVs of list 0 and list 1 can be derived according to the affine models of equations 22 - 24 using a_l0', b_l0', c_l0', d_l0', and a_l1', b_l1', c_l1', d_l1'. Then, the sub-block MVs can be derived, and motion compensation can be performed at the sub-block level.

[0151] Furthermore, the bilateral matching cost (regardless of whether it is based on SAD or SATD, or MRSAD or MRSATD) can be calculated at either the sub-block level or the CB level. When the bilateral matching cost is calculated at the sub-block level, after the motion compensation of each sub-block, the cost of that sub-block is calculated. After obtaining the sub-block costs, the CB level cost is calculated by summing up all the sub-block costs.

[0152] When the bilateral matching cost is calculated at the CB level, the motion compensation of each sub-block is first performed to obtain the L1 and L0 predictors for the entire CB, and then the overall cost is calculated. To reduce the computational complexity, in some embodiments, not all sub-blocks or all samples of the CB are necessarily considered in the bilateral matching cost calculation. Since only the differences of some sub-blocks or some samples are calculated, the motion compensation of the sub-blocks or samples not considered in the cost calculation can also be skipped.

[0153] Alternatively, the VVC standard encoder and decoder configure one or more processors of the computing system to perform two-parameter refinement in the affine model. In the case of the four-parameter affine model, since b is equal to -c and d is equal to a, the refinement can also follow this constraint, i.e., offset_b is equal to -offset_c and offset_d is equal to offset_a. Therefore, the encoder and decoder only need to perform an affine parameter offset search for offset_a and offset_b, and derive offset_b and offset_d according to offset_a and offset_b.

[0154] Since the affine parameter offset search for the two-parameter offset is not more computationally complex than the affine parameter offset search for the four-parameter offset, the constraint that offset_b is equal to -offset_c and offset_d is equal to offset_a can also be applied in the DMVR applied to the six-parameter affine model. The refinement of each parameter proceeds according to the following equations 43, 44, 45, 46, 47, 48, 49, and 50. a_l0' = a_l0 + offset_a a_l1' = a_l1 - offset_a b_l0' = b_l0 - offset_c b_l1' = b_l1 + offset_c c_l0' = c_l0 + offset_c c_l1' = c_l1 - offset_c d_l0' = d_l0 + offset_a d_l1' = d_l1 - offset_a

[0155] Furthermore, the four-parameter refinement can also be applied to a four-parameter model or a six-parameter model. For both the four-parameter affine model and the six-parameter affine model, the same refinements as the above equations 35, 36, 37, 38, 39, 40, 41, and 42 are applied.

[0156] In some examples, the VVC standard encoder and decoder configure one or more processors of a computing system to apply the refined MV search method in affine parameter offset search. For example, as shown in FIG. 3A or FIG. 3B, a 3×3 square search or a 3×3 cross search can be applied to yield an optimal parameter offset. In four-parameter refinement, the square search or the cross search is performed in four-dimensional space. A 3×3×3×3 square search or a 3×3×3×3 cross search can be applied to obtain an optimal parameter offset. In the case of a 3×3×3×3 square search, there are 80 neighboring positions to be searched for each center position, and in the case of a 3×3×3×3 cross search, there are 8 neighboring positions to be searched for each center position, which is much less than that of a 3×3×3×3 square search.

[0157] Assuming that the parameter offset of the current center position is (offset_a, offset_b, offset_c, offset_d), the eight neighboring positions to be searched in the 3×3×3×3 cross search are (offset_a + step_a, offset_b, offset_c, offset_d), (offset_a - step_a, offset_b, offset_c, offset_d), (offset_a, offset_b + step_b, offset_c, offset_d), (offset_a, offset_b - step_b, offset_c, offset_d), (offset_a, offset_b, offset_c + step_c, offset_d), (offset_a, offset_b, offset_c - step_c, offset_d), (offset_a, offset_b, offset_c, offset_d + step_d), (offset_a, offset_b, offset_c, offset_d - step_d), where step_a, step_b, step_c, and step_d are the search steps for parameters a, b, c, and d, respectively.

[0158] After the optimal parameter offset is returned from the affine parameter offset search, an error surface-based offset estimation can also be applied to further refine the returned parameters with higher accuracy. The search step of the integer search can be a fixed value. According to some examples, since the MV accuracy is 1 / 16 in the ECM and the basic sub-block for affine motion compensation is 4×4, the search step can be 1 / 64 so that the MV difference between two adjacent sub-blocks is 1 / 64 * 4 = 1 / 16, which is the minimum difference for the MV. In some other examples, the search step can be larger than 1 / 64. A larger search step reduces the search iterations and thus reduces the search time, but may sacrifice the refinement accuracy. In another example, the search step depends on the CB size. Let the width of the CB be denoted as w and the height of the CB be denoted as h, and the search step for a and c be step acshown as, and the search steps for d and b are step bd shown as, the search steps can satisfy the following Equation 51. w × step ac = T1 h × step bd = T2 where T1 and T2 are two thresholds, and the thresholds can be 1 / 16, 1 / 8, 1 / 4, or other values. These thresholds define the maximum MV difference between the MVs of any two samples within the block. According to this example, different parameters have different search steps.

[0159] Regarding the cost of each search point, the difference in parameter offsets can also be considered. The cost can be the weighted sum of SAD or SATD (or MRSAD or MRSATD) between two predictors of the coding block and the parameter offset, such as bilCost = w * ParameterOffsetCost + sadCost, where w is the weight, sadCost is the SAD / SATD or mean-removed SAD / SATD cost of the predictor, and ParameterOffsetCost is the cost that depends on the parameter offset of the refined parameter. When w is equal to 0, only sadCost is considered.

[0160] Furthermore, according to an exemplary embodiment of the present disclosure, during the affine parameter offset search, the MV of the affine model can be fixed.

[0161] During any affine parameter offset search, one MV of the affine model is fixed. Unlike the refined MV search, multiple affine parameter offset searches can be performed and different MVs of the affine model can be fixed during different affine parameter offset searches. Any MV at any point in the plane is treated as an affine model MV and can thus be fixed in the parameter search. In one implementation, the CPMV is treated as the affine model MV that is fixed during the search. For example, the upper left CPMV is fixed and the affine model parameters are refined as shown in FIG. 15(a). When there is a change in the parameters, the coding block rotates and zooms in / out, so the upper right CPMV and the lower right CPMV also change. Then, the sub-block MVs are derived from the refined CPMVs (i.e., derived from the refined parameters and the new CPMVs) and motion compensation is performed. FIGS. 15(b) and 15(c) show examples where the upper right CPMV and the lower left CPMV are treated as the affine model MVs that are fixed during the refinement of the four affine model parameters, respectively. Similar to FIG. 15(a), when there is a refinement of the affine model parameters, the coding block rotates and zooms in / out, so the CPMVs that are not fixed change. Multiple affine parameter offset searches can be performed because the non-fixed CPMVs can result in different optimal parameter offsets (as well as different MV refinement offsets).

[0162] In another implementation, various CPMVs, each treated as an affine model MV, are fixed in order for several affine parameter offset searches that are executed in order. In the first affine parameter offset search, the upper left CPMV is fixed and the parameter offset is searched as shown in Fig. 15(a). Along with the refined parameters, the upper right CPMV can also be changed and calculated accordingly. Then, in the second affine parameter offset search, as shown in Fig. 15(b), the refined upper right CPMV is fixed and the parameters are refined again. Along with the refined parameters obtained in the second sub-step, the lower left CPMV can also be changed and calculated accordingly. Then, in the third affine parameter offset search, the refined lower left CPMV is fixed and the parameters are refined again as shown in Fig. 15(c). The same affine parameter offset search can be repeated several times. After the third affine parameter offset search, the first affine parameter offset search can continue with a new upper left CPMV fixed, and the search can continue until several conditions are met. For example, the conditions include, but are not limited to, 1) a preset number of iterations, 2) SAD or SATD (or MRSAD or MRSATD) between the L0 predictor and the L1 predictor that is less than a threshold, 3) the currently fixed CPMV is the same or similar to that in the last iteration, 4) the optimal parameter offset generated by the latest affine parameter offset search is less than a threshold.

[0163] In yet another embodiment, instead of the CPMV, the zero MV found in the plane is treated as the affine model MV that is fixed during refinement. First, the point at (x, y) with the zero MV is derived according to the following Equation 52.

[0164]

Equation

[0165] Assume that the solution is (x1, y1), and thus the affine model can be represented according to Equation 53 below using zero MV.

[0166]

Number

[0167] Then, the affine model parameters a, b, c, and d are searched to find refined values. All of the above refinement methods can be applied in this embodiment.

[0168] As described previously, the affine parameter refinement process is similar to the MV refinement process. The search is performed iteratively, and the iterations collectively constitute one affine parameter offset search. For each iteration, if the bilateral matching cost at the center position is smaller than that at all neighboring positions, the current center position is found as the optimal position and the search ends; otherwise, the neighboring position with the minimum bilateral matching cost is set as the new center position and the search proceeds to the next iteration.

[0169] Therefore, to control the search complexity, the VVC standard encoder and decoder are configured to perform, in each affine parameter offset search, the maximum number of search iterations, where the number of search iterations is configured on both the encoder side and the decoder side, to configure one or more processors of the computing system to perform. The search ends when either the center position has the minimum cost or the maximum search iteration threshold reaches a pre-set maximum iteration threshold. A larger maximum search iteration threshold can result in a larger coding performance gain but can consume longer encoding and decoding times.

[0170] To achieve a good trade-off between complexity and performance, the maximum search iteration threshold can be set according to fixed MVs, QPs, time layers, CB sizes, etc. For example, according to the search processes shown in FIGS. 15(a) to 15(c), in the first affine parameter offset search, the upper left CPMV is fixed, and the maximum search iteration threshold is set to a larger value such as 8 compared to subsequent thresholds. Since it is only when the parameters are refined, the larger search iteration threshold can configure one or more processors of the computing system to improve coding and decoding performance by utilizing a longer calculation time.

[0171] Next, in the second affine parameter offset search, the upper right CPMV is fixed, and the maximum search iteration threshold is set to a smaller value such as 6 compared to the previous threshold. Since the parameters have already been refined in the first affine parameter offset search, the smaller search iteration threshold can save coding and decoding time.

[0172] Next, in the third affine parameter offset search, the lower left CPMV is fixed, and the maximum search iteration threshold is set to a smaller value such as 2 compared to the previous threshold, further saving coding and decoding time. Therefore, in this embodiment, the maximum search iteration threshold is initially set to a larger value and changed to a smaller value in subsequent affine parameter offset searches.

[0173] In some other embodiments, the maximum search iteration threshold for subsequent affine parameter offset search depends on the actual number of search iterations performed during the previous affine parameter offset search. For example, in the first affine parameter offset search, when the upper left CPMV is fixed, the maximum search iteration threshold is set to N. However, during the first affine parameter offset search, the search actually ends in the k-th (k < N) search iteration because the center position has the minimum bilateral matching cost. Next, in the second affine parameter offset search, the maximum search iteration threshold is set to k / 2 (or another value that depends on k and is less than P).

[0174] In contrast, in the first affine parameter offset search, when the maximum number of search iterations enabled by the threshold is performed, in the second affine parameter offset search, the maximum search iteration threshold is set to P, which is a value less than N. A similar method can be applied in the third affine parameter offset search, that is, when the actual number of search iterations performed in the second affine parameter offset search is the maximum number, the maximum search iteration threshold for the third affine parameter offset search is set to L, provided that L is less than P, and when the actual number of search iterations is t, which does not reach the maximum number, the maximum search iteration threshold for the third affine parameter offset search is set to t / 2. Therefore, the maximum search iteration threshold is adaptively determined in the previous search iteration.

[0175] In some other embodiments, to reduce complexity, the search neighborhood locations for the search iteration are adaptively reduced according to the previous search iteration. For example, in a 3×3×3×3 cross-search pattern, there are 8 neighboring locations to be searched in each search iteration. Assuming the current center is (a, b, c, d), the 8 neighboring locations to be checked are, respectively, pa0 = (a + s, b, c, d), pa1 = (a - s, b, c, d), pb0 = (a, b + s, c, d), pb1 = (a, b - s, c, d), pc0 = (a, b, c + s, d), pc1 = (a, b, c - s, d), pd0 = (a, b, c, d + s), and pd1 = (a, b, c, d - s). The bilateral matching costs for the 8 neighboring locations are denoted as cost_pa0, cost_pa1, cost_pb0, cost_pb1, cost_pc0, cost_pc1, cost_pd0, and cost_pd1.

[0176] The VVC standard encoder and decoder configure one or more processors of the computing system to compare cost_pa0 and cost_pa1, and if cost_pa0 is less than cost_pa1, only the positive offset is considered for parameter a in the next iteration, and if cost_pa0 is greater than cost_pa1, only the negative offset is considered for parameter a in the next iteration.

[0177] The VVC standard encoder and decoder configure one or more processors of the computing system to compare cost_pb0 and cost_pb1, and if cost_pb0 is less than cost_pb1, only the positive offset is considered for parameter b in the next iteration, and if cost_pb0 is greater than cost_pb1, only the negative offset is considered for parameter b in the next iteration.

[0178] The VVC standard encoder and decoder configure one or more processors of the computing system to compare cost_pc0 and cost_pc1. If cost_pc0 is less than cost_pc1, only the positive offset is considered for parameter c in the next iteration. If cost_pc0 is greater than cost_pc1, only the negative offset is considered for parameter c in the next iteration.

[0179] The VVC standard encoder and decoder configure one or more processors of the computing system to compare cost_pd0 and cost_pd1. If cost_pd0 is less than cost_pd1, only the positive offset is considered for parameter d in the next iteration. If cost_pd0 is greater than cost_pd1, only the negative offset is considered for parameter d in the next iteration.

[0180] For the current search iteration, assume that cost_pa0 is less than cost_pa1, cost_pb0 is greater than cost_pb1, cost_pc0 is less than cost_pc1, and cost_pd0 is greater than cost_pd1. Then, in the next search iteration, the four neighboring positions to be checked will be (a'+s, b', c', d'), (a', b'-s, c', d'), (a', b', c'+s, d'), and (a', b', c', d'-s), where (a', b', c', d') is the center position of the next search iteration.

[0181] In some other embodiments, the minimum bilateral matching cost of the current search iteration is compared with that of the previous search iteration or with that of the previous search iteration multiplied by a coefficient f. If the minimum cost reduction is small, the search ends. For example, if the cost of the previous search iteration is A and the cost of the current search center means it is also A, the minimum cost of the neighboring position is B at position posb, provided that B < A. According to this search rule, the search proceeds to the next iteration with the search center posb. However, in this embodiment, if A - B < K or B > A * f, the search ends and posb is selected as the optimal position for this search iteration. K and f are preset thresholds. For example, f is a coefficient less than 1 such as 0.95, 0.9, or 0.8.

[0182] QP controls quantization in video coding. At a higher QP, a larger quantization step is used, and thus, a larger distortion is introduced. Therefore, for a higher QP, more search iterations are required in refinement and the coding time increases. To reduce the total coding time, in this embodiment, a smaller maximum search iteration threshold is set for a higher QP than for a lower QP.

[0183] Other methods for reducing complexity can also be used at a high QP. For example, reducing the neighboring positions to be searched, adaptively reducing the search iterations, or ending the search early according to the previous search process can be implemented. Therefore, in this embodiment, different search strategies can be adopted at different QPs.

[0184] Inter-coded frames, such as B-frames or P-frames, have one or more reference frames. The temporal distance between the current frame and the reference frames affects the accuracy of inter-prediction. The temporal distance between two frames in video coding is usually represented by the POC distance. Usually, with a longer POC distance, the inter-prediction accuracy is lower, the motion information accuracy is also lower, and thus further refinement is required. Therefore, in this embodiment, the search process depends on the POC distance between the current frame and the reference frames.

[0185] In the case of hierarchical B-frames, frames with a higher temporal layer have a short POC distance to the reference frames, and frames with a lower temporal layer have a longer POC distance to the reference frames. Therefore, the search process may also depend on the temporal layer of the current frame. For example, affine parameter refinement may be disabled for a high temporal layer because the high temporal layer has a short POC distance to the reference frames and does not require refinement. In another example, a small search iteration threshold is set or the neighboring search positions are reduced for high temporal layer frames.

[0186] Also, other methods for reducing the complexity of parameter refinement can be used for high temporal layer frames. Therefore, in this embodiment, the parameter refinement process depends on the temporal distance or POC distance between the current frame and the reference frames.

[0187] In all of the above embodiments, the affine model parameters are directly refined. However, the affine motion can include translation, rotation, and zooming. Translation is represented by the MV of the affine model, and rotation and zooming are represented by the affine model parameters. In another embodiment, the rotation and zoom motions are explicitly refined. Based on the original affine model, additional rotation and scaling are added. When the original affine model is described as Equation 22 above, rotation using the angle t and scaling using the coefficient k are applied according to Equation 54 below.

[0188]

Equation

[0189] Here, t and k are two parameters that will be searched during the DMVR process. The current search method can be applied to obtain the optimal values of t and k. Then, the sub-block MV is derived according to Equation 49 above, and sub-block-based affine motion compensation is performed to obtain the predictor of the current affine-coded block.

[0190] All existing early termination methods in MV refinement can also be applied in parameter refinement. For example, during refined MV search, if the SAD or SATD (or MRSAD or MRSATD) between two predictors is less than the threshold, the refined MV search ends.

[0191] Similar to the case of MV refinement, PROF may or may not be applied before SAD calculation during the search process of affine model parameters. When PROF is applied, for each search point, after obtaining the predictor, PROF is applied to refine the predictor, and SAD is calculated between two PROF-refined predictors. When PROF is not applied, for each search point, SAD is directly calculated after the L0 predictor and the L1 predictor are obtained.

[0192] Also, to further reduce the search complexity, for each search point, the predictor can be generated using an interpolation filter having a reduced number of taps instead of an 8-tap or 12-tap interpolation filter.

[0193] Affine parameter refinement and MV refinement can be executed in cooperation. In one embodiment, MV refinement and affine parameter refinement are executed continuously. After MV refinement, affine parameter refinement can follow, or after affine parameter refinement, MV refinement can follow.

[0194] In an alternative embodiment, MV refinement and affine parameter refinement are refined in the same process. For the affine model according to Equation 55 below,

[0195] [Number]

[0196] Six parameters (a, b, c, d, and mv0x, mv0y) are refined together. offset_a, offset_b, offset_c, offset_d, offset_mv 0x , and offset_mv 0y The optimal values of are found in the search, and the parameters are refined according to the following Equations 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, and 67, respectively. a_l0' = a_l0 + offset_a a_l1' = a_l1 - offset_a b_l0' = b_l0 + offset_b b_l1' = b_l1 - offset_b c_l0' = c_l0 + offset_c c_l1' = c_l1 - offset_c d_l0' = d_l0 + offset_d d_l1' = d_l1 - offset_d mv 0x _l0' = mv 0x _l0 + offset_mv 0x mv 0x _l1' = mv 0x _l1 - offset_mv 0x mv 0y _l0' = mv 0y _l0 + offset_mv 0y mv 0y _l1' = mv 0y _l1 - offset_mv 0y Here, a_l0, b_l0, c_l0, and d_l0 are the initial parameters of the affine model for reference picture list 0, a_l1, b_l1, c_l1, and d_l1 are the initial parameters of the affine model for reference picture list 1, (mv 0x _l0, mv 0y _l0) is the initial MV of the affine model for reference picture list 0, (mv 0x _l1, mv 0y _l1) is the initial MV of the affine model for reference picture list 1. a_l0', b_l0', c_l0', and d_l0' are the refined parameters of the affine model for reference list 0, a_l1', b_l1', c_l1', and d_l1' are the refined parameters of the affine model for reference list 1, (mv 0x _l0', mv 0y _l0') is the refined MV of the affine model for reference picture list 0, (mv 0x _l1', mv 0y_l1') is the refined MV of the affine model for the reference picture list 1. After the refinement of these six parameters, the CPMV can be recalculated, and then the sub-block MV can be derived. Sub-block level motion compensation can be performed. This embodiment can also be combined with the four-parameter refinement limitation, i.e., offset_b is equal to -offset_c and offset_d is equal to offset_a.

[0197] Those skilled in the art will understand that all of the above aspects of the present disclosure can be implemented simultaneously in any combination thereof, and all aspects of the present disclosure can be implemented in combination as another embodiment of the present disclosure.

[0198] FIG. 16 shows an exemplary system 1600 for implementing the processes and methods described herein for implementing decoder-side motion vector refinement for affine motion compensation.

[0199] The techniques and mechanisms described herein can be implemented by multiple instances of system 1600, as well as by any other computing device, system, and / or environment. System 1600 shown in FIG. 16 is merely an example of a system and does not imply any limitation regarding the use or functionality scope of any computing device utilized to execute the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, game consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, implementations using field programmable gate arrays (FPGA) and application specific integrated circuits (ASIC), and the like.

[0200] System 1600 may include one or more processors 1602 and a system memory 1604 communicatively coupled to the processors 1602. The processors 1602 may execute one or more modules and / or processes to cause the processors 1602 to perform various functions. In some embodiments, the processors 1602 may include a central processing unit (CPU), a graphics processing unit (GPU), both a CPU and a GPU, or other processing units or components known in the art. Further, each of the processors 1602 may have its own local memory, which may also store program modules, program data, and / or one or more operating systems.

[0201] Depending on the exact configuration and type of the system 1600, the system memory 1604 may be volatile such as RAM, non-volatile such as ROM, flash memory, a small hard drive, a memory card, etc., or some combination thereof. The system memory 1604 may include one or more computer-executable modules 1606 executable by the processors 1602.

[0202] The modules 1606 may include, without limitation, an encoder module 1608 and a decoder module 1610. The encoder module 1608 and the decoder module 1610 may be configured to execute any of the methods described above.

[0203] System 1600 may further include an input / output (I / O) interface 1640 for receiving video source data and bitstream data and for outputting decoded pictures to a reference picture buffer and / or a display buffer. System 1600 may also include a communication module 1650 that enables System 1600 to communicate with other devices (not shown) via a network (not shown). The network may include wired media such as the Internet, a wired network, or a direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0204] In some embodiments, a method is provided that includes performing refined MV search for the CPMV of an inter-coded CB and outputting the refined MV of the CB. The method may be performed by a decoder or an encoder.

[0205] Some or all of the operations of the methods described above may be performed by the execution of computer-readable instructions stored on a computer-readable storage medium as defined below. The term "computer-readable instructions" as used herein and in the claims includes routines, applications, application modules, program modules, programs, components, data structures, algorithms, and the like. The computer-readable instructions may be implemented on a variety of system configurations including single-processor or multiprocessor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like.

[0206] A computer-readable storage medium may include volatile memory (such as random-access memory (“RAM”)), and / or non-volatile memory (such as read-only memory (“ROM”), flash memory, etc.). The computer-readable storage medium may also include additional removable storage and / or non-removable storage including, but not limited to, flash memory, magnetic storage, optical storage, and / or tape storage that may provide non-volatile storage such as computer-readable instructions, data structures, program modules, etc.

[0207] A non-transitory computer-readable storage medium is an example of a computer-readable medium. Computer-readable media include at least two types of computer-readable media, namely, computer-readable storage media and communication media. A computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any process or technology for the storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, phase change memory ("PRAM"), static random-access memory ("SRAM"), dynamic random-access memory ("DRAM"), other types of random-access memory ("RAM"), read-only memory ("ROM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory or other memory technologies, compact disk read-only memory ("CD-ROM"), digital versatile disk ("DVD") or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by a computing device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism. The computer-readable storage media employed herein is not to include signals per se that are interpreted as transitory signals, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (such as optical pulses through an optical fiber cable), or electrical signals propagating through a wire.

[0208] Computer-readable instructions stored on one or more non-transitory computer-readable storage media that, when executed by one or more processors, can perform the operations described above with reference to FIGS. 1A-15. Generally, computer-readable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular abstract data type. The order in which the operations are described is not to be construed as limiting, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0209] In some embodiments, a computer program product is provided, the program product including computer program instructions that enable a computer to perform the steps of the method described in any of the embodiments of the present disclosure.

[0210] In some embodiments, a computer program is provided, the computer program enabling a computer to perform the steps of the method described in any of the embodiments of the present disclosure.

[0211] The subject matter has been described in language specific to structural features and / or methodological acts, but it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms for implementing the claims.

Description of Reference Numerals

[0212] 100 Encoding process 102 Image source 104 Intra prediction 106 Inter prediction 108 Mode decision and encoder control setting 110 Subtractor 112 Transformation 114 Quantization 116 Inverse Quantization 118 Inverse Transformation 120, 162 Adder 122, 164 Loop Filter 124 Entropy Coder 150 Decoding Process 152 Entropy Decoder 154 Inverse Quantization 156 Inverse Transformation 158 Intra Prediction 160 Motion Compensation Prediction 200 Decoded Picture Buffer (「DPB」), DPB 202 Current Picture 202A Current Block 204, 206 Collocated Reference Picture 204A, 206A Collocated Reference Block, Reference Block 204B, 206B Offset Block 1600 System 1602 Processor 1604 System Memory 1606 Computer Executable Module, Module 1608 Encoder Module 1610 Decoder Module 1640 Input / Output (I / O) Interface 1650 Communication Module

Claims

1. A computing system comprising: one or more processors; and a computer-readable storage medium communicatively coupled to the one or more processors, the computer-readable storage medium storing computer-readable instructions executable by the one or more processors, the computer-readable instructions, when executed by the one or more processors, performing related operations; wherein the related operations include: performing refined motion vector (MV) search for control point motion vectors (CPMVs) of inter-coded coding blocks (CBs); and outputting the refined MVs of the CBs. A computing system as described above.

2. The operations further include: outputting a refined MV for each of a plurality of CPMVs of the CB; wherein each refined MV is derived from the same MV refinement offset applied to the respective initial CPMV. The computing system according to Claim 1.

3. Performing refined MV search includes: deriving MVs of sub-blocks of the CB based on the CPMVs of the CB; performing sub-block MV refinement for the MVs of the sub-blocks; outputting the refined MVs of the sub-blocks; and outputting the refined MV of the CB based on the refined MVs of the sub-blocks. The computing system according to Claim 1.

4. Performing refined MV search further includes performing a plurality of search iterations of integer sample offset search at a plurality of search points around a search center to output an integer distance refined MV. The computing system according to Claim 2.

5. The operations further include obtaining predicted sample values of the CB based on the MVs of the sub-blocks by fetching pre-interpolated samples of each search point around a search center without performing interpolation. The computing system according to Claim 3.

6. Performing refined MV search includes: deriving affine model parameters based on a plurality of CPMVs of the CB; Performing an affine parameter offset search for the affine model parameters, Outputting an optimal parameter offset, Based on the optimal parameter offset, outputting the refined MV of the CB The computing system according to claim 1, comprising:

7. Deriving affine model parameters based on a plurality of CPMVs, Determining four affine model parameters based on three CPMVs, Determining two affine model parameters based on two CPMVs, or Determining four affine model parameters based on three CPMVs The computing system according to claim 6, comprising one of the above.

8. While fixing the CPMV of the CB, an affine parameter offset search is performed, Outputting the refined MV of the CB is further based on the fixed CPMV. The computing system according to claim 6 or 7.

9. Performing the refined MV search includes performing a plurality of refined MV searches in sequence, and each refined MV search Performing an affine parameter offset search for the affine model parameters while fixing different CPMVs of the CB The computing system according to any one of claims 6 to 8, comprising:

10. The operation is Outputting the refined MV of each of the plurality of CPMVs of the CB Further including Each refined MV is derived from different MV refinement offsets applied to their respective initial CPMVs. The computing system according to claim 1.

11. Each refined MV search further includes performing a plurality of search iterations at a plurality of search points around the search center up to a maximum search iteration threshold, and each maximum search iteration threshold is smaller in each subsequent refined MV search. The computing system according to claim 1.

12. Performing the refined MV search Performing a plurality of search iterations of integer sample offset search at a plurality of search points around the search center to output an integer distance refined MV, To output a sub-pixel accuracy refined delta MV, applying fractional sample refinement to the integer distance refined MV The computing system according to claim 1, further comprising.

13. The computing system according to claim 1, wherein each refined MV search further comprises performing a plurality of search iterations at a plurality of search points around a search center.

14. The computing system according to claim 13, wherein performing an iteration among the plurality of search iterations includes one of searching for a plurality of search points by square search around the search center and searching for a plurality of search points by cross search around the search center.

15. Performing an iteration among the plurality of search iterations is respectively, calculating the bilateral matching cost of each search point among the plurality of search points based on two derived predictors of the CB from the reference picture in the first reference picture list and the reference picture in the second reference picture list; determining a minimum bilateral matching cost among the bilateral matching costs of each search point among the plurality of search points; ending the refined MV search based on a determination that the minimum bilateral matching cost is greater than the minimum bilateral matching cost of the previous iteration among the plurality of search iterations multiplied by a coefficient The computing system according to claim 13, comprising.

16. Performing an iteration among the plurality of search iterations is respectively, calculating the bilateral matching cost of each search point among the plurality of search points based on two derived predictors of the CB from the reference picture in the first reference picture list and the reference picture in the second reference picture list; determining a minimum bilateral matching cost among the bilateral matching costs of each search point among the plurality of search points; ending the refined MV search based on the difference between the minimum bilateral matching cost and the minimum bilateral matching cost of the previous iteration among the plurality of search iterations The computing system according to claim 13, comprising.

17. Performing an iteration among the plurality of search iterations is Calculating the bilateral matching cost for each search point among the plurality of search points, respectively, based on two derived predictors of the CB, one from a reference picture in a first reference picture list and the other from a reference picture in a second reference picture list; Determining a minimum bilateral matching cost among the calculated bilateral matching costs; The computing system according to claim 13, comprising:

18. The computing system according to claim 17, wherein calculating the bilateral matching cost for each search point among the plurality of search points is based on the distance between each respective search point and the search center.

19. The computing system according to claim 17, wherein the bilateral matching cost is calculated based on the sum of absolute differences, the sum of absolute transform differences, the sum of mean removed absolute differences, or the sum of mean removed absolute transform differences between the two derived predictors.

20. The computing system according to claim 17, wherein performing an iteration among the plurality of search iterations further includes not applying prediction refinement by optical flow (PROF) to the two derived predictors before calculating the bilateral matching cost.

21. Performing refined MV search for the CPMV of the inter-coded CB; Outputting the refined MV of the CB; A method comprising:

22. The method further includes: Outputting a refined MV for each of a plurality of CPMVs of the CB; wherein each refined MV is derived from the same MV refinement offset applied to the respective initial CPMV, the method according to claim 21.

23. The step of performing refined MV search includes: Deriving an MV of a sub-block of the CB based on the CPMV of the CB; Performing sub-block MV refinement for the MV of the sub-block; Outputting the refined MV of the sub-block; Outputting the refined MV of the CB based on the refined MV of the sub-block; The method according to claim 21, comprising:

24. ​ The method according to claim 22, wherein the step of performing refined MV search further includes the step of performing a plurality of search iterations of integer sample offset search at a plurality of search points around the search center in order to output an integer distance refined MV.

25. The method includes the step of obtaining the predicted sample value of the CB based on the MV of the sub-block by fetching the pre-interpolated samples of each search point around the search center without performing interpolation. The method according to claim 23.

26. The step of performing refined MV search includes the step of deriving affine model parameters based on a plurality of CPMVs of the CB, the step of performing an affine parameter offset search for the affine model parameters, the step of outputting an optimal parameter offset, and the step of outputting the refined MV of the CB based on the optimal parameter offset. The method according to claim 21.

27. The step of deriving affine model parameters based on a plurality of CPMVs includes one of the steps of determining four affine model parameters based on three CPMVs, determining two affine model parameters based on two CPMVs, or determining four affine model parameters based on three CPMVs. The method according to claim 26.

28. The affine parameter offset search is performed while fixing the CPMV of the CB, and the step of outputting the refined MV of the CB is further based on the fixed CPMV. The method according to claim 26 or 27.

29. The step of performing refined MV search includes the step of performing a plurality of refined MV searches in sequence, and each refined MV search includes the step of performing an affine parameter offset search for the affine model parameters while fixing different CPMVs of the CB. The method according to any one of claims 26 to 28.

30. The method further includes the step of outputting the refined MV of each of the plurality of CPMVs of the CB. ​ The method according to claim 21, wherein each refined MV is derived from a different MV refinement offset applied to each respective initial CPMV.

31. The method according to claim 21, further comprising the step of performing a plurality of search iterations at a plurality of search points around a search center up to a maximum search iteration threshold, wherein each respective maximum search iteration threshold is smaller in each subsequent refined MV search.

32. The step of performing refined MV search comprises performing a plurality of search iterations of integer sample offset search at a plurality of search points around a search center to output an integer distance refined MV; and applying fractional sample refinement to the integer distance refined MV to output a sub-pixel accuracy refined delta MV The method according to claim 21, further comprising.

33. The method according to claim 21, wherein each refined MV search further comprises the step of performing a plurality of search iterations at a plurality of search points around a search center.

34. The method according to claim 33, wherein the step of performing an iteration among the plurality of search iterations comprises one of the steps of searching for a plurality of search points by a square search around the search center and searching for a plurality of search points by a cross search around the search center.

35. The step of performing an iteration among the plurality of search iterations comprises calculating a bilateral matching cost for each search point among the plurality of search points based on two derived predictors of the CB from a reference picture in a first reference picture list and a reference picture in a second reference picture list, respectively; determining a minimum bilateral matching cost among the bilateral matching costs of each search point among the plurality of search points; ending the refined MV search based on a determination that the minimum bilateral matching cost is greater than the minimum bilateral matching cost of the previous iteration among the plurality of search iterations multiplied by a coefficient The method according to claim 33, comprising.

36. The step of performing an iteration among the plurality of search iterations comprises Calculating the bilateral matching cost for each search point among the plurality of search points based on two derived predictors of the CB from the reference picture in the first reference picture list and the reference picture in the second reference picture list, respectively; Determining the minimum bilateral matching cost among the bilateral matching costs of each search point among the plurality of search points; Ending the refined MV search based on the difference between the minimum bilateral matching cost and the minimum bilateral matching cost of the previous iteration among the plurality of search iterations The method according to claim 33, comprising.

37. The step of performing an iteration among the plurality of search iterations is Calculating the bilateral matching cost for each search point among the plurality of search points based on two derived predictors of the CB from the reference picture in the first reference picture list and the reference picture in the second reference picture list, respectively; Determining the minimum bilateral matching cost among the calculated bilateral matching costs The method according to claim 33, comprising.

38. The method according to claim 37, wherein the step of calculating the bilateral matching cost for each search point among the plurality of search points is based on the distance between each search point and the search center.

39. The method according to claim 37, wherein the bilateral matching cost is calculated based on the sum of absolute differences, the sum of absolute transform differences, the sum of mean removed absolute differences, or the sum of mean removed absolute transform differences between the two derived predictors.

40. The method according to claim 37, wherein the step of performing an iteration among the plurality of search iterations further comprises the step of not applying prediction refinement by optical flow (PROF) to the two derived predictors before calculating the bilateral matching cost.

41. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of the device to cause the device to start the method according to any one of claims 21 to 40.

42. A computer program product, wherein the computer program product includes computer program instructions for enabling a computer to execute the method according to any one of claims 21 to 40.

43. A computer program for enabling a computer to execute the method according to any one of claims 21 to 40.