AMVP method and device for video coding and decoding in merge mode

By simplifying the AMVP merge mode, reordering the merge candidates using MV cost, and generating target AMVP merge candidates, solving the problems of large computing volume and high bandwidth of the AMVP merge mode, achieving more efficient coding performance.

CN119999207APending Publication Date: 2025-05-13MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380056167.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-07-28
Filing Date
2023-07-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The AMVP merge mode is computationally expensive, and the bilateral matching algorithm used for MV refinement will significantly increase the system bandwidth and affect the encoding performance.

Method used

The simplified AMVP merge mode is adopted to reorder the merge candidate list through MV cost, generate the target AMVP merge candidate, and use the candidate during the encoding or decoding process.

Benefits of technology

Reduces system bandwidth requirements, improves coding performance, and reduces computing and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999207A_ABST
    Figure CN119999207A_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for video coding and decoding using a simplified AMVP merge mode or using an AMVP merge mode with a BCW. According to one of the methods, a target AMVP candidate is determined from an AMVP candidate list. A merge candidate list is determined. And reordering the merge candidates in the merge candidate list according to the MV cost related to the merge candidates to form a reordered merge candidate list. A target merge candidate is determined from the reordered merge candidate list. And generating a target AMVP combined candidate according to the target AMVP candidate and the target combined candidate. According to another method, a target weight pair is determined from a weight pair set comprising two or more weight pairs. And generating a target AMVP merging candidate according to the target weight pair, and taking the target AMVP merging candidate as a weighted sum of the target AMVP candidate and the target merging candidate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [Cross-reference to related applications]

[0002] The present invention claims priority to U.S. Provisional Patent Application Serial No. 63 / 369,672 filed on July 28, 2022. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety. [Technical field]

[0003] The present application relates to video coding using motion estimation and motion compensation. In particular, the present application relates to a scheme for improving the performance of AMVP merge mode. [Background technology]

[0004] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology-Coded representation of immersive media-Part 3: Versatile video coding, released in February 2021. VVC is developed on the basis of its predecessor HEVC (High Efficiency Video Coding), adding more codec tools to improve codec efficiency and handle various types of video sources, including three-dimensional (3D) video signals.

[0005] Figure 1A A self-adjusting intra / inter video coding system including a loop process is shown. For intra prediction, prediction data is derived based on previously encoded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed at the encoder end, and motion compensation (MC) is performed based on the results of ME to provide prediction data derived from other images and motion data. Switch 114 selects intra prediction 110 or inter prediction 112, and the selected prediction data is provided to adder 116 to form a prediction error, also called a residual. The prediction error is then processed by transform (T) 118 and quantization (Q) 120. The transformed and quantized residual is then encoded by entropy encoder 122 for inclusion in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients will then contain auxiliary information such as motion and coding modes related to intra prediction and inter prediction, as well as other information such as parameters related to the loop filter applied to the underlying image area. As shown Figure 1AAs shown, auxiliary information related to intra prediction 110, inter prediction 112, and in-loop filter 130 is provided to entropy encoder 122. When using inter prediction mode, one or more reference pictures must also be reconstructed at the encoder end. Therefore, the transformed and quantized residuals are processed by inverse quantization (IQ) 124 and inverse transform (IT) 126 to restore the residuals. The residuals are then added back to the prediction data 136 at reconstruction (REC) 128 to reconstruct the video data. The reconstructed video data can be stored in a reference picture buffer 134 and used to predict other frames.

[0006] like Figure 1A As shown, the input video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subjected to various impairments due to the series of processing. Therefore, before the reconstructed video data is stored in the reference picture buffer 134, an in-loop filter 130 is usually applied to the reconstructed video data to improve the video quality. For example, a deblocking filter (DF), sample self-adjusting offset (SAO), and self-adjusting loop filter (ALF) can be used. The loop filter information may need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. Figure 1A In FIG. 1 , an in-loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in a reference picture buffer 134 . Figure 1A The system in is intended to illustrate an example structure of a typical video codec. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, ​​H.264, or VVC.

[0007] like Figure 1B As shown, the decoder may use similar or partially identical functional blocks as the encoder, except for transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. Instead of using the entropy encoder 122, the decoder uses an entropy decoder 140 to decode the video bitstream into quantized transform coefficients and required coding information (such as ILPF information, intra-frame prediction information, and inter-frame prediction information). The intra-frame prediction 150 on the decoder side does not need to perform a pattern search. Instead, the decoder only needs to generate an intra-frame prediction based on the intra-frame prediction information received from the entropy decoder 140. In addition, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from the entropy decoder 140, without performing motion estimation.

[0008] According to VVC, the input image is divided into non-overlapping square block areas called CTUs (Coding Tree Units), similar to HEVC. Each CTU can be divided into one or more smaller coding units (CUs). The resulting CU partitions can be square or rectangular. In addition, VVC also divides CTUs into prediction units (PUs) as units for applying prediction processes such as intra-frame prediction, inter-frame prediction, etc.

[0009] Compared with the HEVC standard, the VVC standard adopts various new coding tools to further improve the coding and decoding efficiency. Among the various new coding tools, some coding tools related to the present application are summarized as follows.

[0010] Bi-prediction with CU-level weights (BCW)

[0011] In HEVC, the bi-prediction signal is generated by averaging two prediction signals and, respectively, these two prediction signals come from two different reference pictures and / or use two different motion vectors. In VVC, the bi-prediction mode goes beyond the scope of simple averaging and allows for a weighted average of the two prediction signals.

[0012] P bi-pred =((8-w)*P 0 +w*P 1 +4)>3

[0013] Weighted average bi-prediction allows five weights, w∈{-2,3,4,5,10}. For each bi-predicted CU, the weight w can be determined in one of two ways: 1) For non-merged CUs, the weight index is signaled after the motion vector difference; 2) For merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW is only applicable to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. The encoder uses a fast search algorithm to find the weight index without significantly increasing the complexity of the encoder. These algorithms are outlined below. For details, see VTM software and JVET-L0646 document (Yu-Chi Su et al., "CE4-related: Generalized bi-prediction improvements combined from JVET-L0197 and JVET-L0296", ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 Joint Video Experts Group (JVET), 12th meeting: Macau, China, October 3-12, 2018, document: JVET-L0646).

[0014] When used in conjunction with AMVR, unequal weighting of 1-pel and 4-pel motion vector precision is conditionally checked only when the current picture is a low-latency picture.

[0015] When used in conjunction with affine, affine ME is performed on unequal weights only when the affine mode is selected as the current best mode.

[0016] Unequal weights are only checked conditionally when the two reference pictures in bi-prediction are the same.

[0017] Based on the picture order count (POC) distance between the current picture and its reference pictures, the coding QP, and the temporal level, unequal weights are not searched when certain conditions are met.

[0018] The BCW weight index codec uses a context codec partition followed by a bypass codec partition. The first context codec partition indicates whether equal weights are used; if unequal weights are used, an additional partition using bypass codecs indicates which unequal weights are used.

[0019] Weighted prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficient coding of video content with attenuation. The VVC standard also adds support for WP. WP allows weight parameters (weights and offsets) to be set for each reference picture in each reference picture list L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures will be applied. WP and BCW are designed for different types of video content. In order to avoid the interaction between WP and BCW (which complicates the design of the VVC decoder), if a CU uses WP, the BCW weight index is not displayed and the weight w is inferred to be 4 (i.e., equal weights are applied). For merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. This applies to both normal merge mode and inherited affine merge mode. For constructed affine merge mode, affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using constructed affine merge mode only needs to be set equal to the BCW index of the first control point MV.

[0020] In VVC, CIIP and BCW cannot be applied to a CU at the same time. When a CU is encoded using CIIP mode, the BCW index of the current CU will be set to 2 (ie, w=4 means equal weight). Equal weight means the default value of the BCW index.

[0021] Bidirectional Optical Flow (BDOF)

[0022] The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF was formerly known as BIO and is included in JEM. Compared to the JEM version, BDOF in VVC is much simpler and requires significantly less computation, especially in terms of the number of multiplications and multiplier size.

[0023] BDOF is used to improve the bi-prediction signal of a CU at the 4×4 sub-block level. BDOF can be used if the CU meets all of the following conditions:

[0024] CU uses "real" dual-prediction mode encoding and decoding, that is, one of the two reference pictures is arranged before the current picture in display order, and the other is arranged after the current picture in display order.

[0025] The distances from the two reference images to the current image (i.e., POC difference) are the same

[0026] Both reference images are short-term reference images

[0027] CU does not use affine mode or SbTMVP merge mode encoding

[0028] CU has more than 64 luma samples

[0029] CU height and CU width are both greater than or equal to 8 luma samples

[0030] BCW Weight Index means equal weight

[0031] WP is not enabled for the current CU

[0032] The current CU does not use CIIP mode

[0033] BDOF is only applied to the luma component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ). Motion refinement is then used to adjust the bi-predicted sample values ​​in the 4×4 sub-block.

[0034] Decoder-side Motion Vector Refinement (DMVR) in VVC

[0035] In order to improve the accuracy of the merge mode motion vector, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bilateral prediction operation, a refined MV is searched for the current block 220 of the current picture 210 around the initial MVs (232 and 234) in the reference picture list L0 22 and the reference picture list L1 214. Figure 2 As shown, according to the initial MV (232 and 234) and the position of the current block 220 in the current picture, the co-located blocks 222 and 224 in L0 and L1 are determined. The BM method calculates the distortion between the two candidate blocks (242 and 244) in the reference picture list L0 and the list L1. The positions of the two candidate blocks (242 and 244) are determined by adding two opposite offsets (262 and 264) to the two initial MVs (232 and 234), thereby deriving two candidate MVs (252 and 254). As shown Figure 2 As shown, the SAD between the candidate blocks (242 and 244) is calculated based on each candidate MV around the initial MV (232 or 234). The candidate MV (252 or 254) with the smallest SAD value becomes the refined MV for generating the bi-prediction signal.

[0036] In VVC, the application of DMVR is limited and only applies to CUs using the following modes and feature codecs:

[0037] CU-level merge mode with bi-predicted MV

[0038] Relative to the current image, one reference image is in the past and the other reference image is in the future

[0039] The distances from the two reference images to the current image (i.e., POC difference) are the same

[0040] Both reference images are short-term reference images

[0041] CU has more than 64 luma samples

[0042] CU height and CU width are both greater than or equal to 8 luma samples

[0043] BCW Weight Index means equal weight

[0044] The current block is not WP enabled

[0045] The current block does not use CIIP mode

[0046] The refined MV obtained through the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction of future picture coding. The original MV is used in the deblocking process and is also used for spatial motion vector prediction of future CU coding and decoding.

[0047] AMVP-Merge mode

[0048] In JVET-X0083 (Zhi Zhang et al., "EE2: Bilateral and template matching AMVP-merge mode (test 3.3)", ITU-T SG16WP 3 and ISO / IEC JTC 1 / SC 29 Joint Video Experts Group (JVET), 24th Meeting, Teleconference, October 6-15, 2021, Document: JVET-X0083), an AMVP merge mode is proposed. In the proposed AMVP merge mode, the bidirectional predictor consists of an AMVP predictor in one direction and a merged predictor in the other direction.

[0049] According to the JVET-X0083 standard, the AMVP part is sent as a regular unidirectional AMVP (i.e., the reference index and MVD signal are sent), with a derived MVP index if template matching is used (such as asserting TM_AMVP), and the MVP index is sent if template matching is disabled. The merge index is not sent, and the merge predictor will be selected from the list of candidates with the smallest template or bilateral matching cost.

[0050] When the selected merged predictor and AMVP predictor meet the DMVR condition, that is, there is at least one past reference picture and one future reference picture relative to the current picture, and the distances of the two reference pictures to the current picture are the same, bilateral matching MV refinement will be applied starting from the merged MV candidate and AMVP MVP. Otherwise, if the template matching function is enabled, template matching MV refinement will be applied to the merged predictor or AMVP predictor with a higher template matching cost.

[0051] For multi-channel DMVR, the AMVP merge mode codec block enables the third channel, i.e., 8x8 sub-PU BDOF refinement.

[0052] Although the AMVP merge mode can improve the encoding and decoding performance, the computational complexity of this motion compensation mode is quite large. In addition, the bilateral matching algorithm used for MV refinement will significantly increase the system bandwidth because the algorithm must access bidirectional reference data. Therefore, the present application discloses a solution to reduce the system bandwidth requirements and / or further improve performance. [Summary of the invention]

[0053] The present application discloses a method and device for performing video encoding and decoding using a simplified AMVP merge mode. The method includes: receiving input data associated with a current block in a current picture; wherein the input data includes pixel data of the current block for encoding at an encoder side, or coded data associated with the current block for decoding at a decoder side; determining a target AMVP (Advanced Motion Vector Prediction) candidate from an AMVP candidate list; determining a merge candidate list; reordering the merge candidates in the merge candidate list according to motion vector (MV) costs associated with the merge candidates to form a reordered merge candidate list; wherein each of the MV costs is measured based on one or more factors including an MV distance, the MV distance being measured between each candidate in the merge candidate list and a corresponding mapped MV; wherein each of the candidates associated with a first reference picture is mapped to the corresponding mapped MV in a second reference picture; determining a target merge candidate from the reordered merge candidate list; generating a target AMVP merge candidate based on the target AMVP candidate and the target merge candidate; and encoding or decoding the current block using a motion candidate list including the target AMVP merge candidate.

[0054] In some embodiments, one or more of the merge candidates with lower MV costs are moved forward in the reordered merge candidate list. In some embodiments, the merge candidates with lower MV costs are assigned to smaller merge indexes. In some embodiments, the one or more factors include reference indices and / or picture order counts (POCs) associated with the first reference picture and the second reference picture. In some embodiments, each of the MV costs is derived from a weighted sum of at least two of the one or more factors.

[0055] Another method includes: determining a target AMVP candidate from an AMVP candidate list; determining a target merge candidate from a merge candidate list; determining a target weight pair from a weight pair set including two or more weight pairs; generating a target AMVP merge candidate based on the target weight pair as a weighted sum of the target AMVP candidate and the target merge candidate; and encoding or decoding the current block using a motion candidate list including the target AMVP merge candidate.

[0056] In some embodiments, the weight index associated with the target weight pair is inherited from the target merge candidate. In some embodiments, the weight index associated with the target weight pair is derived based on a template matching cost, where the template matching cost is calculated based on a first neighboring template of an AMVP reference block and a second neighboring template of a merged reference block. In some embodiments, the weight index associated with the target weight pair is derived based on any one or more of the following: a slice quantization parameter (QP) of an AMVP reference picture and a merged reference picture, a picture order count (POC) difference between the AMVP reference picture and the current picture, a POC between the merged reference picture and the current picture, and a reference picture ID associated with the AMVP reference picture and the merged reference picture.

[0057] In some embodiments, the weight index associated with the target weight pair is emitted in the bitstream or parsed from the bitstream.

[0058] Another method includes: receiving input data associated with a current block in a current picture; wherein the input data includes pixel data of the current block for encoding on the encoder side, or encoding data associated with the current block for decoding on the decoder side; determining a target AMVP (Advanced Motion Vector Prediction) candidate from an AMVP candidate list; wherein the AMVP candidate list includes affine motion candidates; determining a target merge candidate from a merge candidate list; generating a target AMVP merge candidate for a dual prediction candidate by using the target AMVP candidate in one direction and the target merge candidate in another direction; and encoding or decoding the current block using a motion candidate list including the target AMVP merge candidate.

[0059] In some embodiments, the merge candidate list further includes the affine motion candidate.

[0060] In some embodiments, the affine motion candidate can be included in the AMVP candidate list only when the current block is greater than or less than a threshold. In some embodiments, the threshold corresponds to a 32x32 block size. In some embodiments, when the affine motion candidate is allowed to be included in the AMVP candidate list, the AMVP candidate list and the merged candidate list have the same motion type. In some embodiments, the same motion type corresponds to translational motion or sub-block based motion.

Brief Description of the Drawings

[0061] Figure 1A A self-adjusting intra / inter video coding system including a loop process is shown.

[0062] Figure 1B Shown Figure 1A The corresponding decoder of the encoder in .

[0063] Figure 2 An example of decoder-side motion vector refinement (DMVR) in VVC is shown.

[0064] Figure 3 An exemplary flow chart of a video coding system using a simplified AMVP merge mode provided by an embodiment of the present application is shown.

[0065] Figure 4 An exemplary flowchart of a video coding system using a bi-prediction AMVP merge mode with CU-level weight (BCW) provided in an embodiment of the present application is shown.

[0066] Figure 5 An exemplary flow chart of a video encoding system using the AMVP merge mode with affine motion is shown. [Specific implementation method]

[0067] It is easy to understand that the components of the present application, as generally described and illustrated in the figures, can be arranged and designed in various configurations. Therefore, the following more detailed description of the embodiment of the system and method of the present application, as shown in the figures, is not intended to limit the scope of the present application, but only represents the selected embodiment of the present application. "One embodiment", "one embodiment" or similar language mentioned in this specification means that the specific features, structures, or characteristics related to the embodiment may be included in at least one embodiment of the present application. Therefore, the phrases "in one embodiment" or "in one embodiment" that appear in various places in this specification do not necessarily all refer to the same embodiment.

[0068] In addition, the described features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. However, those of ordinary skill in the relevant art will recognize that the present application can be implemented without one or more specific details, and can also be implemented with other methods, components, etc. In other cases, in order to avoid covering up the various aspects of the present application, the present application does not show or describe well-known structures or operations in detail. The drawings can better understand the illustrated embodiments of the present application, wherein the same parts are represented by the same numbers throughout the text. The following description is only carried out by way of example, and only some selected embodiments of the equipment and methods consistent with the scope of the present application are described.

[0069] As mentioned above, the computational complexity of the AMVP merge mode is quite large, and the bilateral matching algorithm used for MV refinement will significantly increase the system bandwidth. Therefore, the present application discloses a solution to reduce the system bandwidth requirement and / or further improve the performance.

[0070] Method 1: Simplify the reordering scheme of the merge candidate list in AMVP merge mode

[0071] In JVET-X0083, the merge candidates of AMVP merge mode will first be reordered by bilateral matching, and then one of the top two candidates after reordering will be selected as the final candidate of AMVP merge mode. An EP (equal probability) partition will be signaled to indicate which candidate is selected. Since bilateral matching reordering of the merge candidate list during AMVP merge may require higher bandwidth, it is recommended to replace this step with other methods.

[0072] In one embodiment, it is suggested to reorder the merge candidate list using MV cost. The motion with lower MV cost will be moved forward and represented by a smaller merge index, where motion refers to motion information related to the candidate, such as motion vector and corresponding reference picture list and reference picture index. For example, the MV on one side (i.e., MvL0 or MvL1) will be mapped to the reference picture on the other side (i.e., L1 or L0) according to the picture order count (POC). In other words, the mapping is performed relative to the current picture so that the MV and the mapped MV are located on both sides of the current picture. Afterwards, the absolute distance between the two actions will be used as the MV cost. The closer the two motions are, the smaller the MV cost. Alternatively, the farther the two motions are, the smaller the MV cost. In another example, not only the distance between the two actions can be used as the MV cost, but also the reference picture index of the two reference pictures or the POC distance between the current picture and the respective reference pictures can be considered. For example, if the distance between the two motions of the two candidates is the same, the reference picture index or POC distance of the two reference pictures is used for determination. For another example, more than one criterion can be used for reordering, and the final MV cost can be calculated by the weighted sum of multiple factors. (i.e., the distance between two motions, the reference picture index, or the POC of the reference picture.) For another example, more than one criterion may be used for reordering and compared by a hierarchical algorithm. For another example, two motion distances may be compared first, and only when the two motion distances are the same, the reference picture index may be compared.

[0073] Method 2: Enable AMVP merging with BCW

[0074] In the original AMVP merge mode disclosed in JVET-X0083, AMVP prediction and merge prediction use the default equal-weight merge. In one embodiment of the present application, it is recommended to support BCW in the AMVP merge mode. The BCW index can be inherited from the merge candidate. In another embodiment, the BCW index is designed based on the template matching (TM) cost. For example, the adjacent templates of the reference block in L0 and L1 of the AMVP merge mode are used to calculate the TM cost. The adjacent templates can be N upper templates and M left templates, where N and M can be any integer greater than 0. For the side with a larger TM cost (L0 or L1), a smaller weight is used in dual prediction mixing. For the other side with a smaller TM cost (L0 or L1), a larger weight is used in dual prediction mixing. For another example, TM refinement or other motion refinement processes are applied in the AMVP merge mode. After the final refined motion is derived, a plurality of BCW indices with different weight pairs are tested for the refined motion. The weight pair with the lowest TM cost is selected as the final weight pair. As another example, BCW index selection can be performed together with the TM refinement process or other motion refinement process. Therefore, during the motion refinement process, BCW indices with different weight pairs are tested under different motion offsets. For example, each offset in the diamond search is tested with a different weight pair. In this way, the best BCW index and offset pair can be selected. In another embodiment, the BCW index signal is sent directly to the decoder to indicate the weight pair of the AMVP merge mode as a regular inter-frame mode. In another embodiment, when the size of the current CU is less than or greater than a preset threshold, a BCW index signal is sent to the decoder to indicate the weight pair of the AMVP merge mode.

[0075] In another embodiment, a BCW index signal is directly sent to the decoder to indicate that the weight pair of the AMVP merge mode is a conventional inter-frame mode. The weight pair with the corresponding BCW index having a smaller weight is used for the AMVP predictor, and the weight pair with the corresponding BCW index having a larger weight is used for the merge predictor. Alternatively, the design can be reversed, that is, the weight pair with the corresponding BCW index having a larger weight is used for the merge predictor.

[0076] In one embodiment, the BCW index is derived according to the QP of the current block, the POC difference between the AMVP reference picture and the merged reference picture, the reference picture index associated with the AMVP reference picture and the merged reference picture, or any combination thereof.

[0077] Method 3: Enable AMVP merging with Affine

[0078] In the original AMVP merge mode disclosed in JVET-X0083, if AMVP merge is used, the affine mode will be disabled. In one embodiment, not only translational motion can be used to predict the AMVP part of the AMVP merge mode, but affine motion can also be used to predict the AMVP part of the AMVP merge mode. In another embodiment, the CU size restriction is used to determine whether the affine mode of AMVP merge is supported. For example, the affine of the AMVP merge mode can only be enabled when the size of the CU is greater than or less than a threshold (such as 32x32). In another embodiment, not only translational motion can be used to predict the merge part of the AMVP merge mode, but sub-block-based merge candidates (i.e., affine motion) can also be used to predict the merge part of the AMVP merge mode. In another embodiment, the motion type of the AMVP part and the merge part of the AMVP merge mode can be further restricted to be the same. For example, if the AMVP merge mode with affine motion candidates is enabled, the motions of the AMVP part and the merge part are both translational motions. For another example, if the AMVP merge mode with affine motion candidates is enabled, the motions in the AMVP and merge parts are both sub-block-based motions.

[0079] Any of the AMVP merging methods proposed above can be implemented in an encoder and / or decoder. For example, any of the proposed methods can be implemented in the interactive encoding and decoding of the encoder (such as Figure 1A Inter-frame prediction 112 in the decoder) and / or the interactive coding module of the decoder (such as Figure 1B Alternatively, any of the proposed methods may be implemented as a circuit coupled to an inter-coding of an encoder and / or a decoder, thereby providing information required for the inter-coding.

[0080] Figure 3An exemplary flow chart of a video encoding and decoding system using a simplified AMVP merge mode provided by an embodiment of the present application is shown. The steps shown in the flow chart can be implemented as program codes that can be executed on one or more processors (e.g., one or more CPUs) on the encoder side. The steps shown in the flow chart can also be implemented based on hardware, such as one or more electronic devices or processors to perform the steps in the flow chart. According to the method: in step 310, input data associated with a current block in a current picture is received, wherein the input data includes pixel data of the current block for encoding on the encoder side, or encoding data associated with the current block for decoding on the decoder side. In step 320, a target AMVP candidate is determined from an AMVP (Advanced Motion Vector Prediction) candidate list. In step 330, a merge candidate list is determined. In step 340, the merge candidates are reordered according to the MV costs associated with the merge candidates in the merge candidate list to form a reordered merge candidate list; wherein each MV cost is measured based on one or more factors including an MV distance, and the MV distance is measured between each candidate in the merge candidate list and a corresponding mapped MV; wherein each candidate associated with a first reference picture on one side of the current picture is mapped to a corresponding mapped MV in a second reference picture on the other side of the current picture. In step 350, a target merge candidate is determined from the reordered merge candidate list. In step 360, a target AMVP merge candidate is generated based on the target AMVP candidate and the target merge candidate. In step 370, the current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

[0081] Figure 4 An exemplary flow chart of a video coding system using a dual-prediction AMVP merge mode with CU-level weights (BCW) provided by an embodiment of the present application is shown. According to the method: in step 410, input data associated with a current block in a current picture is received, wherein the input data includes pixel data of the current block for encoding on the encoder side, or coded data associated with the current block for decoding on the decoder side. In step 420, a target AMVP candidate is determined from an AMVP (Advanced Motion Vector Prediction) candidate list. In step 430, a target merge candidate is determined from a merge candidate list. In step 440, a target weight pair is determined from a weight pair set including two or more weight pairs. In step 450, a target AMVP merge candidate is generated based on the target weight pair as a weighted sum of the target AMVP candidate and the target merge candidate. In step 460, the current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

[0082] Figure 5An exemplary flow chart of a video coding system using an AMVP merge mode with affine motion is shown. According to the method: in step 510, input data associated with a current block in a current picture is received, wherein the input data includes pixel data of the current block for encoding on the encoder side, or coded data associated with the current block for decoding on the decoder side. In step 520, a target AMVP candidate is determined from an AMVP (Advanced Motion Vector Prediction) candidate list, wherein the AMVP candidate list includes affine motion candidates. In step 530, a target merge candidate is determined from a merge candidate list. In step 540, a target AMVP merge candidate is generated for a dual prediction candidate by using a target AMVP candidate in one direction and a target merge candidate in another direction. In step 550, the current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

[0083] The flow chart shown is intended to illustrate an example of video encoding according to the present application. A person of ordinary skill in the art can modify each step, rearrange the steps, split the steps, or merge the steps to implement the present application without departing from the spirit of the present application. In the disclosed content, specific syntax and semantics are used to illustrate the embodiments of the present application. A person of ordinary skill can replace these syntax and semantics with equivalent syntax and semantics to implement the present application without departing from the spirit of the present application.

[0084] The above description is intended to enable a person of ordinary skill in the art to implement the present application in the context of a specific application and its requirements. It is obvious to a person of ordinary skill in the art that various modifications to the described embodiments are possible, and the general principles defined herein may be applied to other embodiments. Therefore, the present application is not intended to be limited to the specific embodiments shown and described, but rather to give the present application the broadest scope consistent with the disclosed principles and novel features. In the above detailed description, various specific details are described to provide a thorough understanding of the present application. Nevertheless, a person of ordinary skill in the art will appreciate that the present application is implementable.

[0085] The above-mentioned embodiments of the present application can be implemented by various hardware, software codes, or a combination of the two. For example, one embodiment of the present application can be a circuit integrated in a video compression chip, or a program code integrated in a video compression software to perform the processing described herein. One embodiment of the present application can also be a program code executed on a digital signal processor (DSP) to perform the processing described herein. The present application may also involve a series of functions performed by a computer processor, a digital signal processor, a microprocessor, or a field programmable gate array (FPGA). These processors can be configured to perform specific tasks of the present application by executing machine-readable software codes or firmware codes that define the specific methods embodied in the present application. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, the different code formats, styles and languages ​​of the software code, and other methods of configuring the code to perform tasks that meet the present application will not depart from the spirit and scope of the present application.

[0086] The present application may be embodied in other specific forms without departing from its spirit or essential features. The described examples should be considered in all respects as illustrative only and not restrictive. Therefore, the scope of the present application is described by the attached invention claims rather than the above description. All changes within the meaning and equivalent range of the invention claims should be included within the scope of the invention claims.

Claims

1. A video encoding and decoding method, the improvement of which comprises: Receiving input data associated with a current block in a current picture; wherein the input data includes pixel data of the current block for encoding at an encoder side, or coded data associated with the current block for decoding at a decoder side; Determine a target AMVP candidate from an AMVP (Advanced Motion Vector Prediction) candidate list; Identify a list of merger candidates; reordering the merge candidates in the merge candidate list according to motion vector (MV) costs associated with the merge candidates to form a reordered merge candidate list; wherein each of the MV costs is measured based on one or more factors including an MV distance measured between each candidate in the merge candidate list and a corresponding mapped MV; wherein each of the candidates associated with the first reference picture is mapped to the corresponding mapped MV in the second reference picture; determining a target merge candidate from the re-ordered merge candidate list; generating a target AMVP merge candidate based on the target AMVP candidate and the target merge candidate; and The current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

2. The method of claim 1, wherein: One or more of the merge candidates with lower MV costs are moved forward in the re-ordered merge candidate list.

3. The method of claim 1, wherein: The merge candidate with a lower MV cost is assigned a smaller merge index.

4. The method of claim 1, wherein: The one or more factors include reference indices and / or picture order counts (POCs) associated with the first reference picture and the second reference picture.

5. The method of claim 1, wherein: Each of the MV costs is derived from a weighted sum of at least two of the one or more factors.

6. A video codec device, the improvement comprising one or more electronic circuits or processors for: Receiving input data associated with a current block in a current picture; wherein, The input data includes pixel data of the current block for encoding at the encoder side, or coded data associated with the current block for decoding at the decoder side; Determine a target AMVP candidate from an AMVP (Advanced Motion Vector Prediction) candidate list; Identify a list of merger candidates; reordering the merge candidates in the merge candidate list according to motion vector (MV) costs associated with the merge candidates to form a reordered merge candidate list; wherein each of the MV costs is measured based on one or more factors including an MV distance measured between each candidate in the merge candidate list and a corresponding mapped MV; wherein each of the candidates associated with the first reference picture is mapped to the corresponding mapped MV in the second reference picture; determining a target merge candidate from the re-ordered merge candidate list; generating a target AMVP merge candidate based on the target AMVP candidate and the target merge candidate; and The current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

7. A video encoding and decoding method, the improvement of which comprises: Receiving input data associated with a current block in a current picture; wherein the input data includes pixel data of the current block for encoding at an encoder side, or coded data associated with the current block for decoding at a decoder side; Determine a target AMVP candidate from an AMVP (Advanced Motion Vector Prediction) candidate list; Determine a target merge candidate from the merge candidate list; determining a target weight pair from a weight pair set including two or more weight pairs; According to the target weight pair, generating a target AMVP merge candidate by a weighted sum of the target AMVP candidate and the target merge candidate; and The current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

8. The method of claim 7, wherein: The weight index associated with the target weight pair is inherited from the target merge candidate.

9. The method of claim 7, wherein: A weight index associated with the target weight pair is derived based on a template matching cost calculated based on a first neighboring template of an AMVP reference block and a second neighboring template of a merged reference block.

10. The method of claim 7, wherein: A weight index associated with the target weight pair is derived based on any one or more of the following: a slice quantization parameter (QP) of an AMVP reference picture and a merged reference picture, a picture order count (POC) difference between the AMVP reference picture and the current picture, a POC between the merged reference picture and the current picture, and a reference picture ID associated with the AMVP reference picture and the merged reference picture.

11. The method of claim 7, wherein: A weight index associated with the target weight pair is emitted in the bitstream or parsed from the bitstream.

12. A video codec device, the improvement comprising one or more electronic circuits or processors for: Receiving input data associated with a current block in a current picture; wherein, The input data includes pixel data of the current block for encoding at the encoder side, or coded data associated with the current block for decoding at the decoder side; Determine a target AMVP candidate from an AMVP (Advanced Motion Vector Prediction) candidate list; Determine a target merge candidate from the merge candidate list; determining a target weight pair from a weight pair set including two or more weight pairs; According to the target weight pair, a target AMVP merge candidate is generated by a weighted sum of the target AMVP candidate and the target merge candidate; as well as The current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

13. A video encoding and decoding method, the improvement of which comprises: Receiving input data associated with a current block in a current picture; wherein the input data includes pixel data of the current block for encoding at an encoder side, or coded data associated with the current block for decoding at a decoder side; Determining a target AMVP candidate from an AMVP (Advanced Motion Vector Prediction) candidate list; wherein the AMVP candidate list includes an affine motion candidate; Determine a target merge candidate from the merge candidate list; generating a target AMVP merge candidate for a bi-prediction candidate by using the target AMVP candidate in one direction and the target merge candidate in another direction; and The current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.

14. The method of claim 13, wherein: The merge candidate list further includes the affine motion candidate.

15. The method of claim 14, wherein: Only when the current block is larger or smaller than the threshold, the affine motion candidate can be included in the AMVP candidate list.

16. The method of claim 15, wherein: This threshold corresponds to a 32x32 block size.

17. The method of claim 13, wherein: When the affine motion candidate is allowed to be included in the AMVP candidate list, the AMVP candidate list and the merge candidate list have the same motion type.

18. The method of claim 17, wherein: The same motion type corresponds to translational motion or sub-block based motion.

19. A video codec device, the improvement comprising one or more electronic circuits or processors for: Receiving input data associated with a current block in a current picture; wherein, The input data includes pixel data of the current block for encoding at the encoder side, or coded data associated with the current block for decoding at the decoder side; Determining a target AMVP candidate from an AMVP (Advanced Motion Vector Prediction) candidate list; wherein the AMVP candidate list includes an affine motion candidate; Determine a target merge candidate from the merge candidate list; generating a target AMVP merge candidate for a bi-prediction candidate by using the target AMVP candidate in one direction and the target merge candidate in another direction; and The current block is encoded or decoded using a motion candidate list including the target AMVP merge candidate.