MMVD and Combining SMVD with Motion and Prediction Models

By extending MMVD and SMVD tools to various motion models and temporal prediction methods, the solution enhances video compression efficiency in bi-predictive scenarios, addressing limitations in existing systems.

JP7794909B2Active Publication Date: 2026-01-06INTERDIGITAL VC HOLDINGS INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024131933
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-17
Filing Date
2024-08-08
Publication Date
2026-01-06
Estimated Expiration
2039-12-16

AI Technical Summary

Technical Problem

Existing video compression systems face challenges in improving bi-prediction efficiency in inter-coded blocks, particularly in extending the use of MMVD and SMVD motion vector coding tools to all motion model derivation and temporal prediction methods supported in current video standards.

Method used

The solution involves extending the use of MMVD and SMVD motion vector coding tools to all motion models and temporal prediction methods supported in the VVC standard, including affine, planar, recursive, triangle partition-based, GBI, and multiple hypothesis prediction methods, by combining these tools with bi-prediction cases.

Benefits of technology

This approach enhances the overall compression performance of video standards by improving the flexibility and efficiency of motion vector coding, particularly in bi-predictive scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794909000016
    Figure 0007794909000016
  • Figure 0007794909000017
    Figure 0007794909000017
  • Figure 0007794909000018
    Figure 0007794909000018
Patent Text Reader

Abstract

To improve bi-prediction in an inter-coded block compared to an existing video compression system.SOLUTION: In a general embodiment, motion modes such as merge with motion vector difference and symmetric motion vector difference are extended to motion models beyond simple translation models, e.g., in combination with merge and alternative temporal motion vector prediction modes. In an embodiment, the use of MMVD and SMVD motion vector coding tools is extended to all motion model derivation and temporal prediction methods supported in proposed video standards to improve overall compression performance. In a particular embodiment, the combination of MMVD or SMVD with affine motion models, ATMVD motion models, planar motion models, regression motion fields, triangular partition-based motion models, GBI temporal prediction methods, LIC temporal prediction methods, and multiple hypothesis prediction methods are described.SELECTED DRAWING: Figure 36
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] At least one of the present embodiments generally relates to a method or apparatus for video encoding or decoding, compression or decompression. [Background technology]

[0002] To achieve high compression efficiency, image and video coding schemes typically utilize prediction, including motion vector prediction, and transforms to exploit spatial and temporal redundancies within the video content. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlations, after which the difference between the original and predicted image, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction. Summary of the Invention [Problem to be solved by the invention]

[0003] The present invention is in the field of video compression and aims to improve bi-prediction in inter-coded blocks compared to existing video compression systems. [Means for solving the problem]

[0004] At least one of the present embodiments relates generally to a method or apparatus for video encoding or decoding, and more particularly to a method or apparatus for coding mode simplification based on a neighborhood sample dependent parametric model.

[0005] According to a first aspect, a method is provided, the method including the steps of: indicating a first motion mode through syntax in a video bitstream; indicating use of a second motion mode through the presence of syntax in the video bitstream, the syntax including, if present, information related to the second motion mode; and encoding a video block using motion information corresponding to the first motion mode and the second motion mode.

[0006] According to a second aspect, a method is provided, the method including the steps of: parsing a video bitstream for whether syntax indicates a first motion mode; parsing the video bitstream for whether syntax indicates the presence of a second motion mode, and if present, determining information related to the second motion mode; obtaining motion information corresponding to the first motion mode; and decoding a block using the motion information.

[0007] According to another aspect, an apparatus is provided, the apparatus comprising a processor, the processor being configured to encode blocks of video or decode a bitstream by performing any of the methods described above.

[0008] According to another general aspect of at least one embodiment, a device is provided that includes an apparatus according to any of the decoding embodiments and at least one of (i) an antenna configured to receive a signal, the signal including a video block; (ii) a band limiter configured to limit the received signal to a band of frequencies that includes the video block; or (iii) a display configured to display an output representing the video block.

[0009] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that includes data content generated according to any of the described encoding embodiments or variations.

[0010] According to another general aspect of at least one embodiment, a signal is provided that includes video data generated according to any of the described encoding embodiments or variations.

[0011] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.

[0012] According to another general aspect of at least one embodiment, there is provided a computer program product including instructions that, when executed by a computer, cause the computer to perform any of the described decoding embodiments or variations.

[0013] These and other aspects, features and advantages of general aspects will become apparent from the following detailed description of illustrative embodiments, which is to be read in connection with the accompanying drawings. [Effects of the Invention]

[0014] A novel method or apparatus for video encoding or decoding, compression or decompression is provided. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 illustrates an example coding tree unit and coding tree concept representing a compressed HEVC picture. [Figure 2] FIG. 10 is a diagram illustrating an example of division of a coding tree unit into coding units, prediction units, and transform units. [Figure 3]1 illustrates a standard, general-purpose video compression scheme. [Figure 4] FIG. 1 illustrates a standard, general-purpose video decompression scheme. [Figure 5] FIG. 1 is a simplified block diagram of the decoding process of the current scheme. [Figure 6] FIG. 10 is an MMVD diagram in the case of bi-prediction mode. [Figure 7] FIG. 10 is a diagram illustrating an example of SMVD (Symmetric Motion Vector Difference). [Figure 8] A diagram showing an example of ATMVP motion prediction for a coding unit. [Figure 9] The collaborative research model is an example of a simple affine model used in VTM. [Figure 10] FIG. 10 illustrates a 4×4 sub-CU based affine motion vector field. [Figure 11] FIG. 10 illustrates an example of a motion vector prediction process for an affine inter-coded coding unit. [Figure 12] FIG. 10 is a diagram showing motion vector prediction candidates in affine merge mode. [Figure 13] FIG. 10 illustrates the spatial derivation of affine motion field control points in the affine merge case. [Figure 14] FIG. 1 illustrates an exemplary planar motion vector prediction process. [Figure 15] FIG. 1 illustrates regression-based motion vector field construction. [Figure 16] FIG. 1 is a diagram of an example of dividing a coding unit into two triangular prediction units (PUs). [Figure 17] FIG. 1 illustrates non-rectangular partitioning and associated OBMC diagonal weighting. [Figure 18] FIG. 1 illustrates how LIC parameters are derived from (a) reconstructed neighboring samples and (b) corresponding collocated reference samples. [Figure 19] FIG. 10 is a diagram illustrating example neighborhood samples used to derive LIC parameters. [Figure 20] FIG. 10 illustrates the spatial locations considered in the computation of non-sub-block STMVP merging candidates. [Figure 21] FIG. 1 is a simplified block diagram of a motion vector decoding process when MMVD and bi-prediction are combined. [Figure 22] FIG. 10 illustrates an example of CPR constraints. [Figure 23] FIG. 10 illustrates the application of constraints to CPR MV in MMVD mode. [Figure 24] FIG. 10 illustrates MMVD mode adaptation in the case of ATMVP. [Figure 25] FIG. 1 is a simplified block diagram of a first version of an affine motion generation process using MMVD. [Figure 26] FIG. 10 is a simplified block diagram of a second version of the affine motion generation process using MMVD. [Figure 27] FIG. 10 illustrates the allowable magnitude limits for the differential MVD, based on the MVD index, used for the first CPMV of the considered affine motion field. [Figure 28] FIG. 10 illustrates an example of a planar MVP mode combined with MMVD. [Figure 29] FIG. 1 is a simplified block diagram of a first version of a PMVD motion generation process using MMVD. [Figure 30] FIG. 10 is a simplified block diagram of a second version of the PMVD motion generation process using MMVD. [Figure 31] FIG. 10 is a simplified block diagram of a third version of the PMVD motion generation process using MMVD. [Figure 32] FIG. 1 illustrates regression-based motion vector field construction. [Figure 33] FIG. 10 is a simplified block diagram of a version of an affine motion generation process using SMVD. [Figure 34] FIG. 1 illustrates a processor-based system for encoding and decoding according to the described aspects. [Figure 35]FIG. 1 is a diagram of one embodiment of an encoding method of the described general aspect. [Figure 36] FIG. 1 is a diagram of one embodiment of a decoding method of the described general aspect. [Figure 37] FIG. 1 shows an embodiment of an apparatus in the general aspect described. DETAILED DESCRIPTION OF THE INVENTION

[0016]

[0001] The embodiments described herein are in the field of video compression and generally relate to video compression and video encoding and decoding. The general aspects described aim to provide a mechanism for manipulating constraints in high-level video coding syntax or in video coding semantics to constrain the possible set of tool combinations.

[0017] To achieve high compression efficiency, image and video coding schemes typically utilize prediction, including motion vector prediction, and transforms to exploit spatial and temporal redundancies within the video content. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlations, after which the difference between the original and predicted image, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction.

[0018] In the HEVC (High Efficiency Video Coding, ISO / IEC 23008-2, ITU-T H.265) video compression standard, motion compensated temporal prediction is used to exploit the redundancy that exists between successive pictures of a video.

[0019] To do so, a motion vector is associated with each prediction unit (PU). Each coding tree unit (CTU) is represented in the compressed domain by a coding tree. As in Figure 1, this is a quadtree division of the CTU, and each leaf is called a coding unit (CU).

[0020] Each CU is then given some intra- or inter-prediction parameters (prediction information). To do so, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. As in Figure 2, the intra- or inter-coding mode is assigned at the CU level.

[0021] In HEVC, exactly one motion vector is assigned to each PU. This motion vector is used for motion-compensated temporal prediction of the considered PU. Therefore, in HEVC, the motion model linking a prediction block and its reference block is simply translation.

[0022] HEVC uses two modes to encode motion data, called AMVP (Adaptive Motion Vector Prediction) and Merge, respectively.

[0023] AMVP essentially consists in signaling the reference picture used to predict the current PU, the motion vector predictor index (taken from a list of two predictors), and the motion vector difference. This document deals with merge mode and therefore does not address AMVP in the following.

[0024] The merge mode consists in signaling and decoding the indexes of some motion data collected in a list of motion data predictors. This list is made up of five candidates and is constructed in the same way on the decoder side and the encoder side. Thus, the merge mode aims to derive some motion information obtained from the merge list. The merge list generally contains motion information associated with some spatial and temporal surrounding blocks available in the decoded state when the current PU is being processed.

[0025] In VTM-3 (VVC Draft 3), a new type of motion vector coding was adopted, called MMVD (Merge with Motion Vector Difference). The MMVD mode basically consists in introducing some motion vector differences that are added to some conventional merge candidates to generate some motion information of the block to encode or decode.

[0026] MMVD improves the coding efficiency of VVC coding systems. However, in VVC Draft 3, it was stated that the MMVD tool applies only in the normal translational merge mode. It may be interesting to combine the MMVD tool with some other inter-coding modes, and in particular with some other motion vector (or motion model, in the case of sub-block-based motion fields) generation tools included in VVC coding systems.

[0027] Another motion vector coding tool proposed for VVC standardization is the so-called Symmetric Motion Vector Coding (SMVD). Like MMVD, SMVD is a new mode to improve the efficiency of motion vector coding. It is implemented within the AMVP mode. However, SMVD only applies to the Translational AMVP case. For MMVD, it may be beneficial to extend it for some motion models beyond the simple translational model.

[0028] The Joint Video Exploration Team (JVET) proposal for a new video compression standard, known as the Joint Exploration Model (JEM), proposed adopting a quadtree-binary tree (QTBT) block partitioning structure for high compression performance. A block in a binary tree (BT) can be divided into two equally sized sub-blocks by splitting it horizontally or vertically in the middle. As a result, BT blocks can have rectangular shapes with unequal width and height, unlike blocks in QT, which always have square shapes with equal height and width. In HEVC, angular intra-prediction directions are defined across 180 degrees, from 45 degrees to -135 degrees, and are maintained in JEM, which makes the definition of the angular direction independent of the target block shape.

[0029] To encode these blocks, intra prediction is used to provide an estimated version of the block using previously reconstructed neighboring samples. The difference between the source block and the prediction is then encoded. In the conventional codecs mentioned above, one line of reference samples is used to the left and above the current block.

[0030] In HEVC (High Efficiency Video Coding, H.265), encoding of frames of a video sequence is based on a quadtree (QT) block partitioning structure. A frame is divided into square coding tree units (CTUs), which all undergo a quadtree-based partitioning into multiple coding units (CUs) based on a rate-distortion (RD) criterion. Each CU is intra-predicted, i.e., it is spatially predicted from causal neighboring CUs, or inter-predicted, i.e., it is temporally predicted from an already decoded reference frame. In I slices, all CUs are intra-predicted, while in P slices and B slices, CUs can be both intra-predicted and inter-predicted. For intra prediction, HEVC defines 35 prediction modes, including one planar mode (indexed as mode 0), one DC mode (indexed as mode 1), and 33 angular modes (indexed as modes 2 through 34). Angle modes are associated with prediction directions ranging from 45 degrees to -135 degrees in a clockwise direction. HEVC supports a quadtree (QT) block partitioning structure, so all prediction units (PUs) have a square shape. Therefore, the definition of prediction angles ranging from 45 degrees to -135 degrees is justified in terms of PU (prediction unit) shape. For a target prediction unit with a size of NxN pixels, the upper reference array and the left reference array are each required to be 2N+1 samples in size to cover the above-mentioned angle range for all target pixels. Considering that the height and width of a PU are of equal length, the equal lengths of the two reference arrays are also natural.

[0031] The present invention is in the field of video compression. It aims to improve bi-prediction in inter-coded blocks compared to existing video compression systems. The present invention also proposes separating the luma coding tree and the chroma coding tree for inter slices.

[0032] In the HEVC video compression standard, a picture is divided into so-called coding tree units (CTUs), whose sizes are typically 64x64, 128x128, or 256x256 pixels. Each CTU is represented in the compressed domain by a coding tree, which is a quadtree decomposition of the CTU, where each leaf is called a coding unit (CU).

[0033] Each CU is then given some intra- or inter-prediction parameters (prediction information). To do so, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. An intra- or inter-coding mode is assigned at the CU level.

[0034] Emerging video compression tools include a coding tree unit representation in the compressed domain proposed to represent picture data in a more flexible way in the compressed domain. The advantage of this more flexible representation of the coding tree is that it provides improved compression efficiency compared to the CU / PU / TU arrangement of the HEVC standard.

[0035] The quadtree plus binary tree (QTBT) coding tool provides this increased flexibility. It is essentially a coding tree that can partition a coding unit in both a quadtree and binary tree manner. Such a coding tree representation of a coding tree unit is illustrated below.

[0036] The division of the coding units is determined at the encoder side through a rate-distortion optimization procedure, which consists in determining the QTBT representation of the CTU with the minimum rate-distortion cost.

[0037] In the QTBT technique, a CU has either a square or rectangular shape. The size of a coding unit is always a power of 2, typically between 4 and 128.

[0038] In addition to this variety of rectangular shapes for the coding unit, this new CTU representation has the following different properties compared to HEVC:

[0039] The QTBT decomposition of a CTU is made up of two stages: first, the CTU is split in a quadtree manner, and then each quadtree leaf can be further split in a binary tree manner.

[0040] One problem solved by this invention is how to extend the use of the MMVD and SMVD motion vector coding tools to all motion model derivation and temporal prediction methods supported in currently proposed video standards, thereby improving the overall compression performance of these proposed standards.

[0041] The basic principle of this disclosure is two-fold.

[0042] - Extending the use of MMVD motion vector coding with all motion models and / or temporal prediction methods supported in VVC Draft 3, as well as other motion models proposed in the VVC standardization process. In particular, this disclosure describes how to combine MMVD with affine motion models, ATMVP motion models, planar motion models, recursive motion fields, triangle partition-based motion models, GBI temporal prediction methods, LIC temporal prediction methods, and multiple hypothesis prediction methods. Enhanced use of MMVD in bi-predictive cases is also provided.

[0043] - Extending the use of the SMVD motion vector coding tool with all motion models and / or temporal prediction methods supported in VVC Draft 3, as well as other motion model generators proposed to the VVC standardization process. In particular, this disclosure describes how to combine SMVD with affine motion models, ATMVP motion models, planar motion models, recursive motion fields, triangle partition-based motion models, GBI temporal prediction methods, LIC temporal prediction methods, and multiple hypothesis prediction methods.

[0044] In the developing VVC video standard, multiple inter prediction modes are supported or proposed.

[0045] Figure 5 shows a simplified block diagram of the inter-decoding process in version 3 of the VVC draft. The process starts with the derivation of an MV predictor (mvp) based on the construction of an MV candidate list (1001), which can differ according to the inter mode used for coding the CU or PU.

[0046] Step 1002 is a motion vector difference (MVd) decoding step. The decoded MVd is used to reconstruct the motion vector of the considered CU / PU through its addition to the motion vector predictor obtained in the previous step (1001).

[0047] The next step 1003 consists in deriving a motion model from the motion vector values ​​produced by the two previous steps. The model can be purely translational or more sophisticated, as will be explained later.

[0048] Prediction of samples within a CU using the decoded MVs is performed for one reference picture in step 1004 and for the other reference picture in case of bi-prediction mode in step 1005. The prediction signal is further refined for one reference picture in step 1006 and for the other reference picture in case of bi-prediction mode in step 1007. When applicable, bi-prediction or mixed intra / inter prediction or combination of PUs within a CU is achieved in step 1008. In step 1009, a final refinement step of the prediction signal is performed.

[0049] Figure 5 also shows the different modes applied to these different steps in VTM3. Tools marked with (*) correspond to tools that have been proposed for adoption but have not been adopted at the VVC Draft 3 stage.

[0050] In inter, two basic modes for deriving MVs are used: merge and AMVP. In both cases, an MV candidate list is derived. This derivation process is generally different for these two modes, and even for a given mode, transformations are applied depending on other settings applied to the CU (e.g., merge vs. AMVP, translational vs. affine model).

[0051] In merge mode, a predictor index is signaled to fetch the MV in the merge candidate list. In addition, in the case where skip mode is not applied, a sample prediction residual is signaled. In AMVP mode, for each reference picture (one reference picture in the case of uni-prediction, two in the case of bi-prediction), a reference frame index, a predictor index, and a motion vector difference (mvd) are signaled. In addition, a sample prediction residual is signaled.

[0052] Table 1 provides an overview of the content of the different candidate lists in the current VTM.

[0053] [Table 1]

[0054] A summary of the inter modes considered in this invention is provided in the table below, together with some indicative numbers regarding performance (in terms of psnrY, in random access configuration, in terms of bit rate variation in percentage).

[0055] [Table 2]

[0056] The following subsections provide further details on the inter modes considered, along with the syntax elements associated with these modes or tools.

[0057] The general aspects described herein focus on the "Mvd Derive" and "Model Derive" steps of Figure 5. A goal of this disclosure is to extend the use of the MMVD and SMVD motion vector coding tools to motion models beyond simple translational motion models.

[0058] [MMVD Motion Vector Difference Encoding Tool Description] MMVD is applied only in merge mode. It uses a merge candidate list. The flag "mmvd_skip" indicates whether MMVD mode is applied. When the mode is applied, the MV difference (mmvd) is constructed as follows: A syntax element (here denoted mmvd_idx) is signaled to construct a correction MV mmvd consisting of the following information: ○ The base MV index, selected by the encoder from among the two first translational merge candidates. ○ An index, denoted mmvd_dir_idx, associated with the direction D (currently four) in the (x,y) coordinate system (a table dir[] consisting of four elements {(0,1),(1,0),(-1,0),(0,-1)} is specified). an index, denoted mmvd_dist_idx, relating to the distance step S from the base MV (currently up to 8 distances are possible, with the specification of a table dist[] with 8 elements {1 / 4 pel, 1 / 2 pel, 1 pel, 2 pel, 4 pel, 8 pel, 16 pel, 32 pel}).

[0059] When the MMVD mode is applied, the MV difference is then calculated as follows: refinementMV=dir[mmvd_dir_idx]×dist[mmvd_dist_idx]

[0060] A single MV difference is signaled even if the CU is bi-predictively coded. In the bi-predictive case, two symmetric MV differences are obtained from a single coded MV. When the temporal distance between the predicted picture and the reference picture differs between reference picture lists L0 and L1, the decoded mmvd is assigned to the MV difference (mvd) associated with the largest temporal distance. The mvd associated with the smallest distance is scaled as a function of the POC distance.

[0061] For example, consider the case where reference picture k (k=0 or 1) is the closest one to the current picture. Let kk=1-k, and define POCref_0, POCref_1, and POCcur to be the picture order counts of reference picture 0, reference picture 1, and the current picture, respectively. The scaling factor is derived as follows: sc=(POCref_kk-POCcur) / (POCref_k-POCcur)

[0062] Then, the refined MV for each of the reference pictures is derived as follows: refinementMV_k=refinementMV refinementMV_kk=sc×refinementMV

[0063] In VVC Draft 3, MMVD is only applied to the translational motion model.

[0064] [Symmetric MVD proposed in JVET-L0370] A symmetric MVD (SMVD) tool is being considered for VVC. Its principle is to encode some motion vector information under the constraint that in the bi-prediction case, the motion information of a CU is made up of two symmetric forward and backward motion vector differences. In VVC Draft 3, the SMVD mode is only applicable to AMVP.

[0065] The encoding of a CU under this constraint, called SMVD mode, is signaled through the CU-level flag symmetrical_mvd_flag.

[0066] This flag is coded when SMVD mode is feasible, i.e., when the prediction mode of the CU is bi-predictive and two reference pictures for the CU are found, such as:

[0067] - The reference pictures for the current CU are searched as the nearest forward and backward reference pictures in the (L0 and L1) or (L1 and L0) reference picture lists, respectively. If not found, SMVD mode is not applicable and symmetrical_mvd_flag is omitted.

[0068] If symmetrical_mvd_flag is signaled and equals true, - For L0 reference pictures, one mvd is signaled, and the mvds for the other reference picture lists are derived as symmetrical ones, i.e. as the opposite of the first one. - As in conventional AMVP mode, two MV predictor indices (one per reference picture list) are signaled.

[0069] [Motion models supported in or proposed for VVC Draft 3] *Translational Model (VVC Draft 3) By default, the motion within a CU is based on a translational MV, which is applied to all samples in the block. *ATMVP (VVC Draft 3)

[0070] In the alternative (advanced) temporal motion vector prediction (ATMVP) scheme illustrated by FIG. 8, one or more temporal motion vector predictors for the current CU are retrieved from a reference picture for the current CU.

[0071] First, the motion data associated with the first candidate in the regular merge candidate list of the current CU, the so-called temporal motion vector, and the associated reference picture index, are obtained.

[0072] Next, the current CU is divided into NxN sub-CUs, where N is generally equal to 4. This is shown in Figure 8. For each NxN sub-block (sub-CU), a motion vector and a reference picture index are identified in the reference picture associated with the temporal MV with the help of a temporal motion vector. From the position of the current sub-CU, the NxN sub-block in the reference picture pointed to by the temporal MV is considered. Its motion data is obtained as the ATMVP motion data prediction for the current sub-CU. It is then converted into the motion vector and reference picture index of the current sub-CU through appropriate motion vector scaling.

[0073] Note that in VVC Draft 3, the ATMVP motion information predictor is part of the sub-block based merge candidate list.

[0074] *Affine motion model (VVC Draft 3) One of the new motion models introduced in VVC is the affine mode, which basically consists in using affine motion fields to represent motion vectors within a CU.

[0075] The motion model used for two or three control point (hereinafter also referred to as control points) motion vectors is illustrated by Figure 9. The four-parameter affine motion field, based on two control points, for each position (x, y) in the considered block is the following motion vector component values:

[0076]

number

[0077] where (v 0x ,v 0y ) and (v 1x ,v 1y ) are the so-called control point motion vectors (CPMVs) used to generate the affine motion field. (v 0x ,v 0y ) is the control point motion vector of the upper left corner. (v 1x ,v 1y ) is the control point motion vector of the top right corner. A six-parameter affine field generated from three CPMVs is also possible, as described in JVET-J0021.

[0078] In practice, to keep complexity reasonable, affine motion is managed on a 4x4 sub-block basis, i.e., the same motion vector is used for each sample within each 4x4 sub-block (sub-CU) of the considered CU (see Figure 10). Affine motion vectors are calculated from the CPMV at the center position of each sub-block. The obtained MV is expressed with 1 / 16 pel precision.

[0079] As a result, the entire block (CU) is temporally predicted through motion compensation of each 4x4 sub-block (sub-CU) with its own MV. In VTM, affine motion compensation can be used in two ways: affine inter (AF_INTER) and affine merge (or merge affine), which are described below.

[0080] -Affine Inter (AF_INTER)- A CU in AMVP mode whose size is greater than 8x8 can be predicted in affine inter mode. This is signaled through a flag, inter_affine_flag, coded at the CU level. Generating an affine motion field for that inter CU involves determining a control point motion vector (CPMV), which is obtained by the decoder through the addition of a motion vector difference plus a control point motion vector prediction (CPMVP). A CPMVP is a pair (for a four-parameter affine model with two CPMVs) or a triplet (for a six-parameter affine model with three CPMVs) of motion vector candidates, which can be inherited from affine neighbors (as in affine merge mode) or constructed from non-affine motion vectors obtained from lists (A, B, C), (D, E), and / or (F, G), respectively, as illustrated in Figure 11. This corresponds to the "virtual candidate" mentioned in Table 1.

[0081] -Affine Merge- In affine merge mode, a CU-level flag indicates whether the CU in merge mode uses affine motion compensation. If so, JEM (search reference software developed by JVET before VVC) selects the first available neighboring CU coded in affine mode from the ordered list of candidate positions (A, B, C, D, E) in Figure 12.

[0082] Once the first neighboring CU in affine mode is obtained, three motion vectors from the top left corner, top right corner, and bottom left corner of the neighboring CU are obtained.

[0083]

number

[0084] is retrieved (see FIG. 13). Based on these three vectors, two or three CPMVs of the top-left corner, top-right corner, and / or bottom-left corner of the current CU are derived as follows:

[0085]

number

[0086] Control point motion vector of the current CU

[0087]

number

[0088] When,is obtained, the motion field within the current CU is calculated on a 4x4 sub-CU basis through the model in Eq.

[0089] In parallel applications, more candidates for the affine merge mode are considered, in which case the best candidate is selected at the encoder through a rate-distortion optimization process, and the index of this best candidate is coded into the bitstream through the merge_idx syntax element.

[0090] The next affine candidate placed in the affine merge candidate list is a "constructed" (or "virtual") affine model candidate, as opposed to an inherited candidate. A constructed candidate is an affine motion field that is calculated based on available neighboring motion vectors around the current CU, including those from neighboring CUs that are not coded in affine mode. Two or three neighboring MVs are derived and used to generate a candidate affine motion field for the current CU.

[0091] The neighborhood MVs used to construct the affine merge candidates include several spatial MVPs and one temporal MVP.

[0092] [Planar motion model (proposed in JVET-L0070)] In JVET-L0070, planar motion vector prediction (PMVP) is proposed as an additional merge mode for VVC codec design. Planar motion vector prediction, shown in Figure 14, is achieved by averaging horizontal and vertical linear interpolations on a 4x4 block basis as follows: These linear interpolations are performed from a generator MV located on the corners (AR, AL, BL) of a CU. Its principle is illustrated in Figure 14. P(x,y)=(H×P h (x,y)+W×P v (x,y)+H×W) / (2×H×W)

[0093] [Regression MVF (RMVF) model] In JVET-L0171, regression-based motion vector prediction is proposed as an additional merge mode for VVC codec design.

[0094] The regression-based motion vector field (RMVF) shown in Figure 15 consists in the following: The principle of RMVF is to use a six-parameter motion model to calculate the motion vectors of the sub-blocks.

[0095]

number

[0096] The motion parameters are calculated based on a row and a column of spatially neighboring 4x4 sub-blocks, using their motion vectors and center positions as inputs for a linear regression method.

[0097] [Triangular partition-based motion model] In VVC Draft 3, a triangular motion partitioning tool was adopted, which allows flexibility in dividing a block into two prediction units, as shown by Figure 16. Essentially, a CU is divided into two triangular prediction units in a diagonal or anti-diagonal direction. Each triangular PU is inter-predicted using its own uni-predictive motion vector and reference picture derived from a dedicated uni-predictive candidate list.

[0098] After predicting the triangular prediction unit, an adaptive weighting process (see FIG. 17) is performed on the diagonal edge. Then, the transform and quantization process is applied to the entire CU. The triangular PU mode is only applicable to the skip mode and merge mode.

[0099] Local Lighting Compensation (LIC) The purpose of LIC is to compensate for illumination changes that may occur between a prediction block and its reference block, which are exploited through motion-compensated (MC) temporal prediction. In this tool, the decoder calculates some prediction parameters based on some reconstructed picture samples located in the left columns and / or upper rows of the current block and some reference picture samples located in the left columns and / or upper rows of the motion-compensated block (Figure 18a).

[0100] In another approach, the predicted samples of the left / upper neighboring reconstructed block (Fig. 18b) can be used.

[0101] In the bi-predictive case, a variant of LIC ("bi-directional LIC") consists in estimating the illumination change between the two current reference blocks in order to derive LIC parameters for the current block.

[0102] The LIC parameters are selected to minimize the mean squared error (MSE) between the samples in Vcur and the corrected samples in Vref(MV). In general, the LIC model is linear, i.e., LIC(x) = a.x + b.

[0103]

number

[0104] s and r correspond to pixel positions in Vcur and Vref(MV), respectively, as shown in FIG.

[0105] As a result, when LIC is used for temporal prediction, a linear LIC model is applied to the motion compensated block to provide the temporal prediction of the current block.

[0106] Generalized Bi-Prediction (GBI) In VVC Draft 3, GBI was adopted, which applies unequal weights to predictors from L0 and L1 in bi-prediction mode. In inter-prediction mode, multiple weight pairs, including equal weight pairs (1 / 2, 1 / 2), are evaluated based on rate-distortion optimization (RDO), and the GBI index of the selected weight pair is signaled to the decoder.

[0107] In AMVP mode, the GBI index, which carries the GBI weight information, is signaled at the CU level.

[0108] In merge mode, the GBI index is inherited from the neighboring CU. In GBI mode, the predicted block is calculated as follows: P GBi =(w0×P L0 +w1×P L1 ) where w0 and w1 are the selected GBI weights. Supported w1 values ​​are generally {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}. Since the sum of w1 and w0 equals 1, the corresponding w0 values ​​are {5 / 4, 5 / 8, 1 / 2, 3 / 8, -1 / 4}. Weight pairs are selected and signaled at the CU level. For non-low latency pictures, the number of weights is reduced. The w1 and w0 values ​​are {3 / 8, 1 / 2, 5 / 8} and {5 / 8, 1 / 2, 3 / 8}, respectively.

[0109] 1.1.1. Non-Subblock Spatial-Temporal Merging Motion Vector Predictor (STMVP) This section describes a prior art method JVET-L0354 for generating spatiotemporal merge candidates, called STMVP. The method basically consists in retrieving two spatial neighboring motion vectors above and to the left of the current PU and a temporal motion vector predictor for the current CU.

[0110] The spatial neighborhood is obtained at a spatial location called Afar, as illustrated in Figure 20. The spatial location of Afar relative to the top-left location of the current PU is given by the coordinates (nbPW × 2, -1), where nbPW is the width of the current block. If a motion vector is not available at location Afar (it does not exist or is not intra-coded), location B1 is considered as the upper neighboring motion vector of the current block.

[0111] The selection of the left neighboring block is similar: if available, the neighboring motion vector at the relative spatial position (-1, 2 × nbPH), denoted Lfar, is selected, where nbPH is the height of the current block. Otherwise, the left neighboring motion vector at position A1 (-1, nbPH-1) is selected, if available.

[0112] Then, the TMVP predictor for the current block is derived in the same way as in HEVC temporal motion vector prediction.

[0113] Finally, the STMVP merge candidate for the current block is calculated as the average of up to three obtained spatial and temporal neighboring motion vectors. Thus, in contrast to the sub-block-based approach of STMVP in JEM, the STMVP candidate here is generated from at most one motion vector per reference picture list. Several variations on the approach of JVET-L0354 have been proposed, for example, in JVET-L0207.

[0114] Proposed embodiments for extending the use of MMVD Combining MMVD and bi-prediction In the current MMVD design, a single set of mmvd data, made up of direction and distance indices, is signaled even in the bi-predictive case.

[0115] [Embodiment 1] - Interdependent signaling of mmvd data for both reference pictures In an embodiment, when a CU is coded in bipredictive (bipred) mode and MMVD mode is enabled for the CU, two sets of mmvd data are signaled, one set for each one of the two mmvds applied to each one of the reference pictures (from lists L0 and L1). In addition, the second set of mmvd data can be derived from the first set of mmvd data, with various possible options.

[0116] This is illustrated in the simplified block diagram of Figure 21. In step 1101, mmvd data is decoded for a first reference picture. If bi-prediction is applied (check in step 1102), mmvd data is decoded for a second reference picture conditional on the mmvd data of the first reference picture (step 1103). Then, in step 1104, a bi-predictive MV is derived from the mmvd data of the first and second reference pictures. If bi-prediction is not applied (check in step 1102), a uni-predictive MV is derived from the mmvd data of the first reference picture in step 1105.

[0117] [Embodiment 1] - Option 1 - Estimated direction for mmvd of second reference picture In one option, only a distance index is coded for both mmvds (mmvd0_dist_idx and mmvd1_dist_idx). The direction of the second mmvd is inferred from the direction of the first mmvd (mmvd0_dir_idx). If both reference pictures are located on the same temporal side with respect to the current picture, mmvd1_dir_idx is set equal to mmvd0_dir_idx. If reference pictures are located on both temporal sides from the current picture, mmvd1_dir_idx is set equal to -mmvd0_dir_idx.

[0118] An example of the relevant simplified syntax is shown below:

[0119] [Table 3]

[0120] [Embodiment 1] - Option 2 - Differential encoding of the mmvd distance of the second reference picture In another option, the distance for the second reference picture is coded as a difference relative to the distance for the first reference picture. An index mmvdd1_dist_idx, which corresponds to the difference relative to mmvd0_dist_idx, is coded.

[0121] An example of the relevant simplified syntax is shown below:

[0122] [Table 4]

[0123] The distance for refined MV0 is distMV0=dist[mmvd0_dist_idx] and the distance for refined MV1 is calculated as distMV1=distMV0+dist[mmvdd1_dist_idx] It is calculated as:

[0124] [Embodiment 1] - Option 3 - Restricting the possible distance values ​​of the mmvd of the second reference picture In an embodiment, the maximum value of the second distance (either the distance signaled by mmvd1_dist_idx or the distance difference as in option 2 signaled by mmvdd1_dist_idx) is made smaller compared to that of mmvd0_dist_idx. For mmvd0_dist_idx, consider that N possible distances are used (e.g., 8, dist[]={1 / 4 pel, 1 / 2 pel, 1 pel, 2 pel, 4 pel, 8 pel, 16 pel, 32 pel}). For mmvdd1_dist_idx, only the first N' values ​​of dist[] can be used, and N' <Nである。

[0125] The following variations are possible: - N' is calculated from N (e.g., N'=N / 2). - N' is calculated from the value of mmvd0_dist_idx (e.g., N'=mmvd0_dist_idx, or N'=mmvd0_dist_idx / 2). - N' depends on the value of the scale parameter "sc" (as explained in the section on MMVD encoding tool description), e.g., N' = sc × N

[0126] [MMVD-CPR combination] In an embodiment, when CPR (Current Picture Referred) mode is applied, MMVD mode is enabled with CPR-specific adaptations.

[0127] The MV used for CPR references the current picture. In VVC Draft 3, a constraint was specified to restrict the MV to points within a constrained region close to the current CU in order to limit memory storage needs. In the current VTM, the constrained region consists of the CTU containing the current CU, as illustrated in Figure 22. Another constraint of the CPR MV is that it has integer precision.

[0128] [Embodiment 2a] - mmvd constraints on integer precision for CPR In an embodiment, the enabled distance for the mmvd is limited to match the MV accuracy enabled for CPR. For example, in current VTMs, the CPR MV accuracy is an integer, and the present invention limits the distance for the mmvd when the CPR mode is used to the full set normally used for the mmvd. dist[]={1 / 4 pel, 1 / 2 pel, 1 pel, 2 pel, 4 pel, 8 pel, 16 pel, 32 pel} Instead of the following set: {1 per, 2 per, 4 per, 8 per, 16 per, 32 per} Consider it to be within.

[0129] In the case of CPR, the maximum value N' (for example, 5) of mmvd_dist_idx is smaller than the maximum value N (for example, 8) of mmvd_dist_idx.

[0130] In an embodiment, this solution is implemented using an offset applied to mmvd_dist_idx as follows: - If MMVD mode is applied, ○When CPR mode is applied, □distMV=dist[mmvd_dist_idx+offset].

[0131] ○Otherwise, □distMV=dist[mmvd_dist_idx].

[0132] To limit to integer precision, the offset is equal to 2.

[0133] [Embodiment 2b] - Limiting the maximum value of mmvd in the case of CPR In an embodiment, the maximum distance for the mmvd in the case where the CPR mode is activated is reduced compared to conventional mmvd use. For example, in current VTMs, the CPR MV precision is an integer, and the present invention reduces the distance for the mmvd when the CPR mode is activated to the full set normally used for the mmvd. dist[]={1 / 4 pel, 1 / 2 pel, 1 pel, 2 pel, 4 pel, 8 pel, 16 pel, 32 pel} Instead of the following set: {1 per, 2 per, 4 per, 8 per} Consider it to be within.

[0134] In the case of CPR, the maximum value N' (for example, 3) of mmvd_dist_idx is smaller than the maximum value N (for example, 8) of mmvd_dist_idx.

[0135] [Embodiment 2c] - Clipping of MV from mmvd in case of CPR In an embodiment, the MVs resulting from the mmvd refinement are clipped so that the motion compensation blocks used for prediction remain within the constrained region. The basic process is illustrated in Figure 23.

[0136] Embodiment 3: Combination of MMVD and ATMVP [Embodiment 3.1] - MMVD used in combination with sub-block-based mode ATMVP In VVC Draft 3, MMVD only applies to the normal translational merge mode, so when sub-block-based merge mode is activated for a given CU, MMVD is not used.

[0137] According to this embodiment 3, when the sub-block based merge mode is activated, the MMVD motion vector representation tool can be used. In particular, the MMVD can be used in combination with the ATMVP motion vector prediction mode.

[0138] When the merge_flag syntax indicates the use of merge mode for the current CU, the use of MMVD is signaled in the same way as in VVC Draft 3. Furthermore, the flag subblock_merge_flag, which indicates the use of sub-block-based merge mode, is coded regardless of the value of the syntax element mmvd_merge_flag, which indicates the use of MMVD for the current block.

[0139] According to the first subembodiment 3.0, in combination with MMVD, affine mode is not allowed. Therefore, if mmvd_merge_flag is on, the merge mode used is inferred to be ATMVP mode. Therefore, the merge_idx syntax element is omitted from the coded bitstream.

[0140] According to another sub-embodiment, in combination with MMVD, affine mode is allowed, and therefore the merge_idx syntax element is coded as it is currently done in VVC draft 3.

[0141] [Embodiment 3.2] - mmvd used as a refinement of inherited sub-block MV of ATMVP In an embodiment, when an ATMVP candidate is selected and MMVD is enabled, mmvd refinement is applied to all MVs inherited for each sub-block. The process is shown in the simplified block diagram of Figure 24. In the first step, mmv data is decoded. In the next step, a global ATMVP MV is identified. Then, for each sub-block of the CU, motion data from the corresponding sub-block identified in the reference picture by the global MV is fetched. These motion data are refined using mmvd data, similar to the conventional MMVD refinement process.

[0142] [Embodiment 4]: Combination of MMVD and affine mode According to an embodiment, MMVD is used in combination with affine modes. Within this embodiment, several variations can be applied, which are listed below.

[0143] [Single MMVD coded for all CPMVs] According to a first variant, a single motion vector difference mmvd is applied to all CPMVs that are coded and used to generate the affine motion field. This results in adding a constant motion vector to the entire affine motion model. This is illustrated in the figure below, which shows a simplified decoding block diagram of the affine motion field generation process when the MMVD mode is activated.

[0144] In the case of bidirectional affine prediction, the same symmetric motion vector difference concept is used as in the translational case.

[0145] [mmvd data encoded for each CPMV] According to a second variant, one motion vector difference mmvd per CPMV is allowed, which increases the flexibility for generating candidate affine motion fields. For the first CPMV, the normal mmvd encoding of the MMVD can be used. Then, for the second and third CPMV, a differential motion vector difference (mmvdd) can be encoded on top of the mmvd associated with the first CPMV. The advantage of such an approach is to limit the rate cost of mmvd signaling while allowing some flexibility in the application of MMVD to affine modes. An example of the relevant simplified syntax in the case of encoding mmvd information for three affine CPMVs is shown below. According to this embodiment, the magnitude of the differential motion vector difference can be constrained to a value smaller than the mmvd of the first CPMV (e.g., the full set of values ​​can be dist[]={1 / 4 pel, 1 / 2 pel, 1 pel, 2 pels, 4 pels, 8 pels, 16 pels, 32 pels}, but limited to dist[]={1 / 4 pel, 1 / 2 pel, 1 pel}). The process is illustrated in the following figure, which shows a simplified decoding block diagram of the affine motion field generation process when MMVD mode is activated.

[0146] [Table 5]

[0147] According to some variants, in the coding of the differential mmvd, the allowed magnitude range is limited according to the distance index associated with the mmvd of the first CPMV. For example, mmvd1_dist_idx and, when applicable, mmvd2_dist_idx are constrained to be lower than mmvd0_dist_idx. The allowed range of the second mmvd in the horizontal and vertical directions can be limited as exemplarily shown in Figure 27.

[0148] According to a further variant, the same direction index is used for all mmvds, and therefore for the second and third CPMVs only distance information can be coded. An example of the relevant simplified syntax in the case of coding mmvd information for three affine CPMVs is shown below:

[0149] [Table 6]

[0150] [Constraining the affine merge candidate list by removing hypothetical candidates] - According to a further modification, since the MMVD mode used for affine introduces some diversity in the set of affine merge candidates, the affine list of affine merge candidates is reduced compared to the affine merge list of VVC Draft 3, leading to a simplified overall codec design. For example, in an embodiment, some constructed (virtual) affine model candidates are removed from the affine merge list. In an embodiment, all constructed (virtual) affine model candidates are removed from the affine merge list, and the affine merge list is made up of only inherited affine candidates.

[0151] [Constraining the distance value of the second (and third) mmvd based on the distance value of the first mmvd] - According to an additional property, for the second CPMV and (when it applies) the third CPMV, the differential encoding of the mmvd (motion vector difference) relative to the mmvd of the first CPMV includes the possibility to encode / decode a distance value of 0 for the differential motion vector difference (dist[]={0 pel, 1 / 4 pel, ...}). In fact, in the existing MMVD, which applies only in the translational motion case, a distance equal to 0 is not supported, since it provides an MV predictor that overlaps with the one of the normal merge.

[0152] [Imposing symmetry constraints on affine model parameters in the bidirectional affine case] According to a further characteristic, in the bidirectional affine case, the derivation of the second mmvd (associated with the second reference picture) as a function of the first one (associated with the first reference picture) uses a symmetry constraint similar to the symmetry constraint between two bi-predictive mmvds imposed in the translational case. To do so, several approaches may be possible. In one first approach, the mmvd of the mmvd associated with the second reference picture list is inferred directly from the mmvd associated with the first reference picture list, e.g., the second mmvd is derived by scaling the first mmvd, where the scaling takes into account the temporal distance between the reference picture and the current picture. In another approach, the MVD of the affine model associated with the second reference picture list is estimated from the first one by imposing symmetry on the affine model parameters (typically linked to angles and scaling factors) compared to the MVD of the affine model associated with the first reference picture. To do so, the scaling factor and angle values ​​associated with the first affine model are calculated. They are then transformed to satisfy the symmetry constraint, and the affine model parameters are then inverse transformed to provide a second affine model of the current CU for bidirectional affine prediction of the CU. This takes the following form:

[0153] For the block under consideration, it is assumed that a four-parameter model is used.

[0154]

number

[0155] Then the rotation angle and scaling parameters are obtained as follows:

[0156]

number

[0157] The angles and scaling factors are then transformed as follows:

[0158]

number

[0159] where k is a scaling factor that depends on the temporal distance between the current picture and its two reference pictures. Finally, the a' and b' affine model parameters for the second reference picture are easily calculated from a' and b'.

[0160] In the simplified version, only the scaling factor is changed, resulting in a simple scaling of the MVd.

[0161] [Embodiment 5]: Combination of MMVD and planar motion vector prediction According to embodiment 5, MMVD is used in combination with a planar motion vector prediction (PMVP) mode.

[0162] This can take one of the following forms:

[0163] [Single mmvd refining each of the MVs in the PMVP motion field] According to the first basic approach, in the case where PMVP mode is used for the current CU, the MMVD motion vector difference is coded. The mmvd is applied to the PMVP generated motion field, which is first generated from its generator MV. Thus, the mmvd is used as an additive offset for each MV of the PMVP motion field. This is illustrated in the following figure, which shows a simplified decoding block diagram of the PMVD motion generation process.

[0164] [A single mmvd used to refine at least one of the MVs used to generate the PMVD motion field] According to another variant, a motion vector difference is applied to at least one motion vector used to generate the planar motion field (generator MV). For example, a single MVD can be coded and applied to the AL (top left) motion vector in Figure 28 before generating the planar MV field. This is illustrated in the figure below, which shows a simplified decoding block diagram of the PMVD motion generation process. According to another variant, the MVD is coded and applied to the BR motion vectors before generating the planar MV field. According to another variant, the MVD is coded and applied to the AR motion vectors before generating the planar MV field. According to another variant, the MVD is coded and applied to the BL motion vectors before generating the planar MV field.

[0165] [Some mmvds used to refine some MVs, some used to generate PMVD motion fields] According to another embodiment, before generating the planar motion field, several MVDs are coded and applied to one or more of the AL, BL, AR and BR motion vectors. This is illustrated in the following figure, which shows a simplified decoding block diagram of the PMVD motion generation process. According to another embodiment, when several MVDs are coded and applied to multiple MVs among AL, BL, AR and BR motion vectors, the first coded MVD is coded as in MMVD, and the following ones are coded in a differential way based on the previously coded ones in order to limit the rate cost associated with the MVD coding process. - Following the transformation, the set of allowed distances for MVD in the planar case is changed compared to the MMVD tool currently used in VVC Draft 3. For example, a constrained range of allowed MVD distances is possible. According to the variant, the number of allowed MVD orientations is changed compared to the existing MMVD system in VVC Draft 3. For example, when MMVD is used in combination with planar MV prediction, an enhanced set of MVD angles can be supported.

[0166] [Embodiment 6]: Combining MMVD with regression-based motion vector fields According to embodiment 6, the MMVD is used in combination with a regression-based six-parameter motion field, which is introduced in the Regression MVF Model section. This can take one of the following forms:

[0167] [RMVF: A single mmvd that refines each of the MVs in the motion field] According to the first basic approach, the MMVD motion vector difference is coded in the case where the RMVF mode is used for the current CU and is applied to the RMVF-generated motion field. Thus, the MVd is used as an additive offset to the RMVF motion field. A similar block diagram as shown in Figure 29 can adequately illustrate the simplified process for generating an RMVF motion field according to this variant.

[0168] [A single mmvd used to refine at least one of the MVs used to generate the RMVF motion field] According to another variant, one motion vector difference is applied to at least one motion vector used to generate the RMVF motion field, for example, a single MVD can be coded and applied to the motion vector to the left of the current block before generating the RMVF motion field. According to another variant, one MVD is coded and applied to the upper neighboring motion vectors before generating the regression-based MV field. A similar block diagram as shown in Figure 30 can adequately illustrate the simplified process for generating an RMVF motion field according to this variant.

[0169] [2 mmvds used to refine 2 MVs used to generate RMVF motion fields] According to another embodiment, before generating the regression-based motion field, two MVDs are coded and applied to the top and left MVs respectively. According to another embodiment, when two MVDs are coded and applied as described above, the first coded MVD is coded as in MMVD, and the next one is coded in a differential way based on the first one in order to limit the rate cost associated with the MVD coding process. According to the variant, the set of allowed distances for MVD in the regression base case is changed compared to the MMVD tool currently used in VVC Draft 3. For example, a constrained range of allowed MVD distances is possible. - Following the transformation, the number of allowed MVD orientations is changed compared to the existing MMVD system in VVC Draft 3. For example, when MMVD is used in combination with RMVF, an enhanced set of MVD angles can be supported. A similar block diagram as shown in FIG. 31 can adequately describe the simplified process for generating RMVF motion fields according to these transformations.

[0170] [Embodiment 7]: Combination of MMVD with LIC In an embodiment, MMVD and LIC can be enabled together, in which case, as in the case of VVC Draft 3, the coded motion vector difference is applied to the motion vector of the considered CU. Thus, as in the case of the current LIC tool, the LIC linear model parameters associated with the temporal prediction of the current CU are derived from the merge candidate used.

[0171] According to a variant, if the magnitude of the MVD exceeds a certain threshold, the LIC flag of the current CU can be set to false, i.e., LIC temporal prediction refinement can be deactivated for the current CU. Indeed, intuitively, if the motion vectors of the current CU are significantly different from the motion vectors of the merge candidate used to derive the MV of the current CU, the LIC linear model used for the merge candidate CU can be irrelevant to the current CU.

[0172] For example, if the MVD size is 16 or 32 or greater, the LIC mode can be forced to 0 for the current CU.

[0173] According to a further variant, the above mentioned threshold for the magnitude of the MVD can depend on the current CU size.

[0174] [Embodiment 8]: Combination of MMVD and GBI In VVC Draft 3, when a CU is coded in translational merge mode, its motion vector is derived from the selected merge candidate and from its GBI index. This means that the GBI weight of the CU, which serves as a reference for deriving the current CU motion data, is still used for the current CU, including the case where MMVD is activated for the current CU.

[0175] In this current embodiment, an adaptation of the GBI weight for the current CU can be applied based on the magnitude of the MVD used for the current merged CU.

[0176] For example, if the magnitude of the MVD is above a certain threshold, the GBI can be reset to the default GBI weight (1 / 2, 1 / 2) for the current CU. Indeed, intuitively, if the motion vector of the current CU is significantly different from the motion vector of the merge candidate used to derive the MV of the current CU, the GBI weight used for the merge candidate CU can be irrelevant to the current CU.

[0177] For example, if the MVD magnitude is 16 or 32 or greater, the GBI weights can be forced to (1 / 2, 1 / 2) for the current CU.

[0178] According to a further variant, the above mentioned threshold for the magnitude of the MVD can depend on the current CU size.

[0179] [Embodiment 9]: MMVD and Triangular Motion Partition According to an embodiment, the MMVD motion vector coding tool is used in combination with a triangular motion partitioning tool.

[0180] According to a first variant, for each triangular partition, one MVD is coded according to the VVC Draft 3 MVD coding system of the MMVD.

[0181] According to another variant, a single MVD is coded and used in common by two triangular motion partitions.

[0182] According to another variant, when two MVDs are coded, the second one is coded in a differential way relative to the first one, in which case the second MVD can be constrained to a smaller allowed magnitude range in the same way as for the affine case (see section on embodiment 4).

[0183] According to a more advanced embodiment, in the case where one MVD is coded for each partition, the second MVD is coded in a conditional manner based on the relative values ​​of the motion vectors of the first and second partitions. For example, the second refined MVD can be constrained so that it is not too close to the refined MV of the first triangular partition. In fact, if it is too close, the overall prediction of the current CU may behave very similarly to the normal translational motion compensated prediction of the entire rectangular block.

[0184] [Embodiment 10]: MMVD and multiple hypothesis (MH) prediction In current video standards, such as VVC Draft 3, a new prediction mode called multiple hypothesis consists in the combined prediction of merge / skip temporal predicted blocks and intra predicted blocks. However, in VVC Draft 3, MMVD and MH prediction cannot be used together.

[0185] In this embodiment, it is possible to use MMVD and multiple hypothesis prediction modes in a combined manner, which essentially consists in applying a merge motion vector difference to the merge or skip candidate motion vectors used to derive the motion information of the current coding unit.

[0186] The advantage of this embodiment is improved coding efficiency.

[0187] [Embodiment 11]: MMVD and spatio-temporal motion vector prediction (STMVP) In VVC Draft 3, a motion vector prediction mode called STMVP is proposed, as described in the section on non-subblock spatiotemporal merge motion vector predictors. STMVP motion vector candidates are generally additional merge candidates that can be part of a translational merge candidate list. It essentially predicts a single motion vector for the current CU. Therefore, MMVD can be applied to STMVP candidates in a straightforward manner.

[0188] However, MMVD can also be applied to only one of the three spatial and temporal MV predictors used to compute the STMVP candidate.

[0189] According to an embodiment, the motion vector difference of the MMVD is applied to one or two spatial motion vector predictors before calculating the average motion vector between these two spatial motion vector predictors and the temporal motion vector predictor.

[0190] According to an embodiment, the motion vector difference of the MMVD is applied to the temporal motion vector predictor before calculating the average motion vector between this temporal MV predictor and one or two spatial motion vector predictors.

[0191] Suggested embodiments for extending the use of SMVD Embodiment 12: SMVD replaced by MMVD in AMVP mode As mentioned in the Symmetric MVC section, the SMVD motion vector coding mode is only applied in AMVP mode, while the MMVD motion vector representation mode is only applied in merge mode. In this embodiment, the codec design is harmonized. Only one motion vector coding mode, MMVD and SMVD, is proposed for the overall design to handle both the symmetric bidirectional motion case and the small magnitude motion vector case in the context of merging. The following two variants are proposed:

[0192] [Use of mmvd syntax to encode MV differences in both MMVD and SMVD modes] In AMVP mode, when symmetric mode is on, the motion vector difference is coded using the MMVD MVD coding syntax shown in the MMVD Motion Vector Difference Coding Tool Description section. Thus, if symmetric MVD mode is active for the considered CU, the conventional MVD coding method of VVC Draft 3 in AMVP is replaced.

[0193] [Use of VTM3 MVd syntax to encode MV differences in both MMVD and SMVD modes] In AMVP mode, when symmetric mode is on, motion vector differences are coded as is currently done in VVC Draft 3 in AMVP. However, in the merge case, when MMVD mode is on, MMVD motion vector differences are coded in the AMVP way.

[0194] [Embodiment 13]: SMVD combined with affine motion model According to an embodiment, the use of SMVD is extended to the affine AMVP case, which can take the form:

[0195] The first CPMV of the considered affine CU is coded according to the SMVD mode of the section on symmetric MVD, just as in the conventional translational AMVP case. Then, the MVDs of other CPMVs of the considered affine CU are differentially coded with respect to the MVD of the first CPMV, just as is currently done in the affine AMVP case. A corresponding simplified block diagram is shown in the following figure.

[0196] According to a further variant, the differential MVDs of the second CPMV and optionally the third CPMV are differentially coded with respect to the MVD of the first CPMV, but also coded under a symmetric mode constraint, which allows to reduce the rate cost of the second and third CPMVs.

[0197] [Embodiment 14]: SMVD combined with triangular motion partitioning In AMVP, when triangular partitions are used, the symmetric MVD mode can be extended to the triangular partition case.

[0198] In such a case, in an embodiment, the two unidirectional MVDs of the first and second triangular partitions may be symmetric with respect to each other.

[0199] [Embodiment 15]: SMVD combined with multi-hypothesis prediction mode In an embodiment, in addition to merge mode, multiple hypotheses can be used in AMVP mode, in which case SMVD mode can be used in combination with multiple hypothesis prediction mode in cases where the inter component of this inter / intra hybrid prediction uses bi-prediction.

[0200] [Embodiment 16]: SMVD combined with planar motion model In an embodiment, the planar motion model can be used in AMVP mode in addition to its current use in merge mode, in which case motion differences can be used and applied to motion vectors surrounding the current CU and used to generate a motion field for the current CU.

[0201] Furthermore, SMVD mode can be used in combination with this AMVP planar motion model case, in which some symmetric bidirectional motion vector differences can be applied to the motion vectors around the current CU that are used to generate the planar MV field of the current CU. Such an embodiment is similar to that in the section covering embodiment 5, but in the AMVP+symmetric mode case.

[0202] [Embodiment 17]: SMVD combined with regression-based motion model In an embodiment, a regression-based motion model can be used in AMVP mode in addition to its current use in merge mode, in which case motion differentials can be used and applied to motion vectors surrounding the current CU and used to generate a regression-based motion field for the current CU.

[0203] Furthermore, SMVD mode can be used in combination with this AMVP regression-based motion model case, in which several symmetric bidirectional motion vector differences can be applied to the motion vectors around the current CU that are used to generate the regression-based MV field for the current CU. Such an embodiment is similar to that in the section covering embodiment 5, but in the AMVP+symmetric mode case.

[0204] [Embodiment 18]: SMVD combined with ATMVP motion model Exclusive due to AMVP and merging In an embodiment, the ATMVP motion model can be used in AMVP mode in addition to its current use in merge mode, in which case motion differentials can be used and applied to the motion vectors contained in the ATMVP predicted motion vectors of the current CU.

[0205] Furthermore, SMVD mode can be used in combination with this AMVP ATMVP motion model case, in which some symmetric bidirectional motion vector differentials can be applied to the ATMVP motion vector derived for the current CU. Such an embodiment is similar to that in the section covering embodiment 3, but in the AMVP+ATMVP mode case.

[0206] [Embodiment 19]: SMVD combined with spatio-temporal motion vector prediction (STMVP) In current video standard proposals, such as VVC Draft 3, the STMVP motion vector predictor (see the section on STMVP) is proposed only in the merge case. However, in AMVP mode, SMVD can be enabled. In that case, SMVD can be applied to STMVP candidates in a straightforward manner.

[0207] However, SMVD can also be applied to only one of the three spatial and temporal MV predictors used to compute the STMVP candidate.

[0208] According to an embodiment, the SMVD motion vector difference is applied to one or two spatial motion vector predictors before calculating the average motion vector between these two spatial motion vector predictors and the temporal motion vector predictor.

[0209] According to an embodiment, the SMVD motion vector difference is applied to the temporal motion vector predictor before calculating the average motion vector between this temporal MV predictor and one or two spatial motion vector predictors. 1.1.2. [Embodiment 20]: Modified SMVD mode to impose symmetry on the final pair of bidirectional motion vectors. Translational case This section presents an embodiment in which the SMVD mode is modified so that symmetry constraints are imposed not on the motion vector difference MVd but on the final motion vector used to bidirectionally predict a block in translational AMVP mode. Indeed, as described in the section covering symmetric MVD, in SMVD, for the first motion vector of a bidirectional MV pair, an MVD is coded, and then the MVD of the second motion vector is inferred from the first MVD as a scalable opposite MVD vector as a function of the temporal distance between the current picture and its reference picture. Here, in the proposed embodiment, the modified SMVD mode is such that the MVD of the first vector is coded in the same way as in existing SMVD. Next, the first motion vector is reconstructed as the sum of the MV predictor and the decoded motion vector difference. Finally, the reconstructed second motion vector is inferred from the reconstructed first motion vector as a scalable opposite MVD as a function of the temporal distance between the current picture and its reference picture.

[0210] In this way, the proposed new SMVD mode ensures that a predicted block is on the same line as its two reference blocks, one behind and one ahead of it. This approach is expected to provide improved coding efficiency compared to the existing SMVD in the section covering symmetric MVD.

[0211] [Embodiment 21]: Modified SMVD mode to impose symmetries on the final affine model (rotation, scaling / zoom, and translation) According to an embodiment, the SMVD mode is used in combination with affine AMVP, in which case symmetry constraints can also be imposed on the affine MVD in a straightforward manner.

[0212] According to another approach, symmetry constraints can be imposed on the affine model parameters (angles and scaling factors) in a similar way to the corresponding properties in the section on embodiment 4 relating to the combination of MMVD and affine.

[0213] One embodiment of a method 3500 under the general aspects described herein is shown in Figure 35. The method begins at start block 3501, with control passing to block 3510 for indicating a first motion mode through syntax in the video bitstream. Control passes from block 3510 to block 3520 for indicating the use of a second motion mode through the presence of syntax in the video bitstream, and including information related to the second motion mode, if present. Control passes from block 3520 to block 3530 for encoding the video block using motion information corresponding to the first and second motion modes.

[0214] Another embodiment of a method 3600 under the general aspects described herein is shown in Figure 36. The method begins at start block 3601, with control passing to block 3610 for parsing a video bitstream for syntax indicating a first motion mode. From block 3610, control passes to block 3620 for parsing the video bitstream for syntax indicating the presence of a second motion mode and, if present, determining information related to the second motion mode. From block 3620, control passes to block 3630 for obtaining motion information corresponding to the first motion mode. From block 3630, control passes to block 3640 for decoding a block using the motion information.

[0215] 37 shows one embodiment of an apparatus 3700 for encoding, decoding, compressing, or decompressing video data using coding mode simplification based on a neighboring sample-dependent parametric model. The apparatus comprises a processor 3710 and may be interconnected through at least one port to a memory 3720. Both the processor 3710 and the memory 3720 may also have one or more additional interconnections to external connections.

[0216] The processor 3710 is also configured to insert information into or receive information in a bitstream, and to perform compression, encoding, or decoding using any of the described aspects.

[0217] This application describes various aspects, including tools, features, embodiments, models, techniques, etc. Many of these aspects are described with specificity, often in a manner that may sound limiting, at least to indicate their individual characteristics. However, this is for clarity of description and does not limit the applicability or scope of the aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Furthermore, aspects can also be combined and interchanged with aspects described in previous applications.

[0218] The aspects described and contemplated in this application can be implemented in many different forms. While Figures 3, 4, and 34 provide some embodiments, other embodiments are contemplated, and the descriptions of Figures 3, 4, and 34 do not limit the breadth of the embodiments. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.

[0219] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image," "picture," and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, while the term "decoded" is used on the decoder side.

[0220] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions can be varied or combined.

[0221] Various methods and other aspects described in this application can be used to modify modules, e.g., intra-prediction, entropy encoding, and / or decoding modules (160, 360, 145, 330), of video encoder 100 and decoder 200, as shown in Figures 3 and 4. Furthermore, the aspects are not limited to VVC or HEVC, but can be applied, for example, to other standards and recommendations, whether existing or developed in the future, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.

[0222] In this application, various numerical values ​​are used, and the particular values ​​are for illustrative purposes only and the described aspects are not limited to these particular values.

[0223] 3 illustrates an encoder 100. Although variations of this encoder 100 are contemplated, the encoder 100 is described below for clarity without describing all possible variations.

[0224] Before being encoded, the video sequence may pass through a pre-encoding process (101) that, for example, applies a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or performs a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.

[0225] In encoder 100, a picture is encoded by encoder elements as described below. The encoded picture is partitioned (102) and processed, for example, in units of CUs. Each unit is encoded, for example, using either intra mode or inter mode. When a unit is encoded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to encode the unit and indicates the intra / inter decision, for example, by a prediction mode flag. A prediction residual is calculated, for example, by subtracting (110) the prediction block from the original image block.

[0226] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process.

[0227] The encoder decodes the encoded blocks to provide references for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (155) to reconstruct an image block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / sample adaptive offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).

[0228] Figure 4 illustrates a block diagram of a video decoder 200. In the decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding pass that is reciprocal to the encoding pass as described in Figure 3. The encoder 100 also generally performs video decoding as part of encoding the video data.

[0229] In particular, the decoder's input includes a video bitstream, which may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned. Thus, the decoder can divide (235) the picture according to the decoded picture's partitioning information. The transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. An image block is reconstructed by combining (255) the decoded prediction residual with a prediction block. The prediction block can be obtained (270) from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).

[0230] The decoded picture may further pass through a post-decoding process (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.

[0231] FIG. 34 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 1000 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected home appliances, and servers. The elements of system 1000 may be embodied, singly or in combination, in a single integrated circuit (IC), multiple ICs, and / or individual components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or individual components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to perform one or more of the aspects described herein.

[0232] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein, for example, to implement various aspects described herein. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a nonvolatile memory device). The system 1000 includes a storage device 1040, which may include nonvolatile and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. The storage device 1040 may include, by way of non-limiting example, an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.

[0233] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software, as is known to those skilled in the art.

[0234] Program code loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and then loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and arithmetic logic.

[0235] In some embodiments, memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used, for example, to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, JVET, a new standard being developed by the Joint Video Experts Team).

[0236] Input to the elements of system 1000 can be provided through various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals, e.g., transmitted wirelessly by a broadcast station, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples, not shown in FIG. 34, include composite video.

[0237] In various embodiments, the input devices of block 1130 have associated respective input processing elements as known in the art. For example, the RF section may be associated with appropriate elements to (i) select a desired frequency (also referred to as selecting a signal or band-limiting a signal to a band of frequencies), (ii) downconvert the selected signal, (iii) band-limit again to a narrower band of frequencies to select (for example) a signal frequency band, which in some embodiments may be referred to as a channel, (iv) demodulate the downconverted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs a variety of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium, filter them to a desired frequency band, downconvert them, and filter them again to perform frequency selection. Various embodiments rearrange the order of the above-described (and other) elements, delete some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0238] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within processor 1010, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, in a separate interface IC or within processor 1010, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030, operating in combination with memory and storage elements to process the data stream, as desired, for presentation on an output device.

[0239] The various elements of system 1000 may be provided within an integrated housing in which the various elements are interconnected and capable of transmitting data between them using suitable connection configurations, such as internal buses known in the art, including an Inter-IC (I2C) bus, wiring, and printed circuit boards.

[0240] The system 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented within a wired and / or wireless medium, for example.

[0241] In various embodiments, data is streamed or otherwise provided to system 1000 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received over communication channel 1060 and communication interface 1050, which are adapted for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that delivers data over the HDMI connection of input block 1130. Still other embodiments provide streamed data to system 1000 using the RF connection of input block 1130. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0242] System 1000 can provide output signals to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. Display 1100 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 1100 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. Display 1100 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). In various example embodiments, other peripheral devices 1120 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 to provide functionality based on the output of system 1000. For example, a disc player performs the function of playing the output of the system 1000 .

[0243] In various embodiments, control signals are signaled between system 1000 and display 1100, speaker 1110, or other peripheral device 1120 using signaling such as AV.Link, Consumer Electronic Control (CEC), or other communication protocols that allow inter-device control with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000 in an electronic device, such as a television. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.

[0244] For example, if the RF section of input 1130 is part of a separate set-top box, display 1100 and speakers 1110 can alternatively be separate from one or more of the other components. In various embodiments in which display 1100 and speakers 1110 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0245] The embodiments may be implemented by computer software executed by the processor 1010, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type appropriate to the technology environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type appropriate to the technology environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0246] Various implementations include decoding. As used herein, "decoding" can encompass all or part of the processes performed, for example, on a received encoded sequence to generate a final output suitable for display. In various embodiments, such processes include one or more of processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include, or alternatively include, the processes performed by the decoders of the various implementations described herein.

[0247] As a further example, in one embodiment, "decoding" refers to entropy decoding only, in another embodiment, "decoding" refers to differential decoding only, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.

[0248] Various implementations include encoding. Similar to the above description of "decoding," as used herein, "encoding" can encompass all or part of the processes performed, for example, on an input video sequence to generate an encoded bitstream. In various embodiments, such processes include one or more of processes typically performed by an encoder, such as partitioning, differential encoding, transforming, quantization, and entropy encoding. In various embodiments, such processes also include, or alternatively include, the processes performed by the encoder of the various implementations described herein.

[0249] As a further example, in one embodiment, "encoding" refers to entropy encoding only, in another embodiment, "encoding" refers to differential encoding only, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.

[0250] Note that syntax elements, as used herein, are descriptive terms, so they do not preclude the use of other syntax element names.

[0251] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.

[0252] Various embodiments may refer to parametric models or rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is usually considered, often subject to computational complexity constraints. This can be measured through a rate-distortion optimization (RDO) metric, or through least mean squares (LMS), mean absolute error (MAE), or other such measures. Rate-distortion optimization is usually formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist for solving the rate-distortion optimization problem. For example, an approach can be based on extensive testing of all encoding options, including all considered modes or coding parameter values, with a thorough evaluation of their coding costs and the associated distortion of the reconstructed signal after encoding and decoding. To save on encoding complexity, faster approaches can also be used, particularly those that use approximated distortion calculations based on predicted or predicted residual signals rather than reconstructed ones. A mixture of these two approaches can also be used, such as by using approximated distortion for only some of the possible encoding options and full distortion for others. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches utilize any of a variety of techniques to perform optimization, but the optimization does not necessarily involve a complete evaluation of both the coding cost and the associated distortion.

[0253] The implementations and aspects described herein may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of implementation (e.g., described only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in a processor, which generally refers to a processing device and includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the transfer of information between end users.

[0254] References to "one embodiment" or "embodiment," or "one implementation" or "implementation," and other variations thereof, mean that particular features, structures, characteristics, etc. described in connection with an embodiment are included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment," or "in one implementation" or "in an implementation," and any other variations thereof, appearing in various places throughout this application are not necessarily all referring to the same embodiment.

[0255] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0256] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0257] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" generally includes in various ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0258] For example, it should be understood that the use of any of the following " / ," "and / or," and "at least one of" in the cases of "A / B," "A and / or B," and "at least one of A and B" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded to include as many items as are listed, as would be apparent to one skilled in the relevant arts and technology.

[0259] Also, as used herein, the term "signal" refers, among other things, to indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals a specific one of multiple transforms, coding modes, or flags. Thus, in embodiments, the same transform, parameter, or mode is used at both the encoder and decoder sides. Thus, for example, an encoder may transmit specific parameters to a decoder so that the decoder can use the same specific parameters (explicit signaling). Conversely, if the decoder already has certain parameters, etc., it may not transmit them, but may use signaling to simply enable the decoder to know and select the specific parameters (implicit signaling). By avoiding the transmission of any actual functions, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the above relates to the verb form of the word "signal," the word "signal" can also be used as a noun in this specification.

[0260] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0261] We have described numerous embodiments across various claim categories and types. The features of these embodiments may be provided alone or in any combination. Furthermore, embodiments may include one or more of the following features, devices, or aspects across various claim categories and types, alone or in any combination. • A process or device that uses MMVD and / or SMVD in conjunction with an affine motion model. A process or device that uses MMVD and / or SMVD with alternate temporal motion vector prediction. A process or device that uses MMVD and / or SMVD in conjunction with bidirectional optical flow. • A process or device that uses MMVD and / or SMVD with motion vector references to the current picture. A process or device that uses MMVD and / or SMVD in conjunction with generalized bi-prediction. • A process or device that uses MMVD and / or SMVD in conjunction with local illumination compensation. • A process or device that uses MMVD and / or SMVD in conjunction with a multiple hypothesis merge / intra combination mode. • A process or device that uses MMVD and / or SMVD along with merging with MVD or Ultimate MV Expression. • A process or device that uses MMVD and / or SMVD in conjunction with a per-subblock motion vector field based on a regression model. A process or device that uses bi-predictively coded, symmetric MVD, only one MVD, as well as MMVD and / or SMVD. • A process or device that uses MMVD and / or SMVD in conjunction with triangular partitioning. A process or device that combines MMVD and bi-prediction. • A process or device that combines MMVD with CPR. • A process or device that combines an MMVD and an ATMVP. • A process or device that combines MMVD with affine modes. • A process or device that combines MMVD with planar motion vector prediction. • A process or device that combines MMVD with a regression-based motion vector field. • A process or device that combines an MMVD with a GBI. • A process or device that combines MMVD with spatio-temporal motion vector prediction. - In AMVP mode, a process or device replaces the SMVD by the MMVD. • A process or device that combines SMVD with an affine motion model. • A process or device that combines SMVD with a planar motion model. • A process or device that combines SMVD with a regression-based motion model. • A process or device that combines SMVD with ATMVP motion models. • A process or device that combines SMVD with spatio-temporal motion vector prediction. • A process or device that modifies the SMVD mode to impose symmetry on bidirectional motion vectors in the translational case. • A process or device that modifies the SMVD modes to impose symmetries on the final affine model, including rotation, scaling, zooming, or translation. • A bitstream or signal that includes one or more of the described syntax elements or variations thereof. A bitstream or signal including syntax signaling information generated according to any of the described embodiments. - Creating and / or transmitting and / or receiving and / or decoding according to any of the described embodiments. - A method, process, apparatus, instruction storage medium, data storage medium, or signal according to any of the described embodiments. Insertion into the signaling of syntax elements that allow the decoder to determine the coding mode in a way that corresponds to that used by the encoder. • Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof. - A television, set-top box, cell phone, tablet, or other electronic device that performs a conversion method according to any of the described embodiments. ●A television, set-top box, cell phone, tablet, or other electronic device that performs the conversion method determination according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). ●A television, set-top box, cell phone, tablet, or other electronic device that selects, band-limits, or adjusts (e.g., using a tuner) a channel to receive a signal containing an encoded image and performs a conversion method according to any of the described embodiments. • A television, set-top box, cell phone, tablet, or other electronic device that receives a signal containing an encoded image wirelessly (e.g., using an antenna) and performs a conversion method. [Industrial Applicability]

[0262] The present invention can be used in communications. [Explanation of symbols]

[0263] 102, 102a-102d WTRU 104, 113 RAN 106 Core Network 108 PSTN 110 Internet 112 other networks 160a, 160b, 160c eNodeB 118 processors 120 Transceiver (transmitter / receiver) 122 Antenna

Claims

1. parsing a video bitstream for syntax indicating use of merge mode with motion vector difference; obtaining at least one motion vector difference from the syntax; refining at least one control point motion vector of an affine motion model with said at least one motion vector difference; decoding the block using the refined control point motion vector; A method for providing the above.

2. obtaining at least one motion vector difference from the syntax includes obtaining a single motion vector difference from the syntax; 2. The method of claim 1 , wherein the step of refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with the single motion vector difference.

3. obtaining at least one motion vector difference from the syntax includes obtaining one motion vector difference from the syntax for each control point motion vector of the affine motion model; 2. The method of claim 1, wherein the step of refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with a corresponding motion vector difference.

4. 4. The method of claim 3, wherein obtaining one motion vector difference for each control point motion vector of the affine motion model from the syntax comprises decoding, for each control point, a syntax element representing a direction and a syntax element representing a distance.

5. 4. The method of claim 3, wherein obtaining one motion vector difference for each control point motion vector of the affine motion model from the syntax includes decoding, for a first control point, a syntax element representing a direction and a syntax element representing a distance, and, for all other control points, a syntax element representing a difference of the distance from the first control point and a syntax element representing a direction.

6. 6. A method according to claim 4 or 5, wherein the range of magnitude of the motion vector difference distance for one control point is limited by the magnitude of the motion vector difference distance for another control point.

7. 7. The method of claim 1, wherein the step of decoding a block using the refined control point motion vectors comprises deriving an affine merge candidate list that includes only affine merge candidates inherited from affine neighboring blocks.

8. 1. An apparatus comprising one or more processors and at least one memory coupled to said processors, The one or more processors: parsing a video bitstream for syntax indicating use of merge mode with motion vector difference; obtaining at least one motion vector difference from the syntax; refining at least one control point motion vector of an affine motion model with said at least one motion vector difference; decoding the block using the refined control point motion vectors; 1. An apparatus configured to perform the steps of:

9. obtaining at least one motion vector difference from the syntax includes obtaining a single motion vector difference from the syntax; 9. The apparatus of claim 8, wherein the step of refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with the single motion vector difference.

10. obtaining at least one motion vector difference from the syntax includes obtaining one motion vector difference from the syntax for each control point motion vector of the affine motion model; 9. The apparatus of claim 8, wherein the step of refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with a corresponding motion vector difference.

11. 11. The apparatus of claim 10, wherein obtaining one motion vector difference for each control point motion vector of the affine motion model from the syntax comprises decoding, for each control point, a syntax element representing a direction and a syntax element representing a distance.

12. 11. The apparatus of claim 10, wherein obtaining one motion vector difference for each control point motion vector of the affine motion model from the syntax comprises decoding, for a first control point, a syntax element representing a direction and a syntax element representing a distance, and for all other control points, a syntax element representing a difference of the distance from the first control point and a syntax element representing a direction.

13. 13. An apparatus according to claim 11 or 12, wherein the range of magnitude of the motion vector difference distance for one control point is limited by the magnitude of the motion vector difference distance for another control point.

14. 14. The apparatus of claim 8, wherein the step of decoding a block using the refined control point motion vectors comprises deriving an affine merge candidate list that includes only affine merge candidates inherited from affine neighboring blocks.

15. indicating, via syntax in the video bitstream, the use of merge mode with motion vector difference; obtaining at least one motion vector difference; refining at least one control point motion vector of an affine motion model with said at least one motion vector difference; encoding syntax into the video bitstream indicating the at least one motion vector differential; encoding a block using the refined control point motion vector; A method for providing the above.

16. the step of obtaining at least one motion vector difference includes obtaining a single motion vector difference; 16. The method of claim 15, wherein refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with the single motion vector difference.

17. obtaining at least one motion vector difference includes obtaining one motion vector difference for each control point motion vector of the affine motion model; 16. The method of claim 15, wherein refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with a corresponding motion vector difference.

18. 18. The method of claim 17, wherein encoding syntax indicating the at least one motion vector difference into the video bitstream comprises encoding, for each control point, a syntax element representing a direction and a syntax element representing a distance.

19. 18. The method of claim 17, wherein encoding syntax indicating the at least one motion vector difference into the video bitstream comprises encoding, for a first control point, a syntax element representing a direction and a syntax element representing a distance, and, for all other control points, a syntax element representing the difference in distance from the first control point and a syntax element representing a direction.

20. 20. A method according to claim 18 or 19, wherein the range of magnitude of the motion vector difference distance for one control point is limited by the magnitude of the motion vector difference distance for another control point.

21. 21. A method according to claim 15, wherein encoding a block using the refined control point motion vectors comprises deriving an affine merge candidate list that includes only affine merge candidates inherited from affine neighboring blocks.

22. 1. An apparatus comprising one or more processors and at least one memory coupled to said processors, The one or more processors: indicating, via syntax in the video bitstream, the use of merge mode with motion vector difference; obtaining at least one motion vector difference; refining at least one control point motion vector of an affine motion model with said at least one motion vector difference; encoding syntax into the video bitstream indicating the at least one motion vector differential; decoding the block using the refined control point motion vector; 1. An apparatus configured to perform the steps of:

23. the step of obtaining at least one motion vector difference includes obtaining a single motion vector difference; 23. The apparatus of claim 22, wherein refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with the single motion vector difference.

24. obtaining at least one motion vector difference includes obtaining one motion vector difference for each control point motion vector of the affine motion model; 23. The apparatus of claim 22, wherein refining at least one control point motion vector of an affine motion model with the at least one motion vector difference comprises refining each control point motion vector of the affine motion model with a corresponding motion vector difference.

25. 25. The apparatus of claim 24, wherein encoding syntax indicating the at least one motion vector difference into the video bitstream comprises encoding, for each control point, a syntax element representing a direction and a syntax element representing a distance.

26. 25. The apparatus of claim 24, wherein encoding syntax indicating the at least one motion vector difference into the video bitstream comprises encoding, for a first control point, a syntax element representing a direction and a syntax element representing a distance, and, for all other control points, a syntax element representing the difference in distance relative to the first control point and a syntax element representing a direction.

27. 27. Apparatus according to claim 25 or 26, wherein the range of magnitude of the motion vector difference distance for one control point is limited by the magnitude of the motion vector difference distance for another control point.

28. 28. The apparatus of claim 22, wherein encoding a block using the refined control point motion vectors comprises deriving an affine merge candidate list that includes only affine merge candidates inherited from affine neighboring blocks.

Citation Information

Patent Citations

  • Video coding apparatus, video coding method, video coding program, transmission apparatus, transmission method, and transmission program

    JP2018196136A

  • Method, apparatus and computer program for video encoding

    JP2021524176A

  • Video signal processing method and apparatus using sub-block based motion compensation

    JP2022505578A

  • Interaction between intra block copy modes and inter predictors

    JP2022508177A

  • Hybrid motion vector coding modes for video coding

    US20130070855A1