Combinations of MMVD and SMVD with motion and predictive models

By extending MMVD and SMVD motion vector coding to various motion models and time prediction methods, the solution enhances video compression efficiency and flexibility, addressing limitations in existing systems and improving overall performance.

JP2026048900APending Publication Date: 2026-03-17INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing video compression systems face challenges in improving compression efficiency, particularly in intercoded blocks, and the application of motion vector coding tools like MMVD and SMVD is limited to translational models, lacking flexibility in motion model derivation and time prediction methods.

Method used

The proposed solution extends the use of MMVD and SMVD motion vector coding tools to all motion models and time prediction methods supported in the VVC standard, including affine, planar, recursive, triangular partition-based, GBI, and multiple hypothesis prediction methods, enhancing their application beyond simple translational models.

Benefits of technology

This extension improves the overall compression performance of video standards by providing enhanced flexibility and efficiency in motion vector coding, particularly in bi-prediction modes, leading to improved video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048900000001_ABST
    Figure 2026048900000001_ABST
Patent Text Reader

Abstract

Compared to existing video compression systems, it improves biprediction in intercoded blocks. [Solution] A general embodiment extends motion models beyond simple translational models by combining motion modes, such as merge with motion vector difference and symmetric motion vector difference, with, for example, merge and alternative time motion vector prediction modes. Embodiments extend the use of MMVD and SMVD motion vector coding tools to all motion model derivation and time prediction methods supported in the proposed video standard in order to improve overall compression performance. Specific embodiments describe combining MMVD or SMVD with affine motion models, ATMVD motion models, planar motion models, regressive motion fields, triangular partition-based motion models, GBI time prediction methods, LIC time prediction methods, and multiple hypothesis prediction methods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] At least one of these embodiments relates to a method or apparatus for video encoding or decoding, compression or decompression. [Background technology]

[0002] To achieve high compression efficiency, image and video coding schemes typically utilize prediction, including motion vector prediction, and transformation to leverage spatial and temporal redundancy within video content. Generally, intra or inter-prediction is used to utilize intra-frame or inter-frame correlation, and then the difference between the original and predicted images—often called prediction error or prediction residual—is transformed, quantized, and entropy-encoded. To reconstruct the video, the compressed data is decoded by the reverse process corresponding to entropy encoding, quantization, transformation, and prediction. [Overview of the project] [Problems that the invention aims to solve]

[0003] The present invention is in the field of video compression and aims to improve biprediction in intercoded blocks compared to existing video compression systems. [Means for solving the problem]

[0004] At least one of these embodiments relates to a method or apparatus for video encoding or decoding, and more particularly to a method or apparatus for simplifying encoding modes based on a neighbor-sample-dependent parametric model.

[0005] According to a first aspect, a method is provided. The method includes the steps of: indicating a first motion mode through syntax in a video bitstream; indicating the use of a second motion mode through the presence of syntax in a video bitstream, which, if present, includes information relating to the second motion mode; and encoding a video block using motion information corresponding to the first motion mode and the second motion mode.

[0006] A method is provided according to a second aspect of the method, which includes the steps of: analyzing a video bitstream to determine whether the syntax indicates a first motion mode; analyzing the video bitstream to determine whether the syntax indicates the presence of a second motion mode, and if present, determining information relating to the second motion mode; obtaining motion information corresponding to the first motion mode; and decoding a block using the motion information.

[0007] In another embodiment, an apparatus is provided. The apparatus comprises a processor. The processor can be configured to encode blocks of video or decode a bitstream by performing one of the methods described above.

[0008] According to another general aspect of at least one embodiment, a device is provided comprising an apparatus according to one of the decoding embodiments, and at least one of the following: (i) an antenna configured to receive a signal, the signal including a video block; (ii) a band limiter configured to limit the received signal to a bandwidth of frequencies including the video block; or (iii) a display configured to display an output representing the video block.

[0009] According to another general aspect of at least one embodiment, a non-temporary computer-readable medium is provided which includes data content generated according to any of the described encoding embodiments or variations.

[0010] According to another general aspect of at least one embodiment, a signal is provided which includes video data generated according to any of the described encoding embodiments or variations.

[0011] According to another general aspect of at least one embodiment, the bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.

[0012] According to another general aspect of at least one embodiment, a computer program product is provided which includes instructions, and when the program is executed by a computer, causes the computer to execute one of the described decoding embodiments or variations.

[0013] The features and advantages of these and other embodiments, as well as those of general embodiments, will become apparent from the following detailed description of exemplary embodiments, which should be read in conjunction with the accompanying drawings. [Effects of the Invention]

[0014] The present invention provides a novel method or apparatus for video encoding or decoding, compression or decompression. [Brief explanation of the drawing]

[0015] [Figure 1] This figure shows an example of an encoded tree unit and an encoded tree concept representing a compressed HEVC picture. [Figure 2] This figure shows an example of dividing a coding tree unit into coding units, prediction units, and transformation units. [Figure 3]It is a diagram showing a standard and general-purpose video compression scheme. [Figure 4] It is a diagram showing a standard and general-purpose video decompression scheme. [Figure 5] It is a simplified block diagram of the decoding process of the current scheme. [Figure 6] It is a MMVD diagram in the case of the bi-prediction mode. [Figure 7] It is a diagram showing an example of SMVD (symmetric motion vector difference). [Figure 8] It is a diagram showing an example of ATMVP motion prediction for a coding unit. [Figure 9] It is an example of a co-investigation model, a simple affine model used in VTM. [Figure 10] It is a diagram showing a 4×4 sub-CU-based affine motion vector field. [Figure 11] It is a diagram showing an example of a motion vector prediction process for an affine-inter-coded coding unit. [Figure 12] It is a diagram showing motion vector prediction candidates in the affine merge mode. [Figure 13] It is a diagram showing the spatial derivation of affine motion field control points in the case of affine merge. [Figure 14] It is a diagram showing an exemplary planar motion vector prediction process. [Figure 15] It is a diagram showing regression-based motion vector field construction. [Figure 16] It is a diagram of an example of dividing a coding unit into two triangular prediction units (PUs). [Figure 17] It is a diagram showing non-rectangular partitioning and associated OBMC diagonal weighting. [Figure 18] It is a diagram showing how the LIC parameter is derived from (a) reconstructed neighboring samples and (b) corresponding collocated reference samples. [Figure 19] It is a diagram showing an example of neighboring samples used for LIC parameter derivation. [Figure 20] This figure shows the spatial locations considered in the calculation of non-subblock STMVP merge candidates. [Figure 21] This is a simplified block diagram of the motion vector decoding process when MMVD and biprediction are combined. [Figure 22] This figure shows an example of a CPR constraint. [Figure 23] This figure shows the application of constraints to CPR MV in MMVD mode. [Figure 24] This figure shows the MMVD mode adaptation in the ATMVP case. [Figure 25] This is a simplified block diagram of the first version of the affine motion generation process using MMVD. [Figure 26] This is a simplified block diagram of the second version of the affine motion generation process using MMVD. [Figure 27] This figure shows the allowable size limits for differential MVD, based on the MVD index, used for the first CPMV of the affine motion field under consideration. [Figure 28] This figure shows an example of a planar MVP mode combined with MMVD. [Figure 29] This is a simplified block diagram of the first version of the PMVD motion generation process using MMVD. [Figure 30] This is a simplified block diagram of the second version of the PMVD motion generation process using MMVD. [Figure 31] This is a simplified block diagram of the third version of the PMVD motion generation process using MMVD. [Figure 32] This figure shows the construction of a regression-based motion vector field. [Figure 33] This is a simplified block diagram of a version of the affine motion generation process using SMVD. [Figure 34] This figure shows a processor-based system for encoding and decoding in the described embodiment. [Figure 35]This is a diagram of one embodiment of the encoding method of the general form described. [Figure 36] This is a diagram of one embodiment of the decoding method of the general form described. [Figure 37] This figure shows one embodiment of the apparatus in the general aspects described. [Modes for carrying out the invention]

[0016] The embodiments described herein are in the field of video compression and generally relate to video compression as well as video encoding and decoding. The general embodiments described aim to provide a mechanism for manipulating constraints in high-level video coding syntax or video coding semantics in order to restrict the possible set of tool combinations.

[0017] To achieve high compression efficiency, image and video coding schemes typically utilize prediction, including motion vector prediction, and transformation to leverage spatial and temporal redundancy within video content. Generally, intra or inter-prediction is used to utilize intra-frame or inter-frame correlation, and then the difference between the original and predicted images—often called prediction error or prediction residual—is transformed, quantized, and entropy-encoded. To reconstruct the video, the compressed data is decoded by the reverse process corresponding to entropy encoding, quantization, transformation, and prediction.

[0018] In the HEVC (High Efficiency Video Coding, ISO / IEC 23008-2, ITU-T H.265) video compression standard, motion-compensated time prediction is used to take advantage of the redundancy present between consecutive pictures in a video.

[0019] To do this, a motion vector is associated with each prediction unit (PU). Each coding tree unit (CTU) is represented by a coding tree in the compressed region. As shown in Figure 1, this is a quadtree partition of the CTU, where each leaf is called a coding unit (CU).

[0020] Each CU is then given several intra or inter-prediction parameters (prediction information). To do this, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. As shown in Figure 2, the intra or inter-coding mode is assigned at the CU level.

[0021] In HEVC, exactly one motion vector is assigned to each PU. This motion vector is used for the motion-compensated time prediction of the PU under consideration. Therefore, in HEVC, the motion model that connects the prediction block to its reference block is simply translation.

[0022] HEVC uses two modes to encode motion data: AMVP (Adaptive Motion Vector Prediction) and merging.

[0023] AMVP essentially involves signaling a reference picture used to predict the current PU, a motion vector predictor index (taken from a list of two predictors), and the motion vector difference. This document deals with merge modes and therefore does not address AMVP below.

[0024] The merge mode essentially involves signaling and decoding indices of several motion data collected within a list of motion data predictors. This list is constructed from five candidates and is built in the same manner on both the decoder and encoder sides. Thus, the merge mode aims to derive some motion information obtained from the merge list. The merge list generally contains motion information associated with several spatial and temporal peripheral blocks that are available in the decoded state when the current PU is being processed.

[0025] VTM-3 (VVC Draft 3) adopted a new type of motion vector coding called MMVD (Merge with Motion Vector Difference). The essence of MMVD mode is to introduce several motion vector differences, which are added to some conventional merge candidates, in order to generate some motion information for the block being encoded or decoded.

[0026] MMVD improves the coding efficiency of VVC coding systems. However, VVC Draft 3 states that the MMVD tool is only applicable in the normal translational merge mode. It is interesting to consider combining the MMVD tool with several other intercoding modes, and in particular with several other motion vector (or motion model in the case of subblock-based motion fields) generation tools included in the VVC coding system.

[0027] Another motion vector coding tool proposed for VVC standardization is so-called symmetric motion vector coding (SMVD). Like MMVD, SMVD is a new mode for improving the efficiency of motion vector coding. It is attempted within the AMVP mode. However, SMVD is only applicable to the translational AMVP case. With regard to MMVD, it can be beneficial to extend it to several motion models beyond simple translational models.

[0028] In the JVET (Joint Video Research Team) proposal for a new video compression standard known as the Joint Research Model (JEM), it was proposed to accept a quadtree-binary tree (QTBT) block partitioning structure for high compression performance. In a binary tree (BT), a block can be divided into two equally sized subblocks by splitting it horizontally or vertically down the middle. As a result, BT blocks can have a rectangular shape with unequal width and height, unlike blocks in QT where blocks always have a square shape with equal height and width. In HEVC, the angular intra-prediction direction is defined over 180 degrees from 45 degrees to -135 degrees, and in JEM, these are maintained, which makes the definition of the angular direction independent of the target block shape.

[0029] To encode these blocks, intra-prediction is used to provide an estimated version of the block using previously reconstructed neighbor samples. The difference between the source block and the prediction is then encoded. In the conventional codecs described above, one line of reference samples is used to the left and above the current block.

[0030] In HEVC (High Efficiency Video Coding, H.265), the encoding of frames in a video sequence is based on a quadtree (QT) block partitioning structure. Frames are divided into square coding tree units (CTUs), which all undergo quadtree-based partitioning into multiple coding units (CUs) based on a rate-distortion (RD) criterion. Each CU is intra-predicted, i.e., spatially predicted from its causative neighboring CUs, or inter-predicted, i.e., temporally predicted from an already decoded reference frame. In I-slices, all CUs are intra-predicted, while in P-slices and B-slices, CUs can be both intra-predicted and inter-predicted. For intra-prediction, HEVC defines 35 prediction modes, including one planar mode (indexed as mode 0), one DC mode (indexed as mode 1), and 33 angular modes (indexed as modes 2-34). The angular mode is associated with the prediction direction, ranging from 45 degrees to -135 degrees in a clockwise direction. Since HEVC supports a quadtree (QT) block partitioning structure, all prediction units (PUs) are square in shape. Thus, the definition of prediction angles from 45 degrees to -135 degrees is justified in terms of the shape of the PU (prediction unit). For a target prediction unit of size N×N pixels, the upper and left reference arrays are each 2N+1 samples in size, and they are required to cover the aforementioned angular range for all target pixels. Given that the height and width of the PU are of equal length, the equality of the lengths of the two reference arrays is also a given.

[0031] This invention lies in the field of video compression. It aims to improve biprediction in intercoded blocks compared to existing video compression systems. The invention also proposes separating the lumar-coded tree and the chromar-coded tree for interslices.

[0032] In the HEVC video compression standard, a picture is divided into so-called coding tree units (CTUs), which are typically 64x64, 128x128, or 256x256 pixels in size. Each CTU is represented in the compression region by a coding tree, which is a quadtree division of the CTU, with each leaf being called a coding unit (CU).

[0033] Each CU is then given several intra or inter-prediction parameters (prediction information). To do this, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. The intra or inter-coding mode is assigned at the CU level.

[0034] The newly emerging video compression tools include the proposal of an encoded tree unit representation in the compression region to represent picture data in a more flexible way within the compression region. The advantage of this more flexible representation of the encoded tree is that it provides improved compression efficiency compared to the CU / PU / TU arrangement of the HEVC standard.

[0035] The quadtree-plus-binary (QTBT) coding tool offers this enhanced flexibility. Its essence lies in being a coding tree that can divide coding units in both quadtree and binary ways. Such coding tree representations of coding tree units are illustrated.

[0036] The partitioning of the coding units is determined on the encoder side through a rate distortion optimization procedure, which essentially involves determining the QTBT representation of the CTU that has the lowest rate distortion cost.

[0037] In QTBT technology, the CU (Code Unit) can be either square or rectangular in shape. The size of the coding unit is always a power of 2, generally ranging from 4 to 128.

[0038] In addition to this diversity in the rectangular shape of the coding unit, this new CTU representation has the following different characteristics compared to HEVC.

[0039] The QTBT decomposition of a CTU is performed in two stages: first, the CTU is divided in a quadtree manner, and then each quadtree leaf can be further divided in a binary tree manner.

[0040] One problem solved by the present invention is how to extend the use of MMVD and SMVD motion vector coding tools to all motion model derivation and time prediction methods supported in currently proposed video standards, thereby improving the overall compression performance of these proposed standards.

[0041] The fundamental principles of this disclosure consist of two aspects.

[0042] - To extend the use of MMVD motion vector coding with all motion models and / or time prediction methods supported in VVC Draft 3, as well as with other motion models proposed in the VVC standardization process. In particular, this disclosure describes how MMVD can be combined with affine motion models, ATMVP motion models, planar motion models, recursive motion fields, triangular partition-based motion models, GBI time prediction methods, LIC time prediction methods, and multiple hypothesis prediction methods. Enhanced use of MMVD in the case of biprediction is also provided.

[0043] - To extend the use of the SMVD motion vector coding tool with all motion models and / or time prediction methods supported in VVC Draft 3, as well as with other motion model generators proposed in the VVC standardization process. In particular, this disclosure describes how SMVD can be combined with affine motion models, ATMVP motion models, planar motion models, recursive motion fields, triangular partition-based motion models, GBI time prediction methods, LIC time prediction methods, and multiple hypothesis prediction methods.

[0044] In the VVC video standards currently under development, multiple interpredictive modes are supported or proposed.

[0045] Figure 5 shows a simplified block diagram of the interdecoding process in version 3 of the VVC draft. The process begins with the derivation of an MV predictor (mvp) based on the construction of an MV candidate list (1001), which can differ according to the intermode used for encoding the CU or PU.

[0046] Step 1002 is a motion vector difference (MVd) decoding step. The decoded MVd is used to reconstruct the motion vector of the CU / PU under consideration by adding it to the motion vector predictor obtained in the preceding step (1001).

[0047] The next step, 1003, essentially involves deriving a motion model from the motion vector values ​​derived from the two preceding steps. The model can be simply translational, or more sophisticated, as will be explained later.

[0048] Using the decoded MV, the prediction of samples within the CU is performed in step 1004 for one reference picture in the case of bi-prediction mode, and in step 1005 for the other reference picture. The predicted signal is further refined in step 1006 for one reference picture in the case of bi-prediction mode, and in step 1007 for the other reference picture. Where applicable, bi-prediction, or a combination of intra / inter-prediction, or PU within the CU is achieved in step 1008. In step 1009, the final refinement step of the predicted signal is performed.

[0049] Figure 5 also shows the different modes applied to these different steps in VTM3. Tools indicated with (*) correspond to tools that have been proposed for adoption but were not adopted in the VVC Draft 3 stage.

[0050] In the interface, two basic modes for deriving the MV, merge and AMVP, are used. In both cases, a list of MV candidates is derived. This derivation process generally differs for these two modes, and even for a given mode, modifications are applied depending on other settings applied to the CU (e.g., merge or AMVP, translational model or affine model).

[0051] In merge mode, the predictor index is signaled to fetch MVs in the merge candidate list. In addition, the sample prediction residual is signaled in cases where skip mode is not applied. In AMVP mode, for each reference picture (one reference picture in the single prediction case, two in the dual prediction case), the reference frame index, predictor index, and motion vector difference (mvd) are signaled. In addition, the sample prediction residual is signaled.

[0052] Table 1 provides an overview of the contents of different candidate lists in the current VTM.

[0053] [Table 1]

[0054] A summary of the intermodes considered in this invention, along with some representational figures regarding performance (as a percentage of bitrate fluctuations in a random access configuration for psnrY), is provided in the following table.

[0055] [Table 2]

[0056] The following subsections provide further details about the intermodes under consideration, along with the syntax elements associated with these modes or tools.

[0057] The general embodiments described herein focus on the “Mvd Derivation” and “Model Derivation” steps in Figure 5. The goal of this disclosure is to extend the use of MMVD and SMVD motion vector coding tools to motion models beyond simple translational motion models.

[0058] [MMVD Motion Vector Difference Coding Tool Description] MMVD is applied only in merge mode. It uses a merge candidate list. The flag "mmvd_skip" indicates whether MMVD mode is applied. When the mode is applied, the MV difference (mmvd) is constructed as follows: - The syntax element (referred to here as mmvd_idx) is signaled to construct a corrected MV mmvd consisting of the following information: ○The base MV index selected by the encoder from among the two first translational merge candidates. ○ An index denoted as mmvd_dir_idx, associated with direction D (currently 4) in the (x,y) coordinate system (specified is a table dir[] consisting of 4 elements {(0,1),(1,0),(-1,0),(0,-1)}). ○ An index denoted as mmvd_dist_idx related to the distance step S from the base MV (currently, up to 8 distances are possible, and this involves specifying a table dist[] consisting of 8 elements {1 / 4 Pel, 1 / 2 Pel, 1 Pel, 2 Pel, 4 Pel, 8 Pel, 16 Pel, 32 Pel}).

[0059] When MMVD mode is applied, the MV difference is then calculated as follows: refinementMV=dir[mmvd_dir_idx]×dist[mmvd_dist_idx]

[0060] Even when the CU is encoded in biprediction, a single MV difference is signaled. In the biprediction case, two symmetric MV differences are obtained from a single encoded MV. When the temporal distance between the predicted picture and the reference picture differs between reference picture lists L0 and L1, the decoded mmvd is assigned to the MV difference (mvd) associated with the largest temporal distance. The mvd associated with the smallest distance is scaled as a function of the POC distance.

[0061] For example, consider the case where reference picture k (k=0 or 1) is the closest to the current picture. Define kk=1-k, and POCref_0, POCref_1, and POCcur as the picture order counts of reference picture 0, reference picture 1, and the current picture, respectively. The scaling factor is derived as follows: sc=(POCref_kk-POCcur) / (POCref_k-POCcur)

[0062] Next, the refined MV for each of the reference pictures is derived as follows: refinementMV_k=refinementMV refinementMV_kk=sc×refinementMV

[0063] In VVC Draft 3, MMVD is applied only to translational motion models.

[0064] [Symmetric MVD proposed in JVET-L0370] A symmetric MVD (SMVD) tool is being considered for VVC. Its principle is to encode some motion vector information under the constraint that, in the case of biprediction, the motion information of the CU is made up of two symmetric forward motion vector differences and backward motion vector differences. In VVC Draft 3, the SMVD mode is applied only to AMVP.

[0065] Under these constraints, known as SMVD mode, the encoding of the CU is signaled through the CU level flag symmetrical_mvd_flag.

[0066] This flag is encoded when the SMVD mode is feasible, i.e., when the prediction mode of the CU is biprediction and two reference pictures for the CU are found, as follows:

[0067] - The reference picture for the current CU is searched for as the nearest forward and backward reference picture in the (L0 and L1) or (L1 and L0) reference picture lists, respectively. If not found, SMVD mode is not applicable and symmetrical_mvd_flag is omitted.

[0068] If symmetrical_mvd_flag is signaled and equals true, - For the L0 reference picture, one mvd is signaled, and for the other reference picture list, the mvd is derived symmetrically, i.e., as the opposite of the first one. - As in the conventional AMVP mode, two MV predictor indices (one per reference picture list) are signaled.

[0069] [Motion models supported in VVC Draft 3, or proposed for VVC] *Translational Model (VVC Draft 3) By default, motion within a CU is based on translational MV, which is applied to all samples within the block. *ATMVP (VVC Draft 3)

[0070] In the alternative (advanced) temporal motion vector prediction (ATMVP) method shown in Figure 8, one or more temporal motion vector predictors for the current CU are taken from a reference picture for the current CU.

[0071] First, the so-called time-motion vector and its associated reference picture index are obtained as motion data associated with the first candidate in the current CU's normal merge candidate list.

[0072] Next, the current CU is divided into N×N subCUs, where N is generally equal to 4. This is shown in Figure 8. For each N×N subblock (subCU), the motion vector and reference picture index are identified in the reference picture associated with the time MV, with the help of the time motion vector. The N×N subblocks in the reference picture, pointed to by the time MV from the position of the current subCU, are considered. Their motion data is obtained as ATMVP motion data predictions for the current subCU. It is then transformed into the motion vector and reference picture index of the current subCU through appropriate motion vector scaling.

[0073] Note that in VVC Draft 3, the ATMVP motion information predictor is part of the subblock-based merge candidate list.

[0074] *Affine motion model (VVC Draft 3) One of the new motion models introduced in VVC is the affine mode, which essentially involves using an affine motion field to represent motion vectors within the CU.

[0075] A motion model used for two or three (hereinafter also called control point) control point motion vectors is illustrated in Figure 9. The four-parameter affine motion field based on two control points for each position (x,y) within the block under consideration has the following motion vector component values.

[0076]

number

[0077] Here, (v 0x ,v 0y ) and (v 1x ,v 1y ) is the so-called control point motion vector (CPMV) used to generate the affine motion field. (v 0x ,v 0y ) is the motion vector of the control point in the upper left corner. (v 1x ,v 1y ) is the motion vector of the control point in the upper right corner. A 6-parameter affine field generated from three CPMVs is also possible, as described in JVET-J0021.

[0078] In practice, to maintain reasonable complexity, affine motion is managed on a 4x4 subblock basis; that is, the same motion vector is used for each sample within each 4x4 subblock (subCU) of the CU under consideration (see Figure 10). The affine motion vector is calculated from the CPMV at the center position of each subblock. The obtained MV is expressed with 1 / 16 Pell precision.

[0079] As a result, the entire block (CU) is predicted in time through motion compensation of each 4x4 subblock (subCU), each having its own MV. In VTM, affine motion compensation can be used in two ways, affine inter (AF_INTER) and affine merge (or merge affine), as described below.

[0080] -AF_INTER- CUs in AMVP mode whose size is larger than 8x8 can be predicted in affine intermode. This is signaled through the flag inter_affine_flag, which is encoded at the CU level. Generating the affine motion field for that interCU involves determining the control point motion vector (CPMV), which is obtained by the decoder through the addition of the motion vector difference plus the control point motion vector prediction (CPMVP). The CPMVP is a pair (for a 4-parameter affine model with two CPMVs) or a triple (for a 6-parameter affine model with three CPMVs) of motion vector candidates, which can be inherited from affine neighbors (as in affine merge mode) or constructed from non-affine motion vectors obtained from lists (A, B, C), (D, E), and / or (F, G), respectively, as illustrated in Figure 11. This corresponds to the "virtual candidates" mentioned in Table 1.

[0081] -Affin Merge- In affine merge mode, the CU level flag indicates whether the CU in merge mode utilizes affine motion compensation. If so, JEM (search and reference software developed by JVET prior to VVC) selects the first available neighbor CU encoded in affine mode from the ordered list of candidate positions (A, B, C, D, E) in Figure 12.

[0082] Once the first neighboring CU in affine mode is acquired, three motion vectors are generated from the upper left corner, upper right corner, and lower left corner of the neighboring CU.

[0083]

number

[0084] However, these are extracted (see Figure 13). Based on these three vectors, two or three CPMVs for the top-left, top-right, and / or bottom-left corners of the current CU are derived as follows:

[0085]

number

[0086] Current CU control point motion vector

[0087]

number

[0088] However, once acquired, the current motion field within the CU is calculated on a 4x4 subCU basis through the model of Equation 1.

[0089] In parallel applications, more candidates for affine merge mode are considered. In this case, the encoder selects the best candidate through a rate-distortion optimization process, and the index of this best candidate is encoded in the bitstream through the merge_idx syntax element.

[0090] The next affine candidate to be placed in the affine merge candidate list is a “constructed” (or “virtual”) affine model candidate, as opposed to an inherited candidate. A constructed candidate is an affine motion field computed based on available neighbor motion vectors around the current CU, including those from neighbors that are not encoded in affine mode. Two or three neighbor MVs are used to derive and generate a candidate affine motion field for the current CU.

[0091] The neighborhood MVP used to construct the affine merge candidate includes several spatial MVPs and one temporal MVP.

[0092] [(Plane motion model proposed in JVET-L0070)] In JVET-L0070, planar motion vector prediction (PMVP) has been proposed as an additional merge mode for VVC codec design. The planar motion vector prediction shown in Fig. 14 is achieved by averaging horizontal and vertical linear interpolations on a 4×4 block basis as follows. These linear interpolations are performed from the generator MVs located on the corners (AR, AL, BL) of the CU. Its principle is illustrated in Fig. 14. P(x,y)=(H×P h (x,y)+W×P v (x,y)+H×W) / (2×H×W)

[0093] [Regression MVF (RMVF) model] In JVET-L0171, regression-based motion vector prediction has been proposed as an additional merge mode for VVC codec design.

[0094] The regression-based motion vector field (RMVF) shown in Fig. 15 essentially consists of the following. The principle of RMVF is to use a 6-parameter motion model to calculate the motion vectors of sub-blocks.

[0095] [Number]

[0096] The motion parameters are calculated based on one row and one column of 4×4 sub-blocks that are spatially adjacent, using their motion vectors and center positions as the input to the linear regression method.

[0097] [Triangle partition-based motion model] In VVC Draft 3, a triangular motion partitioning tool was adopted. This allows for flexibility in dividing a block into two prediction units, as shown in Figure 16. Essentially, a CU is divided into two triangular prediction units in a diagonal or inverse diagonal direction. Each triangular PU is interpreted using its own single prediction motion vector and reference picture, derived from a dedicated single prediction candidate list.

[0098] After predicting the triangular prediction units, an adaptive weighting process (see Figure 17) is performed on the diagonal edges. Subsequently, transformation and quantization processes are applied to the entire CU. The triangular PU mode is applied only to the skip mode and merge mode.

[0099] [Local Lighting Compensation (LIC)] The purpose of LIC is to compensate for lighting changes that may occur between the prediction block and its reference block, which are utilized through motion-compensated (MC) time prediction. In this tool, the decoder calculates several prediction parameters based on several reconstructed picture samples located in the left column and / or upper row of the current block and a reference picture sample located in the left column and / or upper row of the motion-compensated block (Figure 18a).

[0100] In an alternative approach, predicted samples of neighboring reconstructed blocks on the left / upper side (Figure 18b) can be used.

[0101] In the case of biprediction, the variation of LIC ("bidirectional LIC") is essentially about estimating the illumination change between the two current reference blocks in order to derive the LIC parameters for the current block.

[0102] The LIC parameters are selected to minimize the mean squared error difference (MSE) between the sample in Vcur and the corrected sample in Vref(MV). Generally, the LIC model is linear, i.e., LIC(x) = a.x + b.

[0103]

number

[0104] s and r correspond to the pixel positions within Vcur and Vref(MV), respectively, as shown in Figure 19.

[0105] As a result, when LIC is used for time prediction, a linear LIC model is applied to the motion-compensated block to provide a time prediction for the current block.

[0106] [Generalized Bifurcation (GBI)] In VVC Draft 3, GBI was adopted. In dual prediction mode, it applies unequal weights to predictors from L0 and L1. In inter-prediction mode, multiple weight pairs, including equal weight pairs (1 / 2, 1 / 2), are evaluated based on rate-distortion optimization (RDO), and the GBI index of the selected weight pair is signaled to the decoder.

[0107] In AMVP mode, the GBI index that carries GBI weight information is signaled at the CU level.

[0108] In merge mode, the GBI index is inherited from neighboring CUs. In GBI mode, the predicted blocks are calculated as follows: P GBi =(w0×P L0 +w1×P L1 ) Here, w0 and w1 are the selected GBI weights. The supported w1 values ​​are generally {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}. Since the sum of w1 and w0 is equal to 1, the corresponding w0 values ​​are {5 / 4, 5 / 8, 1 / 2, 3 / 8, -1 / 4}. Weight pairs are selected and signaled at the CU level. For non-low latency pictures, the number of weights is reduced. The w1 and w0 values ​​are {3 / 8, 1 / 2, 5 / 8} and {5 / 8, 1 / 2, 3 / 8}, respectively.

[0109] 1.1.1. Non-subblock spacetime merge motion vector predictor (STMVP) This section describes prior art method JVET-L0354, called STMVP, for generating space-time merge candidates. The method essentially involves extracting two spatial neighbor motion vectors, one above and one to the left of the current PU, and a time motion vector predictor for the current CU.

[0110] Spatial neighbors are obtained at a spatial location called Afar, as illustrated in Figure 20. The spatial location of Afar, relative to the top-left position of the current PU, is given by coordinates (nbPW × 2, -1), where nbPW is the width of the current block. If a motion vector is not available (does not exist or is not intra-encoded) at location Afar, location B1 is considered the upper neighbor motion vector of the current block.

[0111] The selection of the left neighboring block is similar. If available, the neighboring motion vector at the relative spatial position (-1, 2 × nbPH), denoted as Lfar, is selected, where nbPH is the height of the current block. If not available, the left neighboring motion vector at position A1 (-1, nbPH-1), if available, is selected.

[0112] Next, the TMVP predictor in the current block is derived in the same way as in the time motion vector prediction of HEVC.

[0113] Finally, the STMVP merge candidate for the current block is calculated as the average of up to three acquired spatial and temporal neighbor motion vectors. Thus, the STMVP candidate here is constructed from at most one motion vector per reference picture list, in contrast to the subblock-based approach to STMVP in JEM. Several variations of the JVET-L0354 method have been proposed, for example, in JVET-L0207.

[0114] [Proposed Embodiments for Extending the Use of MMVD] Combination of MMVD and biprediction In the current MMVD design, a single set of mmvd data, constructed from the direction index and distance index, is signaled even in the case of biprediction.

[0115] [Embodiment 1] - Interdependence signaling of mmvd data for both reference pictures In this embodiment, when the CU is encoded in bipred mode and the MMVD mode is enabled for the CU, two sets of mmvd data are signaled: one set for each of the two mmvds, applied to each of the reference pictures (from lists L0 and L1). In addition, a second set of mmvd data can be derived from the first set of mmvd data, with a variety of possible options.

[0116] This is illustrated in the simplified block diagram in Figure 21. In step 1101, the mmvd data is decoded for the first reference picture. If biprediction is applied (check in step 1102), the mmvd data is decoded for the second reference picture, conditional on the mmvd data of the first reference picture (step 1103). Then, in step 1104, the biprediction MV is derived from the mmvd data of the first and second reference pictures. If biprediction is not applied (check in step 1102), in step 1105, the single prediction MV is derived from the mmvd data of the first reference picture.

[0117] [Embodiment 1] - Option 1 - Estimated orientation for the mmvd of the second reference picture In one option, only the distance index is encoded for both mmvd(mmvd0_dist_idx and mmvd1_dist_idx). The direction of the second mmvd is inferred from the direction of the first mmvd (mmvd0_dir_idx). If both reference pictures are located on the same time side with respect to the current picture, mmvd1_dir_idx is set to equal mmvd0_dir_idx. If the reference pictures are located on both time sides from the current picture, mmvd1_dir_idx is set to equal -mmvd0_dir_idx.

[0118] An example of the related simplified syntax is shown below.

[0119] [Table 3]

[0120] [Embodiment 1] - Option 2 - Differential encoding of the mmvd distance of the second reference picture In another option, the distance to the second reference picture is encoded as the difference to the distance to the first reference picture. The index mmvdd1_dist_idx, corresponding to the difference to mmvd0_dist_idx, is encoded.

[0121] An example of the related simplified syntax is shown below.

[0122] [Table 4]

[0123] The distance for refined MV0 is, distMV0=dist[mmvd0_dist_idx] The distance for refined MV1 is calculated as follows: distMV1=distMV0+dist[mmvdd1_dist_idx] It is calculated as follows.

[0124] [Embodiment 1] - Option 3 - Limitation of possible distance values ​​for the mmvd of the second reference picture In this embodiment, the maximum value of the second distance (which is either the distance signaled by mmvd1_dist_idx or the distance difference as in Option 2, signaled by mmvdd1_dist_idx) is made smaller than that of mmvd0_dist_idx. For mmvd0_dist_idx, we assume that N possible distances are used (e.g., 8, dist[]={1 / 4 Pel, 1 / 2 Pel, 1 Pel, 2 Pel, 4 Pel, 8 Pel, 16 Pel, 32 Pel}). For mmvdd1_dist_idx, only the first N' values ​​of dist[] can be used, and N' <Nである。

[0125] The following variations are possible. - N' is calculated from N (for example, N' = N / 2). - N' is calculated from the value of mmvd0_dist_idx (for example, N'=mmvd0_dist_idx, or N'=mmvd0_dist_idx / 2). - N' depends on the value of the scale parameter "sc" (as described in the section on MMVD encoding tools). For example, N' = sc × N

[0126] [MMVD-CPR combination] In one embodiment, when CPR (Current Picture Reference) mode is applied, MMVD mode is enabled with CPR-specific adaptations.

[0127] The MV used for CPR refers to the current picture. In VVC Draft 3, constraints were stipulated to limit the MV to a point in a constrained region close to the current CU in order to limit the memory storage needs. In the current VTM, as illustrated in Figure 22, the constrained region consists of the CTU containing the current CU. Another constraint of the CPR MV is that it must have integer precision.

[0128] [Embodiment 2a] - mmvd constraint to integer precision in the case of CPR In embodiments, the enabled distance for mmvd is limited to one that conforms to the MV precision enabled for CPR. For example, in current VTMs, the CPR MV precision is an integer, and the present invention stipulates that when CPR mode is used, the distance for mmvd is the full set typically used for mmvd. dist[]={1 / 4 Pell, 1 / 2 Pell, 1 Pell, 2 Pell, 4 Pell, 8 Pell, 16 Pell, 32 Pell} Instead, the following set, namely, {1 Pel, 2 Pel, 4 Pel, 8 Pel, 16 Pel, 32 Pel} It is considered to be inside.

[0129] In the case of CPR, the maximum value N' (e.g., 5) of mmvd_dist_idx is less than the maximum value N (e.g., 8) of mmvd_dist_idx.

[0130] In one embodiment, this solution is implemented using an offset applied to mmvd_dist_idx, as follows: - When MMVD mode is applied, ○When CPR mode is applied, □distMV=dist[mmvd_dist_idx+offset].

[0131] ○In other cases, □distMV=dist[mmvd_dist_idx].

[0132] To limit the precision to integers, the offset is equal to 2.

[0133] [Embodiment 2b] - Limitation of the maximum value of mmvd in the case of CPR In the embodiment, the maximum distance for mmvd in the case where CPR mode is activated is made smaller compared to conventional mmvd usage. For example, in current VTM, the CPR MV precision is an integer, and the present invention makes the distance for mmvd when CPR mode is activated a full set of mmvd that is normally used for mmvd. dist[]={1 / 4 Pell, 1 / 2 Pell, 1 Pell, 2 Pell, 4 Pell, 8 Pell, 16 Pell, 32 Pell} Instead, the following set, namely, {1 Pel, 2 Pel, 4 Pel, 8 Pel} It is considered to be inside.

[0134] In the case of CPR, the maximum value N' of mmvd_dist_idx (e.g., 3) is less than the maximum value N of mmvd_dist_idx (e.g., 8).

[0135] [Embodiment 2c] - Clipping MV from MMVD in the case of CPR In this embodiment, the MV resulting from the mmvd refinement is clipped so that the motion compensation block used for prediction remains within a constrained region. The basic process is illustrated in Figure 23.

[0136] Embodiment 3: Combination of MMVD and ATMVP [Embodiment 3.1] - MMVD used in combination with subblock base mode ATMVP In VVC Draft 3, MMVD applies only to the normal translational merge mode. Therefore, when a subblock-based merge mode is activated for a given CU, MMVD is not used.

[0137] According to this embodiment 3, the MMVD motion vector representation tool can be used when the subblock-based merge mode is activated. In particular, MMVD can be used in combination with the ATMVP motion vector prediction mode.

[0138] When the merge_flag syntax indicates the use of a merge mode for the current CU, the use of MMVD is signaled in the same way as in VVC Draft 3. Furthermore, the flag subblock_merge_flag, which indicates the use of a subblock-based merge mode, is encoded regardless of the value of the syntax element mmvd_merge_flag, which indicates the use of MMVD for the current block.

[0139] According to the first sub-embodiment 3.0, affine mode is not permitted in combination with MMVD. Therefore, if mmvd_merge_flag is on, it is inferred that the merge mode used is ATMVP mode. Consequently, the merge_idx syntax element is omitted from the encoded bitstream.

[0140] According to another sub-embodiment, affine mode is permitted in combination with MMVD. Thus, the merge_idx syntax element is encoded as is currently done in VVC Draft 3.

[0141] [Embodiment 3.2] - mmvd used as refinement of the inherited subblock MV of ATMVP In this embodiment, when an ATMVP candidate is selected and MMVD is activated, mmvd refinement is applied to all MVs, which are inherited for each subblock. The process is shown in the simplified block diagram of Figure 24. In the first step, the mmv data is decoded. In the next step, the global ATMVP MV is identified. Then, for each subblock of the CU, motion data from the corresponding subblock, identified in the reference picture by the global MV, is fetched. This motion data is refined using the mmvd data, similar to the conventional MMVD refinement process.

[0142] [Embodiment 4]: Combination of MMVD and affine mode According to this embodiment, MMVD is used in combination with affine mode. Several modifications can be applied within this embodiment, which are listed below.

[0143] [A single MMVD encoded for all CPMVs] - Following the first variation, a single motion vector difference mmvd is applied to all CPMVs, which are encoded and used to generate the affine motion field. This results in adding a constant motion vector to the entire affine motion model. This is illustrated in the following figure, which shows a simplified decoding block diagram of the affine motion field generation process when the MMVD mode is activated.

[0144] ○In the case of bidirectional affine prediction, the same symmetric motion vector difference concept is used as in the translational case.

[0145] [MMVD data encoded for each CPMV] - Following the second variation, one motion vector difference mmvd is allowed for each CPMV, which improves the flexibility for generating candidate affine motion fields. For the first CPMV, the normal mmvd coding of MMVD can be used. Then, for the second and third CPMVs, a differential motion vector difference (mmvdd) can be coded on top of the mmvd associated with the first CPMV. The advantage of such a technique is that it limits the rate cost of mmvd signaling while allowing some flexibility in the application of MMVD to affine modes. An example of the relevant simplified syntax in the case of coding mmvd information for three affine CPMVs is shown below. According to this embodiment, the magnitude of the differential motion vector difference can be restricted to a value smaller than the mmvd of the first CPMV (for example, the full set of values ​​can be dist[]={1 / 4 Pel, 1 / 2 Pel, 1 Pel, 2 Pel, 4 Pel, 8 Pel, 16 Pel, 32 Pel}, but it can be restricted to dist[]={1 / 4 Pel, 1 / 2 Pel, 1 Pel}). The process is illustrated in the following figure, which shows a simplified decoding block diagram of the affine motion field generation process when the MMVD mode is activated.

[0146] [Table 5]

[0147] - According to some modifications, in differential mmvd encoding, the allowed range of magnitudes is limited according to the distance index associated with the mmvd of the first CPMV. For example, mmvd1_dist_idx, and when applicable, mmvd2_dist_idx, are constrained to be lower than mmvd0_dist_idx. The allowed range of the second mmvd in the horizontal and vertical directions can be limited as illustrated in Figure 27.

[0148] - Following further modifications, the same directional index is used for all mmvds, and therefore, for the second and third CPMVs, only distance information can be encoded. An example of the relevant simplified syntax in the case of encoding mmvd information for three affine CPMVs is shown below.

[0149] [Table 6]

[0150] [Constraints on the affine merge candidate list due to the removal of virtual candidates] - Following further modifications, the MMVD mode used for affines introduces some diversity within the set of affine merge candidates, so the affine list of affine merge candidates is reduced and simplified overall compared to the affine merge list in VVC Draft 3, leading to a more compact and simplified overall codec design. For example, in one embodiment, several constructed (virtual) affine model candidates are removed from the affine merge list. In another embodiment, all constructed (virtual) affine model candidates are removed from the affine merge list, and the affine merge list is made up only of inherited affine candidates.

[0151] [Constraints on the second (and third) mmvd distance values ​​based on the first mmvd distance value] - According to additional properties, for the second CPMV and (when applicable) the third CPMV, the differential encoding of the mmvd (motion vector difference) to the mmvd of the first CPMV includes the possibility of encoding / decoding a distance value of 0 for the differential motion vector difference (dist[]={0 Pel, 1 / 4 Pel,...}). In practice, this applies only in the translational motion case, and in existing MMVDs, distances equal to 0 are not supported because it provides an MV predictor that overlaps with that of the normal merge.

[0152] [Imposing symmetry constraints on affine model parameters in the case of bidirectional affine models] - According to further properties, in the bidirectional affine case, the derivation of the second mmvd (associated with the second reference picture) as a function of the first one (associated with the first reference picture) uses symmetry constraints, similar to the symmetry constraints imposed between the two bipredictive mmvds in the translational case. To do so, several techniques may be possible. ○In one first method, the mmvd of the mmvd associated with the second reference picture list is directly inferred from the mmvd associated with the first reference picture list. For example, the second mmvd is derived by scaling the first mmvd, and the scaling takes into account the temporal distance between the reference picture and the current picture. ○In another method, the MVD of the affine model associated with the second list of reference pictures is estimated from the first by imposing that the affine model parameters (generally tied to angles and scaling factors) are symmetric with respect to the MVD of the affine model associated with the first reference picture. To do so, the scaling factors and angle values ​​associated with the first affine model are computed. They are then transformed to satisfy the symmetry constraint, and then the affine mode parameters are inversely transformed to provide a second affine model of the current CU for bidirectional affine prediction of the CU. This takes the following form:

[0153] We assume that a four-parameter model is used for the block under consideration.

[0154]

number

[0155] Next, the rotation angle and scaling parameters are obtained as follows:

[0156]

number

[0157] Next, the angle and scaling factor are converted as follows:

[0158]

number

[0159] Here, k is a scaling factor that depends on the temporal distance between the current picture and its two reference pictures. Finally, the affine model parameters a' and b' for the second reference picture are easily calculated from a' and b'.

[0160] In the simplified version, only the scaling factor is changed, resulting in simpler scaling for MVd.

[0161] [Embodiment 5]: Combination of MMVD and Planar Motion Vector Prediction According to Embodiment 5, MMVD is used in combination with Planar Motion Vector Prediction (PMVP) mode.

[0162] This can take one of the following forms:

[0163] [A single MMVD that refines each of the MVs in the PMVP movement field] - Following the first basic method, in the case where the PMVP mode is used for the current CU, the MMVD motion vector difference is encoded. mmvd is applied to the PMVP-generated motion field, which is first generated from its generator MV. Thus, mmvd is used as an additive offset for each MV in the PMVP motion field. This is illustrated in the following figure, which shows a simplified decoding block diagram of the PMVD motion generation process.

[0164] [A single mmvd used to refine at least one MV, which is used to generate a PMVD motion field] - Following another variation, one motion vector difference is applied to at least one motion vector used to generate the planar motion field (generator MV). For example, a single MVD can be encoded and applied to the AL (top left) motion vector in Figure 28 before generating the planar MV field. This is illustrated in the following figure, which shows a simplified decoding block diagram of the PMVD motion generation process. - Following another variation, the MVD is encoded and applied to the BR motion vector before generating the planar MV field. - Following another variation, the MVD is encoded and applied to the AR motion vector before generating the planar MV field. - Following another variation, the MVD is encoded and applied to the BL motion vector before generating the plane MV field.

[0165] [Used to generate PMVD motion fields, used to refine several MVs, several MMVDs] - According to another embodiment, before generating the planar motion field, several MVDs are encoded and applied to one or more MVs from among the AL, BL, AR, and BR motion vectors. This is illustrated in the following figure, which shows a simplified decoding block diagram of the PMVD motion generation process. - According to another embodiment, when several MVDs are encoded and applied to multiple MVs among the AL, BL, AR, and BR motion vectors, the first MVD to be encoded is encoded as in MMVD. The next one is encoded in a differential manner based on the previously encoded one in order to limit the rate cost associated with the MVD encoding process. - Following the transformation, the set of allowable distances for MVD in the planar case will change compared to the MMVD tool currently used in VVC Draft 3. For example, the constrained range of allowable MVD distances may be changed. -According to the deformation, the number of allowed MVD orientations is changed compared to the existing MMVD system in VVC Draft 3. For example, when MMVD is used in combination with planar MV prediction, an enhanced set of MVD angles can be supported.

[0166] [Embodiment 6]: Combination of MMVD and regression-based motion vector field According to Embodiment 6, MMVD is used in combination with a regression-based six-parameter motion field introduced in the section of the regression MVF model. This can take one of the following forms:

[0167] [A single MMVD that refines each individual MV in the RMVF movement field] - Following the first basic method, the MMVD motion vector difference is encoded in the case where the RMVF mode is used for the current CU and applied to the RMVF-generated motion field. Thus, MVd is used as an additive offset to the RMVF motion field. A similar block diagram, as shown in Figure 29, can adequately illustrate the simplified process for generating the RMVF motion field according to this variation.

[0168] [A single mmvd used to refine at least one MV, which is used to generate an RMVF motion field] - Following another variation, one motion vector difference is applied to at least one motion vector used to generate the RMVF motion field. For example, a single MVD can be encoded and applied to the motion vectors on the left side of the current block before generating the RMVF motion field. - Following another variation, one MVD is encoded and applied to the upper neighbor motion vector before generating the regression-based MV field. A similar block diagram, as shown in Figure 30, can adequately illustrate the simplified process for generating the RMVF motion field according to this variation.

[0169] [Two mmvd files used to refine two MVs, which are used to generate the RMVF motion field] - According to another embodiment, two MVDs are encoded and applied to the upper and left MVs, respectively, before generating a regression-based motion field. - According to another embodiment, as described above, when two MVDs are encoded and applied, the first MVD to be encoded is encoded as in MMVD. The next one is encoded in a differential manner based on the first one in order to limit the rate cost associated with the MVD encoding process. - Following the transformation, the set of acceptable distances for MVD in the regression base case will change compared to the MMVD tool currently used in VVC Draft 3. For example, the constrained range of acceptable MVD distances may change. - According to the modification, the number of allowed MVD orientations is changed compared to the existing MMVD system in VVC Draft 3. For example, when MMVD is used in combination with RMVF, it can support an enhanced set of MVD angles. - Similar block diagrams, such as those shown in Figure 31, can adequately illustrate a simplified process for generating the RMVF motion field according to their variations.

[0170] [Embodiment 7]: Combination of MMVD with LIC In this embodiment, MMVD and LIC can be enabled together. In this case, as in the case of VVC Draft 3, the encoded motion vector difference is applied to the motion vector of the CU under consideration. Thus, as in the case of the current LIC tool, the LIC linear model parameters associated with the time prediction of the current CU are derived from the merge candidate used.

[0171] Following the transformation, if the magnitude of the MVD exceeds a certain threshold, the LIC flag for the current CU can be set to false, that is, LIC time prediction refinement can be deactivated for the current CU. In fact, intuitively, if the motion vector of the current CU differs significantly from the motion vector of the merge candidate used to derive the MV of the current CU, the LIC linear model used for the merge candidate CU can be irrelevant to the current CU.

[0172] For example, if the MVD size is 16 or 32 or greater, the LIC mode can be forced to 0 for the current CU.

[0173] Following further modifications, the thresholds mentioned above regarding the size of the MVD can depend on the current CU size.

[0174] [Embodiment 8]: Combination of MMVD with GBI In VVC Draft 3, when a CU is encoded in translational merge mode, its motion vector is derived from the selected merge candidate and its GBI index. This means that the GBI weights of the CU, which serve as a reference for deriving the current CU motion data, are used unchanged for the current CU as well, including in cases where MMVD is activated for the current CU.

[0175] In this current embodiment, the GBI weight adaptation for the current CU can be applied based on the size of the MVD used for the current merged CU.

[0176] For example, if the magnitude of the MVD exceeds a certain threshold, the GBI can be reset to the default GBI weights (1 / 2, 1 / 2) for the current CU. In fact, intuitively, if the motion vector of the current CU differs significantly from the motion vector of the merge candidate used to derive the MV of the current CU, the GBI weights used for the merge candidate CU can be independent of the current CU.

[0177] For example, if the size of the MVD is 16 or 32 or greater, the GBI weights can be forced to (1 / 2, 1 / 2) for the current CU.

[0178] Following further modifications, the thresholds mentioned above regarding the size of the MVD can depend on the current CU size.

[0179] [Embodiment 9]: MMVD and triangular motion partition According to the embodiment, the MMVD motion vector coding tool is used in combination with a triangular motion partitioning tool.

[0180] Following the first variation, one MVD is encoded for each triangular partition according to the MMVD VVC Draft 3 MVD encoding system.

[0181] Following another variation, a single MVD is encoded and used in common by two triangular motion partitions.

[0182] According to another variation, when two MVDs are encoded, the second one is encoded in a differential manner relative to the first. In that case, the second MVD can be constrained to a smaller range of acceptable size in the same way as for the affine case (see section 4 of Embodiment).

[0183] In a more advanced embodiment, in a case where one MVD is encoded for each partition, a second MVD is encoded in a conditional manner based on the relative values ​​of the motion vectors of the first and second partitions. For example, the second MVD can be constrained so that the second refined MV is not too close to the refined MV of the first triangular partition. In fact, if they are too close, the overall prediction of the current CU may behave very similarly to a normal translational motion compensated prediction of the entire rectangular block.

[0184] [Embodiment 10]: MMVD and Multiple Hypothesis (MH) Prediction In current video standards such as VVC Draft 3, a new prediction mode called multiple hypothesis essentially involves a combination of merge / skip time prediction blocks and intra prediction blocks. However, in VVC Draft 3, MMVD and MH prediction cannot be used together.

[0185] In this embodiment, it is made possible to use MMVD and multiple hypothesis prediction modes in a combined manner. Essentially, this involves applying the merge motion vector difference to the merge or skip candidate motion vectors used to derive motion information for the current coding unit.

[0186] The advantage of this embodiment is the further improved coding efficiency.

[0187] [Embodiment 11]: MMVD and Space-Time Motion Vector Prediction (STMVP) In VVC Draft 3, a motion vector prediction mode called STMVP is proposed, as described in the section on non-subblock spacetime merge motion vector predictors. An STMVP motion vector candidate is an additional merge candidate that can generally be part of the translational merge candidate list. Its essence lies in predicting a single motion vector for the current CU. Therefore, MMVD can be applied to STMVP candidates in a direct manner.

[0188] However, MMVD can also be applied to only one of the three spatial and temporal MV predictors used to compute the STMVP candidate.

[0189] According to the embodiment, the motion vector difference of the MMVD is applied to one or two spatial motion vector predictors before calculating the average motion vector between these two spatial motion vector predictors and the time motion vector predictor.

[0190] According to the embodiment, the MMVD motion vector difference is applied to the time motion vector predictor before calculating the average motion vector between this time MV predictor and one or two spatial motion vector predictors.

[0191] [Proposed Embodiments for Extending the Use of SMVD] Embodiment 12: SMVD replaced by MMVD in AMVP mode As described in the Symmetric MVC section, the SMVD motion vector coding mode is applicable only in AMVP mode, while the MMVD motion vector representation mode is applicable only in merge mode. In this embodiment, the codec design is harmonized. Only one of the motion vector coding modes, MMVD and SMVD, is proposed for the overall design to handle both the case of symmetric bidirectional motion and the case of small-magnitude motion vectors in relation to merging. The following two variations are proposed.

[0192] [Using mmvd syntax to encode MV differences in both MMVD and SMVD modes] - In AMVP mode, when symmetric mode is enabled, the motion vector difference is encoded using the MMVD MVD coding syntax shown in the MMVD Motion Vector Difference Coding Tool Description section. Therefore, when symmetric MVD mode is active for the CU under consideration, the conventional MVD coding method of VVC Draft 3 in AMVP is replaced.

[0193] [Using the VTM3 MVd syntax to encode the MV difference in both MMVD and SMVD modes] - In AMVP mode, when symmetric mode is on, the motion vector difference is encoded as currently done in VVC Draft 3 in AMVP. However, in merge cases, when MMVD mode is on, the MMVD motion vector difference is encoded in the AMVP way.

[0194] [Embodiment 13]: SMVD combined with an affine motion model According to the embodiment, the use of SMVD is extended to affine AMVP cases, which can take the following form:

[0195] The first CPMV of the affine CU under consideration is encoded according to the SMVD mode in the section on symmetric MVD, as in the conventional translational AMVP case. Next, the MVDs of the other CPMVs of the affine CU under consideration are encoded differentially with respect to the MVD of the first CPMV, as is currently done in the affine AMVP case. A corresponding simplified block diagram is shown in the figure below.

[0196] Following further modifications, the differential MVD of the second CPMV, and optionally the third CPMV, is differentially encoded with respect to the MVD of the first CPMV, but also under symmetric mode constraints. This allows for a reduction in the rate cost of the second and third CPMVs.

[0197] [Embodiment 14]: SMVD combined with triangular motion partition In AMVP, when triangular partitions are used, the symmetric MVD mode can be extended to the triangular partition case.

[0198] In such cases, in the embodiment, the two unidirectional MVDs of the first and second triangular partitions can be symmetrical to one another.

[0199] [Embodiment 15]: SMVD combined with multiple hypothesis prediction mode In this embodiment, multiple hypotheses can be used in AMVP mode in addition to merge mode. In this case, the SMVD mode can be used in combination with the multiple hypothesis prediction mode, in cases where the inter component of this inter / intra composite prediction uses dual prediction.

[0200] [Embodiment 16]: SMVD combined with a planar motion model In the embodiment, the planar motion model can be used in AMVP mode in addition to its current use in merge mode. In that case, the motion difference can be used and applied to the motion vectors surrounding the current CU and can be used to generate the motion field of the current CU.

[0201] Furthermore, the SMVD mode can be used in combination with this AMVP planar motion model case. In that case, several symmetric bidirectional motion vector differences can be applied to the motion vectors around the current CU, which are used to generate the planar MV field of the current CU. Such embodiments are similar to those in the section covering Embodiment 5, but are within the AMVP+ symmetric mode case.

[0202] [Embodiment 17]: SMVD combined with a regression-based motion model In some embodiments, a regression-based motion model can be used in AMVP mode, in addition to its current use in merge mode. In that case, a motion difference can be used and applied to the motion vectors surrounding the current CU, and can be used to generate a regression-based motion field of the current CU.

[0203] Furthermore, the SMVD mode can be used in combination with this AMVP regression-based motion model case. In that case, several symmetric bidirectional motion vector differences can be applied to the motion vectors around the current CU, which are used to generate the regression-based MV field of the current CU. Such embodiments are similar to those in the section covering Embodiment 5, but are within the AMVP+ symmetric mode case.

[0204] [Embodiment 18]: SMVD combined with ATMVP motion model Exclusive due to AMVP and merge In some embodiments, the ATMVP motion model can be used in AMVP mode in addition to its current use in merge mode. In that case, motion differences can be used and applied to the motion vectors included in the ATMVP predicted motion vector of the current CU.

[0205] Furthermore, the SMVD mode can be used in combination with this AMVP + ATMVP motion model case. In that case, several symmetric bidirectional motion vector differences can be applied to the ATMVP motion vectors derived for the current CU. Such embodiments are similar to those in the section covering Embodiment 3, but are within the AMVP + ATMVP mode case.

[0206] [Embodiment 19]: SMVD combined with spatial-time motion vector prediction (STMVP) In current video standards proposals, such as VVC Draft 3, the STMVP motion vector predictor (see the section on STMVP) is proposed only in merge cases. However, STMV can be enabled in AMVP mode. In that case, SMVD can be applied directly to the STMVP candidate.

[0207] However, SMVD can also be applied to only one of the three spatial and temporal MV predictors used to compute the STMVP candidate.

[0208] According to the embodiment, the SMVD motion vector difference is applied to one or two spatial motion vector predictors before calculating the average motion vector between these two spatial motion vector predictors and the time motion vector predictor.

[0209] According to the embodiment, the motion vector difference of SMVD is applied to the temporal motion vector predictor before calculating the average motion vector between this temporal MV predictor and one or two spatial motion vector predictors. 1.1.2. [Embodiment 20]: Modified SMVD mode for imposing symmetry on the final pair of bidirectional motion vectors. Translational case This section presents an embodiment in which the SMVD mode is modified such that the symmetry constraint is imposed on the final motion vectors used to bidirectionally predict a block in the translational AMVP mode rather than on the motion vector difference MVd. In fact, as explained in the section covering symmetric MVD, in SMVD, for the first motion vector of a pair of bidirectional MVs, the MVD is encoded and then the MVD of the second motion vector is inferred from the first MVD as a scalable, opposite MVD vector as a function of the temporal distance between the current picture and its reference picture. Here, in the proposed embodiment, the modified SMVD mode is such that the MVD of the first vector is encoded as in the existing SMVD. Next, the first motion vector is reconstructed as the sum of the MV predictor and the decoded motion vector difference. Finally, the reconstructed second motion vector is inferred from the reconstructed first motion vector as a scalable, opposite MV as a function of the temporal distance between the current picture and its reference picture.

[0210] Thus, the proposed new SMVD mode ensures that the predicted block is on the same line as the two reference blocks behind and in front of it. This approach is expected to provide improved coding efficiency compared to the existing SMVD in the section covering symmetric MVD.

[0211] [Embodiment 21]: Modified SMVD mode for imposing symmetry on the final affine model (rotation, scaling / zooming, and translation) According to an embodiment, the SMVD mode is used in combination with the affine AMVP. In that case, the symmetry constraint can be directly imposed on the fin MVD as well.

[0212] According to another approach, the symmetry constraint can be imposed on the affine model parameters (angles and scaling factors) in a similar way to the corresponding characteristics of the section regarding Embodiment 4 related to the combination of MMVD and affine.

[0213] An embodiment of method 3500 under the general aspects described herein is shown in FIG. 35. The method starts at start block 3501, and control proceeds to block 3510 to indicate a first motion mode through the syntax in the video bitstream. Control proceeds from block 3510 to block 3520 to indicate the use of a second motion mode through the presence of the syntax in the video bitstream, and if present, to include information related to the second motion mode. Control proceeds from block 3520 to block 3530 to encode the video block using the motion information corresponding to the first and second motion modes.

[0214] Another embodiment of method 3600 under the general aspects described herein is shown in FIG. 36. The method starts at start block 3601, and control proceeds to block 3610 to analyze the video bitstream as to whether the syntax indicates a first motion mode. Control proceeds from block 3610 to block 3620 to analyze the video bitstream as to whether the syntax indicates the presence of a second motion mode, and if present, to determine information related to the second motion mode. Control proceeds from block 3620 to block 3630 to obtain the motion information corresponding to the first motion mode. Control proceeds from block 3630 to block 3640 to decode the block using the motion information.

[0215] Figure 37 shows one embodiment of the apparatus 3700 for encoding, decoding, compressing, or decompressing video data using a simplified encoding mode based on a neighbor-sample-dependent parametric model. The apparatus comprises a processor 3710 which can be interconnected to memory 3720 through at least one port. Both the processor 3710 and memory 3720 may also have one or more additional interconnections to external connections.

[0216] The processor 3710 is also configured to insert information into or receive information in a bitstream, and to perform compression, encoding, or decoding using any of the described embodiments.

[0217] This application describes various embodiments, including tools, features, embodiments, models, and methods. Many of these embodiments are described in detail, and often in a manner that may sound restrictive, at least in order to illustrate their individual characteristics. However, this is for clarity in the description and does not limit the application or scope of those embodiments. In fact, all of the different embodiments can be combined and interchangeable to provide further embodiments. Furthermore, embodiments can also be combined and interchangeable with embodiments described in previous applications.

[0218] The embodiments described and intended in this application can be carried out in many different ways. Figures 3, 4, and 34 provide several embodiments, but other embodiments are intended, and the description of Figures 3, 4, and 34 does not limit the scope of embodiments. At least one of the embodiments generally relates to video encoding and decoding, and at least one other embodiment generally relates to transmitting a generated or encoded bitstream. These and other embodiments can be carried out as a computer-readable storage medium storing thereon instructions for encoding or decoding video data according to any of the methods, apparatus, or methods described, and / or a computer-readable storage medium storing thereon a bitstream generated according to any of the methods described.

[0219] In this application, the terms “reconstructed” and “decoded” are interchangeable, the terms “pixel” and “sample” are interchangeable, and the terms “image,” “picture,” and “frame” are interchangeable. While not always necessary, the term “reconstructed” is typically used on the encoder side, while the term “decoded” is typically used on the decoder side.

[0220] In this specification, various methods are described, each of which includes one or more steps or actions for achieving the described method. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of any particular steps and / or actions may be changed or combined.

[0221] Various methods and other embodiments described herein can be used to modify modules of the video encoder 100 and decoder 200, such as the intra-predictive, entropy coding, and / or decoding modules (160, 360, 145, 330) shown in Figures 3 and 4. Furthermore, these embodiments are not limited to VVC or HEVC and can be applied to other standards and recommendations, whether existing or future, as well as extensions to any such standards and recommendations (including VVC and HEVC). Unless otherwise noted or technically excluded, the embodiments described herein can be used individually or in combination.

[0222] Various numerical values ​​are used in this application. Certain values ​​are illustrative, and the embodiments described are not limited to these specific values.

[0223] Figure 3 illustrates encoder 100. While variations of encoder 100 are intended, encoder 100 will be described below without explaining all expected variations for clarity.

[0224] Before encoding, the video sequence may undergo a pre-encoding process (101) which involves, for example, applying a color conversion to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.

[0225] In encoder 100, the picture is encoded by encoder elements as described below. The encoded picture is partitioned (102) and processed in units of, for example, CUs. Each unit is encoded using, for example, either intra-mode or inter-mode. When a unit is encoded in intra-mode, it performs intra-prediction (160). In inter-mode, motion estimation (175) and compensation (170) are performed. The encoder determines either intra-mode or inter-mode to use to encode a unit (105), and indicates the intra / inter-determination, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting the prediction block from the original image block (110).

[0226] Subsequently, the predicted residual is transformed (125) and quantized (130). The quantized transformation coefficients, as well as the motion vector and other syntax elements, are entropi-encoded (145) to output the bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is encoded directly without the application of a transformation or quantization process.

[0227] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are inversely quantized (140) and inversely transformed (150) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (155) to reconstruct the image blocks. For example, an in-loop filter (165) is applied to the reconstructed picture to reduce encoding artifacts by performing deblocking / SAO (sample adaptive offset) filtering. The filtered image is stored in a reference picture buffer (180).

[0228] Figure 4 illustrates a block diagram of the video decoder 200. In the decoder 200, the bitstream is decoded by the decoder elements as described below. The video decoder 200 generally performs a decoding path which is the opposite of the encoding path described in Figure 3. The encoder 100 also generally performs video decoding as part of encoding the video data.

[0229] In particular, the decoder input includes a video bitstream that can be generated by the video encoder 100. The bitstream is first entropi-decoded (230) to obtain transformation coefficients, motion vectors, and other encoded information. Picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture (235) according to the partitioning information of the decoded picture. The transformation coefficients are inversely quantized (240) and inversely transformed (250) to decode the prediction residuals. The image blocks are reconstructed by combining the decoded prediction residuals and prediction blocks (255). Prediction blocks can be obtained (270) from intra-predictions (260) or motion-compensated predictions (i.e., inter-predictions) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).

[0230] The decoded picture can then undergo a post-decoding process (285), such as an inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or a reverse remapping process that performs the reverse of the remapping process performed in the pre-encoding process (101). The post-decoding process can use metadata derived in the pre-encoding process and signaled in the bitstream.

[0231] Figure 34 illustrates a block diagram of an example of a system in which various embodiments and models are implemented. System 1000 can be materialized as a device comprising various components described below and configured to perform one or more of the embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and servers. The elements of System 1000 can be materialized individually or in combination as a single integrated circuit (IC), multiple ICs, and / or individual components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 1000 are distributed across multiple ICs and / or individual components. In various embodiments, System 1000 is communicably coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, System 1000 is configured to perform one or more of the embodiments described in this document.

[0232] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein, for example, to implement various embodiments described in this document. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the Art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040 which may include, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drives, and / or optical disk drives, and may include non-volatile and / or volatile memory. The storage device 1040 may, in non-limiting examples, include internal storage devices, mounted storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0233] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 1030 can include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device can include one or both of an encoding module and a decoding module. In addition, the encoder / decoder module 1030 can be implemented as a separate element of the system 1000 or can be incorporated within the processor 1010 as a combination of hardware and software, as is known to those skilled in the art.

[0234] To execute the various aspects described in this document, the program code loaded on the processor 1010 or the encoder / decoder 1030 can be stored in the storage device 1040 and then loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 can store one or more of various items during the execution of the processes described in this document. Such stored items can include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstream, matrix, variable, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0235] In some embodiments, the internal memory of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory outside the processing device (for example, the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be memory 1020 and / or storage device 1040, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used, for example, to store the operating system of a television. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations for MPEG-2 (MPEG stands for Moving Picture Experts Group, also known as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Multipurpose Video Coding, a new standard being developed by JVET, a Joint Video Experts Team).

[0236] Inputs to the elements of system 1000 can be provided through various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals, for example, transmitted wirelessly by a broadcasting station; (ii) a component (COMP) input terminal (or a set of COMP input terminals); (iii) a universal serial bus (USB) input terminal; and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other examples not shown in Figure 34 include composite video.

[0237] In various embodiments, the input devices of block 1130 have associated input processing elements, as known in the Art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also said to select a signal, or band-limit a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting it to a narrower frequency band to select a signal frequency band (e.g., which may be called a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments may include one or more elements for performing these functions, e.g., frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to the baseband. In one embodiment of the set-top box, the RF section and its associated input processing elements perform frequency selection by receiving an RF signal transmitted over a wired (e.g., cable) medium, filtering it to a desired frequency band, down-converting it, and filtering it again. Various embodiments rearrange the order of the elements described above (and others), remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0238] In addition, the USB and / or HDMI terminals may include their respective interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented, if necessary, in a separate input processing IC or within the processor 1010. Similarly, aspects of USB or HDMI interface processing can be implemented, if necessary, in a separate interface IC or within the processor 1010. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in combination with memory and storage elements to process the data stream, if necessary, for presentation on an output device.

[0239] Various elements of system 1000 can be provided within an integrated housing. Within the integrated housing, the various elements are interconnected and can transmit data between them using internal buses known in the art, such as an inter-IC (I2C) bus, wiring, and printed circuit boards, with appropriate connection configurations.

[0240] System 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, transceivers configured to transmit and receive data on the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0241] In various embodiments, data is streamed to system 1000 or otherwise provided using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received on a communication channel 1060 and a communication interface 1050 adapted for Wi-Fi communication. The communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide the streaming data to system 1000 using a set-top box that distributes data over the HDMI connection of input block 1130. Yet another embodiment provides the streaming data to system 1000 using the RF connection of input block 1130. As shown above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, e.g., cellular networks or Bluetooth networks.

[0242] System 1000 can provide output signals to a variety of output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. In various embodiments, the display 1100 may include, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 may be for a television, tablet, laptop, cell phone (mobile phone), or other device. The display 1100 may also be integrated with other components (for example, in a smartphone) or separate (for example, an external monitor for a laptop). In various examples of embodiments, the other peripheral devices 1120 may include one or more of a standalone digital video disc (or digital multi-purpose disc) (DVR in either term), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of System 1000. For example, the disc player performs the function of playing the output of system 1000.

[0243] In various embodiments, control signals are signaled between the system 1000 and the display 1100, speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices can be communicatively coupled to the system 1000 via dedicated connections through their respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to the system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speaker 1110 can be integrated into a single unit along with other components of the system 1000 in an electronic device, such as a television. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (TCon) chip.

[0244] For example, if the RF section of input 1130 is part of a separate set-top box, the display 1100 and speaker 1110 can, alternatively, be separated from one or more of the other components. In various embodiments where the display 1100 and speaker 1110 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0245] The embodiments can be implemented by computer software implemented by the processor 1010, by hardware, or by a combination of hardware and software. In a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be any type appropriate for the technical environment and, in a non-limiting example, can be implemented using any appropriate data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 can be any type appropriate for the technical environment and, in a non-limiting example, can include one or more of microprocessors, general-purpose computers, dedicated computers, and processors based on multicore architectures.

[0246] Various implementations include decoding. As used in this application, “decoding” can encompass all or part of a process performed on, for example, a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process typically includes one or more processes performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by a decoder in the various implementations described in this application.

[0247] As further examples, in one embodiment, “decoding” refers only to entropy decoding; in another embodiment, “decoding” refers only to differential decoding; and in yet another embodiment, “decoding” refers to a combination of entropy decoding and differential decoding. Whether the term “decoding process” is intended to refer specifically to a subset of operations or to a broader decoding process in general will be clear from the context of the specific description and will be well understood by those skilled in the art.

[0248] Various implementations include encoding. As with the above-mentioned description of “decoding,” “encoding,” as used in this application, can encompass all or part of the processes performed on, for example, an input video sequence to produce an encoded bitstream. In various embodiments, such processes typically include one or more processes performed by an encoder, e.g., partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also include, or alternatively, processes performed by the encoders of the various implementations described in this application.

[0249] As further examples, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of differential encoding and entropy encoding. Whether the term “encoding process” is intended to refer specifically to a subset of operations or to a broader encoding process in general will be clear from the context of the specific description and will be well understood by those skilled in the art.

[0250] Note that, as used herein, syntax elements are descriptive terms; therefore, they do not preclude the use of other syntax element names.

[0251] Please understand that when a diagram is presented as a flow chart, it also provides a block diagram of the corresponding device. Similarly, please understand that when a diagram is presented as a block diagram, it also provides a flow chart of the corresponding method / process.

[0252] Various embodiments may refer to parametric models or rate-distortion optimization. In particular, during the encoding process, the balance or trade-off between rate and distortion is usually considered, often under constraints of computational complexity. This can be measured through rate-distortion optimization (RDO) metrics, or through least mean squares (LMS), mean absolute error (MAE), or other such measurements. Rate-distortion optimization is usually formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different methods exist for solving rate-distortion optimization problems. For example, a method may be based on extensive testing, involving a complete evaluation of the encoding costs of all encoding options, including all mode or encoding parameter values ​​under consideration, as well as the associated distortions of the reconstructed signals after encoding and decoding. Faster methods can also be used to reduce encoding complexity, in particular by using approximate distortion calculations based on predicted or predicted residual signals rather than reconstructed ones. A mixture of these two methods can also be used, for example, by using approximate distortion for only some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods perform optimization by utilizing one of various techniques, but optimization does not necessarily provide a complete evaluation of both the encoding cost and the associated distortions.

[0253] The implementations and embodiments described herein may be implemented, for example, by methods or processes, apparatus, software programs, data streams, or signals. Even if a feature described is described only in relation to a single form of implementation (for example, only as a method), the implementation of that feature may also be implemented in other forms (for example, apparatus or programs). Apparatus may be implemented, for example, by appropriate hardware, software, and firmware. Methods may be implemented by processors, for example, which generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the transmission of information between end users.

[0254] References to “one embodiment” or “embodiment,” or “one implementation” or “implementation,” and other variations thereof, mean that certain features, structures, and characteristics described in relation to the embodiments are included in at least one embodiment. Therefore, appearances of the phrases “in one embodiment” or “in one embodiment,” or “in one implementation” or “in implementation,” and any other variations, appearing in various places throughout this application, do not necessarily all refer to the same embodiment.

[0255] In addition, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.

[0256] Furthermore, this application may refer to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.

[0257] In addition, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory), one or more of these. Furthermore, "receiving" is generally included in various ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0258] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", please understand that the use of any of the following " / ", "and / or", and "at least one" is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or both options (A and B). As a further example, in the cases of "A, B, and / or C", and "at least one of A, B, and C", such phrasing is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or only the second and third enumerated options (B and C), or all three options (A, B, and C). This can be extended to the number of items listed, as will be obvious to those skilled in the art and those skilled in the relevant technical fields.

[0259] Furthermore, as used herein, the term "signal" refers, among other things, to indicating something to the corresponding decoder. For example, in one embodiment, the encoder signals one particular of several transformations, encoding modes, or flags. Thus, in this embodiment, the same transformation, parameter, or mode is used on both the encoder and decoder sides. For example, the encoder can send certain parameters to the decoder so that the decoder can use the same particular parameters (explicit signaling). Conversely, if the decoder already has certain parameters or other information, signaling can be used without transmission, simply allowing the decoder to know and select those parameters (implicit signaling). Bit saving is achieved in various embodiments by avoiding the transmission of either actual function. It should be understood that signaling can be achieved in various ways. For example, in various embodiments, one or more syntax elements and flags, etc., are used to signal information to the corresponding decoder. The above concerns the verb form of the word "signal," but the word "signal" can also be used as a noun in this specification.

[0260] As will be apparent to those skilled in the art, implements can generate a variety of signals, for example, formatted to carry information that can be stored or transmitted. The information may include, for example, instructions for performing the method, or data generated by one of the described implements. For example, a signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier using the encoded data stream. The information carried by the signal may be, for example, analog information or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0261] We have described numerous embodiments across various claim categories and types. Features of these embodiments can be provided individually or in any combination. Furthermore, embodiments may include one or more of the following features, devices, or aspects across various claim categories and types, individually or in any combination. ● A process or device that uses MMVD and / or SMVD along with an affine motion model. ● A process or device that uses MMVD and / or SMVD, along with alternative time-motion vector prediction. ● A process or device that uses MMVD and / or SMVD along with bidirectional optical flow. ● A process or device that uses MMVD and / or SMVD, along with a motion vector reference to the current picture. ● A process or device using MMVD and / or SMVD, along with generalized biprediction. ● Processes or devices that use MMVD and / or SMVD, along with local illumination compensation. ● Processes or devices that use MMVD and / or SMVD, along with multiple hypothesis merge / intra combination mode. ● A process or device that uses MMVD and / or SMVD, either with MVD merging or with Ultimate MV Expression. ● A process or device that uses MMVD and / or SMVD, along with a motion vector field for each subblock based on a regression model. ● Processes or devices that use symmetric MVD, MMVD, and / or SMVD, encoded with biprediction, along with just one MVD. ● A process or device that uses MMVD and / or SMVD along with a triangular partition. ● A process or device that combines MMVD with dual prediction. ● A process or device that combines MMVD and CPR. ● A process or device that combines MMVD and ATMVP. ● A process or device that combines MMVD and affine mode. ● A process or device that combines MMVD with planar motion vector prediction. ● A process or device that combines MMVD and regression-based motion vector fields. ● A process or device that combines MMVD with GBI. ● A process or device that combines MMVD with spatial-time motion vector prediction. ● In AMVP mode, a process or device that replaces SMVD with MMVD. ● A process or device that combines SMVD with an affine motion model. ● A process or device that combines SMVD with a planar motion model. ● A process or device that combines SMVD with a regression-based motion model. ● A process or device that combines SMVD with ATMVP motion models. ● A process or device that combines SMVD with spatial-time motion vector prediction. ● A process or device that modifies the SMVD mode to impose symmetry on bidirectional motion vectors in the translation case. A process or device that modifies the SMVD mode to impose symmetry on the final affine model, including rotation, scaling, zooming, or translation. ● A bitstream or signal containing one or more of the described syntax elements or variations thereof. ● A bitstream or signal including a syntax for signaling information generated according to any of the described embodiments. ● Creating and / or transmitting and / or receiving and / or decoding in accordance with any of the embodiments described. ● A method, process, apparatus, medium for storing instructions, medium for storing data, or signal, according to any of the embodiments described. ● Insertion of syntax elements into the signaling that allow the decoder to determine the encoding mode in a manner corresponding to that used by the encoder. ● Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntactic elements or variations thereof. ● A television, set-top box, cell phone, tablet, or other electronic device that performs a conversion method according to any of the embodiments described. ● A television, set-top box, cell phone, tablet, or other electronic device that performs a conversion method determination according to any of the embodiments described and displays the resulting image (for example, using a monitor, screen, or other type of display). A television, set-top box, cell phone, tablet, or other electronic device that selects, limits, or adjusts (e.g., using a tuner) channels to receive a signal containing an encoded image, and performs a conversion method according to one of the embodiments described. ● A television, set-top box, cell phone, tablet, or other electronic device that receives a signal containing an encoded image wirelessly (for example, using an antenna) and performs a conversion method. [Industrial applicability]

[0262] This invention can be used for communications. [Explanation of Symbols]

[0263] 102, 102a~102d WTRU 104, 113 RAN 106 Core Network 108 PSTN 110 Internet 112 Other networks 160a, 160b, 160c e-node B 118 processors 120 Transceivers (Transmitters and Receivers) 122 Antenna

Claims

1. A method for decoding, A step of obtaining information indicating that the video block is divided into a first triangular prediction unit and a second triangular prediction unit, The steps include: decoding the video block using the first motion vector for the first triangular prediction unit and the second motion vector for the second triangular prediction unit; The first motion vector is, The steps include acquiring first motion vector difference information associated with a merge mode (MMVD mode) with motion vector difference for the first triangular prediction unit, The steps include: adding the first refined motion vector determined from the first motion vector difference information to the first base motion vector to determine the first motion vector; The method determined by [the specified method].

2. The method of claim 1, wherein the first motion vector difference information includes a direction index and a distance index.

3. The second motion vector is, The steps include acquiring second motion vector difference information associated with the MMVD mode for the second triangular prediction unit, and The second step of determining the second motion vector by adding the second refined motion vector determined from the second motion vector difference information to the second base motion vector. The method of claim 1, determined by...

4. The method of claim 1, wherein a single refined motion vector is used to determine both the first motion vector and the second motion vector, the single refined motion vector being added to the first base motion vector for the first triangular prediction unit and to the second base motion vector for the second triangular prediction unit.

5. The method of claim 3, wherein the second motion vector difference information is differentially encoded with respect to the first motion vector difference information of the first triangular prediction unit.

6. The method of claim 5, wherein the second motion vector difference information is limited to representing a range of magnitude smaller than the range of magnitude permitted for the first motion vector difference information of the first triangular prediction unit.

7. The method of claim 3, wherein the set of permissible values ​​for the second motion vector difference information is limited based on a predetermined neighborhood relationship with the first motion vector.

8. A method for encoding, The steps include dividing the video block into a first triangular prediction unit and a second triangular prediction unit, The steps include encoding the video block using the first motion vector for the first triangular prediction unit and the second motion vector for the second triangular prediction unit, and Equipped with, The first step of determining the first motion vector by adding the first refined motion vector to the first base motion vector in a merge mode with motion vector difference (MMVD mode), The first triangular prediction unit is subjected to the step of encoding first motion vector difference information corresponding to the first refined motion vector, and A method for encoding that further enhances this functionality.

9. The method of claim 8, wherein the first motion vector difference information includes a direction index and a distance index.

10. The steps include determining the second motion vector by adding the second refined motion vector to the second base motion vector in a merge mode with motion vector difference (MMVD mode), The steps include encoding second motion vector difference information corresponding to the second refined motion vector for the second triangular prediction unit, and The method of claim 8, further comprising the above.

11. The method of claim 8, wherein a single refined motion vector is used to determine both the first motion vector and the second motion vector, and the single refined motion vector is added to the first base motion vector for the first triangular prediction unit and to the second base motion vector for the second triangular prediction unit.

12. The method of claim 10, wherein the second motion vector difference information is differentially encoded with respect to the first motion vector difference information of the first triangular prediction unit.

13. The method of claim 12, wherein the second motion vector difference information is limited to representing a range of magnitude smaller than the range of magnitude permitted for the first motion vector difference information of the first triangular prediction unit.

14. The method of claim 10, wherein the set of permissible values ​​for the second motion vector difference information is limited based on a predetermined neighborhood relationship with the first motion vector.

15. A decoding device comprising one or more processors and at least one memory coupled to the processors, The one or more processors described above are: A step of obtaining information indicating that the video block is divided into a first triangular prediction unit and a second triangular prediction unit, The steps include: decoding the video block using the first motion vector for the first triangular prediction unit and the second motion vector for the second triangular prediction unit; The configuration is configured to perform the first motion vector, The steps include acquiring first motion vector difference information associated with a merge mode (MMVD mode) with motion vector difference for the first triangular prediction unit, The steps include: adding the first refined motion vector determined from the first motion vector difference information to the first base motion vector to determine the first motion vector; A device determined by [the specified method / system].

16. The apparatus of claim 15, wherein the first motion vector difference information includes a direction index and a distance index.

17. The second motion vector is, The steps include acquiring second motion vector difference information associated with the MMVD mode for the second triangular prediction unit, and The second step of determining the second motion vector by adding the second refined motion vector determined from the second motion vector difference information to the second base motion vector. The apparatus of claim 15, determined by...

18. The apparatus of claim 15, wherein a single refined motion vector is used to determine both the first motion vector and the second motion vector, and the single refined motion vector is added to the first base motion vector for the first triangular prediction unit and to the second base motion vector for the second triangular prediction unit.

19. The apparatus of claim 17, wherein the second motion vector difference information is differentially encoded with respect to the first motion vector difference information of the first triangular prediction unit.

20. The apparatus of claim 19, wherein the second motion vector difference information is limited to representing a range of magnitude smaller than the range of magnitude permitted for the first motion vector difference information of the first triangular prediction unit.

21. The apparatus of claim 17, wherein the set of permissible values ​​for the second motion vector difference information is limited based on a predetermined neighbor relationship with the first motion vector.

22. An encoding device comprising one or more processors and at least one memory coupled to the processors, The one or more processors described above are: The steps include dividing the video block into a first triangular prediction unit and a second triangular prediction unit, The steps include encoding the video block using the first motion vector for the first triangular prediction unit and the second motion vector for the second triangular prediction unit, and The one or more processors are configured to perform the following: The first step of determining the first motion vector by adding the first refined motion vector to the first base motion vector in a merge mode with motion vector difference (MMVD mode), The first triangular prediction unit is subjected to the step of encoding first motion vector difference information corresponding to the first refined motion vector, and A device further configured to carry out the following.

23. The apparatus of claim 22, wherein the first motion vector difference information includes a direction index and a distance index.

24. The one or more processors described above are: The steps include determining the second motion vector by adding the second refined motion vector to the second base motion vector in a merge mode with motion vector difference (MMVD mode), The steps include encoding second motion vector difference information corresponding to the second refined motion vector for the second triangular prediction unit, and The apparatus of claim 22, further configured to carry out the following:

25. The apparatus of claim 22, wherein a single refined motion vector is used to determine both the first motion vector and the second motion vector, and the single refined motion vector is added to the first base motion vector for the first triangular prediction unit and to the second base motion vector for the second triangular prediction unit.

26. The apparatus of claim 24, wherein the second motion vector difference information is differentially encoded with respect to the first motion vector difference information of the first triangular prediction unit.

27. The apparatus of claim 26, wherein the second motion vector difference information is limited to representing a range of magnitude smaller than the range of magnitude permitted for the first motion vector difference information of the first triangular prediction unit.

28. The apparatus of claim 24, wherein the set of permissible values ​​for the second motion vector difference information is limited based on a predetermined neighborhood relationship with the first motion vector.