Multiple predictor candidates for motion compensation

By constructing a set of multiple predictor candidates and selecting the best predictor based on criteria, the method enhances affine motion compensation in video coding, improving encoding and decoding efficiency.

JP2026071237APending Publication Date: 2026-04-28INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2026-01-09
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing video coding techniques, particularly in HEVC and JEM, face challenges in efficiently selecting predictor candidates for affine motion compensation, leading to suboptimal coding efficiency in video encoding and decoding processes.

Method used

An improved method and apparatus for video coding and decoding that involves constructing a set of multiple predictor candidates, determining corresponding control point motion vectors, and selecting the best predictor candidate based on criteria such as rate distortion, to enhance affine motion compensation.

Benefits of technology

This approach improves coding efficiency by allowing for better selection of predictor candidates, resulting in enhanced video encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071237000001_ABST
    Figure 2026071237000001_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for selecting a predictor candidate from a set of multiple predictor candidates for motion compensation, based on a motion model such as an affine model. [Solution] In the embodiment, the predictor candidates are selected from a set based on a motion model for each of a plurality of predictor candidates, which may be based on criteria such as rate distortion cost. The corresponding motion field is determined based on one or more corresponding control point motion vectors for the block to be encoded or decoded. The corresponding motion field in the embodiment identifies the motion vectors used for predicting the subblocks of the block to be encoded or decoded.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [1] At least one embodiment of the present invention relates in general to a method or apparatus for, for example, video coding or decoding, and more specifically to a method or apparatus for selecting a predictor candidate from a set of multiple predictor candidates for motion compensation based on a motion model, such as an affine model, for a video encoder or video decoder. [Background technology]

[0002] [2] To achieve high compression efficiency, image and video coding schemes generally employ predictions, including motion vector predictions, and transforms to leverage spatial and temporal redundancy in video content. Generally, intra or inter-predictions are used to leverage intra or inter-frame correlations, and the difference between the original image and the predicted image, often expressed as prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the reverse process corresponding to entropy coding, quantization, transformation, and prediction.

[0003] [3] A recent addition to high-compression techniques is the use of motion models based on affine modeling. In particular, affine modeling is used for motion compensation for encoding and decoding video pictures. Generally, affine modeling is a model that uses at least two parameters, such as two control point motion vectors (CPMVs) that represent the motion at each corner of a block of a picture, which allows for the derivation of a motion field about the entire block of the picture to simulate rotation and similarity ratio (zoom). [Overview of the Initiative] [Means for solving the problem]

[0004] [4] According to a general aspect of at least one embodiment, a method for video coding is presented, comprising: determining a set of predictor candidates having a plurality of predictor candidates with respect to a block to be coded in a picture; selecting a predictor candidate from the set of predictor candidates; determining one or more corresponding control point motion vectors with respect to a block with respect to the predictor candidate selected from the set of predictor candidates; determining a corresponding motion field with respect to the selected predictor candidate based on one or more corresponding control point motion vectors, wherein the corresponding motion field identifies motion vectors used for predicting subblocks of the block to be coded; coding the block based on the corresponding motion field with respect to the predictor candidate selected from the set of predictor candidates; and coding an index with respect to the predictor candidate selected from the set of predictor candidates.

[0005] [5] According to another general aspect of at least one embodiment, a method for video decoding is presented, comprising: receiving an index corresponding to a particular predictor candidate with respect to a block to be decoded in a picture; determining one or more corresponding control point motion vectors with respect to the block to be decoded with respect to the particular predictor candidate; determining a corresponding motion field based on a motion model that identifies motion vectors used for predicting subblocks of the block to be decoded based on one or more corresponding control point motion vectors with respect to the particular predictor candidate; and decoding the block based on the corresponding motion field.

[0006] [6] According to another general aspect of at least one embodiment, an apparatus for video coding is presented, comprising: means for determining a set of predictor candidates having a plurality of predictor candidates with respect to a block to be coded in a picture; means for selecting a predictor candidate from the set of predictor candidates; means for determining a corresponding motion field with respect to the selected predictor candidate, based on one or more corresponding control point motion vectors, that identifies a corresponding motion field based on a motion model with respect to the selected predictor candidate, which is used for predicting a subblock of the block to be coded; means for coding a block based on the corresponding motion field with respect to the predictor candidate selected from the set of predictor candidates; and means for coding an index with respect to the predictor candidate selected from the set of predictor candidates.

[0007] [7] According to another general aspect of at least one embodiment, a device for video decoding is presented, comprising: means for receiving an index corresponding to a particular predictor candidate with respect to a block to be decoded in a picture; means for determining one or more corresponding control point motion vectors with respect to a block to be decoded with respect to a particular predictor candidate; means for determining a corresponding motion field based on a motion model that identifies motion vectors used for predicting subblocks of the block to be decoded based on one or more corresponding control point motion vectors with respect to a particular predictor candidate; and means for decoding a block based on the corresponding motion field.

[0008] [8] According to another general aspect of at least one embodiment, an apparatus for video encoding is provided, comprising one or more processors and at least one memory. The one or more processors are configured to determine a set of predictor candidates having a plurality of predictor candidates with respect to a block to be encoded in a picture; select a predictor candidate from the set of predictor candidates; determine one or more corresponding control point motion vectors with respect to a block with respect to the predictor candidate selected from the set of predictor candidates; determine a corresponding motion field with respect to the selected predictor candidate based on the one or more corresponding control point motion vectors, which identifies a motion vector used for predicting a subblock of the block to be encoded, with respect to the selected predictor candidate; encode the block based on the corresponding motion field with respect to the predictor candidate selected from the set of predictor candidates; and encode an index with respect to the predictor candidate selected from the set of predictor candidates. At least one memory relates to at least temporarily storing the encoded block and / or the encoded index.

[0009] [9] According to another general aspect of at least one embodiment, an apparatus for video decoding is provided, comprising one or more processors and at least one memory. The one or more processors are configured to receive an index corresponding to a particular predictor candidate with respect to a block to be decoded in a picture, determine one or more corresponding control point motion vectors with respect to the block to be decoded with respect to the particular predictor candidate, determine a corresponding motion field based on a motion model that identifies motion vectors used for predicting subblocks of the block to be decoded based on the one or more corresponding control point motion vectors with respect to the particular predictor candidate, and decode the block based on the corresponding motion field. At least one memory relates to at least temporary storage of the decoded block.

[0010]

[10] According to another general aspect of at least one embodiment, a method for video coding is presented, comprising: determining a set of predictor candidates with respect to a block to be coded in a picture; determining one or more corresponding control point motion vectors with respect to a block for each of a plurality of predictor candidates in the set of predictor candidates; determining a corresponding motion field for each of the plurality of predictor candidates based on one or more corresponding control point motion vectors, based on a motion model with respect to each of the plurality of predictor candidates in the set of predictor candidates; evaluating a plurality of predictor candidates according to one or more criteria and based on the corresponding motion field; selecting a predictor candidate from the plurality of predictor candidates based on the evaluation; and coding a block based on the predictor candidate selected from the set of predictor candidates.

[0011]

[11] According to another general aspect of at least one embodiment, a method for video decoding is presented, comprising obtaining an index corresponding to a selected predictor candidate with respect to a block in a picture to be decoded. The selected predictor candidate is selected by an encoder by: determining a set of predictor candidates with respect to a block in a picture to be encoded; determining one or more corresponding control point motion vectors with respect to the block to be encoded for each of a plurality of predictor candidates in the set of predictor candidates; determining a corresponding motion field based on a motion model with respect to each of the plurality of predictor candidates in the set of predictor candidates, based on one or more corresponding control point motion vectors for each of the plurality of predictor candidates; evaluating the plurality of predictor candidates according to one or more criteria and based on the corresponding motion field; selecting a predictor candidate from the plurality of predictor candidates based on the evaluation; and encoding an index for the selected predictor candidate from the set of predictor candidates. The method further comprises decoding a block based on the index corresponding to the selected predictor candidate.

[0012]

[12] According to another general aspect of at least one embodiment, the method further comprises evaluating a plurality of predictor candidates according to one or more criteria and based on corresponding motion fields for each of the plurality of predictor candidates, and selecting a predictor candidate from the plurality of predictor candidates based on the evaluation.

[0013]

[13] According to another general aspect of at least one embodiment, the apparatus further comprises means for evaluating a plurality of predictor candidates according to one or more criteria and based on corresponding motion fields for each of the plurality of predictor candidates, and means for selecting a predictor candidate from the plurality of predictor candidates based on the evaluation.

[0014]

[14] According to another general aspect of at least one embodiment, the one or more criteria are based on a rate distortion determination corresponding to one or more of the plurality of predictor candidates in a set of predictor candidates.

[0015]

[15] According to another general aspect of at least one embodiment, decoding or encoding a block based on a corresponding motion field comprises, respectively, decoding or encoding a predictor indicated by a motion vector based on a predictor for a sub-block.

[0016]

[16] According to another general aspect of at least one embodiment, the set of predictor candidates comprises spatial candidates and / or temporal candidates of a block to be encoded or decoded.

[0017]

[17] According to another general aspect of at least one embodiment, the motion model is an affine model.

[0018]

[18] According to another general aspect of at least one embodiment, the corresponding motion field for each position (x, y) within a block to be encoded or decoded is

Number

[0019]

[19] According to another general aspect of at least one embodiment, the number of spatial candidates is five or more.

[0020]

[20] According to another general aspect of at least one embodiment, one or more additional control point motion vectors are added to determine a corresponding motion field based on a function of one or more corresponding control point motion vectors determined.

[0021]

[21] According to other general aspects of at least one embodiment, the function includes 1) mean, 2) weighted mean, 3) unique mean, 4) average, 5) median, or 6) one or more unidirectional parts of 1) to 6) above of the determined one or more corresponding control point motion vectors.

[0022]

[22] According to other general aspects of at least one embodiment, a non-temporary computer-readable medium containing data content generated according to any of the methods or apparatus described above is presented.

[0023]

[23] According to other general aspects of at least one embodiment, a signal is provided comprising video data generated according to any of the methods or apparatus described above.

[0024]

[24] One or more embodiments of the present disclosure also provide a computer-readable storage medium containing instructions for encoding or decoding video data according to any of the methods described above. Embodiments of the present disclosure also provide a computer-readable storage medium containing a bitstream generated according to any of the methods described above. Embodiments of the present disclosure also provide a method and apparatus for transmitting a bitstream generated according to any of the methods described above. Embodiments of the present disclosure also provide a computer program product containing instructions for performing any of the methods described above. [Brief explanation of the drawing]

[0025] [Figure 1]

[25] A block diagram of an embodiment of the HEVC (High Efficiency Video Coding) video encoder is shown. [Figure 2A]

[26] An example image showing HEVC reference sample generation. [Figure 2B]

[27] An example image showing the intra-prediction direction in HEVC. [Figure 3]

[28] A block diagram of an embodiment of the HEVC video decoder is shown. [Figure 4]

[29] An example of the coded tree unit (CTU) and coded tree (CT) concepts for representing a compressed HEVC picture is shown. [Figure 5]

[30] An example is shown in which a coding tree unit (CTU) is divided into coding units (CU), prediction units (PU), and transformation units (TU). [Figure 6]

[31] An example of an affine model is given as a motion model used in the Joint Search Model (JEM). [Figure 7]

[32] An example of a 4x4 subCU-based affine motion vector field used in the Joint Search Model (JEM) is shown. [Figure 8A]

[33] Examples of motion vector prediction candidates for affine interCU are shown. [Figure 8B]

[34] An example of a motion vector prediction candidate in the affine merger mode is shown. [Figure 9]

[35] An example of spatial derivation of affine control point motion vectors in an affine combined mode motion model is shown. [Figure 10]

[36] An example of a method relating to a general aspect of at least one embodiment is shown. [Figure 11]

[37] Other examples of methods relating to a general aspect of at least one embodiment are shown. [Figure 12]

[38] Other examples of methods relating to a general aspect of at least one embodiment are shown. [Figure 13]

[39] Other examples of methods relating to a general aspect of at least one embodiment are shown. [Figure 14]

[40] An example of a known process for evaluating the affine merging mode of interCU in JEM is shown. [Figure 15]

[41] An example of the process for selecting predictor candidates in the affine merger mode in JEM is shown. [Figure 16]

[42] An example of an affine motion field propagated by an affine merger prediction candidate located to the left of the current block being encoded or decoded is shown. [Figure 17]

[43] An example of an affine motion field propagated by affine merger predictor candidates located above and to the right of the current block being encoded or decoded is shown. [Figure 18]

[44] An example of a predictor candidate selection process relating to a general aspect of at least one embodiment is shown. [Figure 19]

[45] An example of a process for constructing a set of multiple predictor candidates relating to a general aspect of at least one embodiment is shown. [Figure 20]

[46] An example of the process for deriving the CPMV of the upper left and upper right corners for each predictor candidate is shown, relating to a general aspect of at least one embodiment. [Figure 21]

[47] An example of an extended set of spatial predictor candidates relating to a general aspect of at least one embodiment is shown. [Figure 22]

[48] ​​Another example of a process for constructing a set of multiple predictor candidates, relating to a general aspect of at least one embodiment. [Figure 23]

[49] Another example of a process for constructing a set of multiple predictor candidates, relating to a general aspect of at least one embodiment. [Figure 24]

[50] An example of how a temporary candidate may be used for a predictor candidate is shown, relating to a general aspect of at least one embodiment. [Figure 25]

[51] An example of a process relating to a general aspect of at least one embodiment is shown, in which average CPMV motion vectors calculated from stored CPMV candidates are added to the final set of CPMV candidates. [Figure 26]

[52] A block diagram of an example apparatus in which various embodiments of the embodiment can be realized is shown. [Modes for carrying out the invention]

[0026]

[53] Figure 1 shows a typical High Efficiency Video Coding (HEVC) encoder 100. HEVC is a compression standard developed by the Joint Team for Video Coding (JCT-VC) (see, for example, “ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services-Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265”).

[0027]

[54] In HEVC, to encode a video sequence having one or more pictures, the pictures are partitioned into one or more slices, each slice may contain one or more slice segments. The slice segments are organized into coding units, prediction units, and transformation units.

[0028]

[55] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” and “coded” may be used interchangeably, and the terms “picture” and “frame” may be used interchangeably. In many but not always cases, the term “reconstructed” is used on the encoder side and “decoded” is used on the decoder side.

[0029]

[56] The HEVC specification distinguishes between “blocks” and “units,” where a “block” refers to a specific area within a sample array (e.g., luma, Y), and a “unit” includes juxtaposed blocks of all encoded color components (Y, Cb, Cr, or monochromatic), syntactic elements, and predictive data associated with the block (e.g., motion vectors).

[0030]

[57] For encoding, a picture is partitioned into square coding tree blocks (CTBs) of a configurable size, and a contiguous set of coding tree blocks is grouped into slices. A coding tree unit (CTU) contains the CTB of the encoded color components. The CTB is the root of a quadtree, which is divided into coding blocks (CBs), and the coding blocks may be partitioned into one or more prediction blocks (PBs), which form the root of a quadtree, which is divided into transform blocks (TBs). Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) contains a set of prediction units (PUs) and transform units (TUs) in a tree structure, where the PUs contain prediction information for all color components, and the TUs contain residual coding syntactic structures for each color component. The sizes of the CBs, PBs, and TBs of the color components are the same as the corresponding CUs, PUs, and TUs. In this application, the term “block” may be used to refer to any of, for example, CTU, CU, PU, ​​TU, CB, PB, and TB. In addition, the term "block" can also be used to refer to macroblocks and partitions specified in H.264 / AVC or other video encoding standards, or more generally, to data arrays of various sizes.

[0031]

[58] In a typical encoder 100, a picture is encoded by encoder elements as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using either intra-mode or inter-mode. If a CU is encoded in intra-mode, the CU performs intra-prediction (160). In inter-mode, motion estimation (175) and compensation (170) are performed. The encoder decides whether to use intra-mode or inter-mode to encode the CU (105), and the intra / inter decision is indicated by a prediction mode flag. The prediction residual is calculated by subtracting the predicted blocks from the original image blocks (110).

[0032]

[59] In intra-mode, the CU is predicted from reconstructed adjacent samples within the same slice. A set of 35 intra-prediction modes is available in HEVC, including DC prediction mode, planar prediction mode, and 33 angular prediction modes. The intra-prediction reference is reconstructed from adjacent rows and columns to the current block. The reference extends horizontally and vertically to twice the block size using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra-prediction, the reference sample may be copied along the direction indicated by the angular prediction mode.

[0033]

[60] The available intra-prediction modes for the current block can be encoded using two different options. If the applicable mode is included in three most likely modes (MPMs), the mode is indicated by an index in the MPM list. Otherwise, the mode is indicated by a fixed-length binarization of the mode index. The three most likely modes are derived from the intra-prediction modes of the adjacent blocks above and to the left.

[0034]

[61] With respect to interCUs, the corresponding coded block is further partitioned into one or more prediction blocks. Interpretation is performed at the PB level, and the corresponding PU contains information about how the interpretation was performed. Motion information (i.e., motion vectors and reference picture indices) can be communicated in two ways, namely "merged mode" and "advanced motion vector prediction (AMVP)".

[0035]

[62] In merge mode, the video encoder or decoder collects a list of candidates based on already encoded blocks, and the video encoder notifies the index of one of the candidates in the list. On the decoder side, the motion vector (MV) and reference picture index are reconstructed based on the notified candidate.

[0036]

[63] The set of possible candidates in the merge mode consists of spatial neighboring candidates, temporal candidates, and generated candidates. FIG. 2A shows the positions of five spatial candidates {a1, b1, b0, a0, b2} for the current block 210, where a0 and a1 are to the left of the current block, and b1, b0, b2 are above the current block. For each candidate position, the availability is checked in the order of a1, b1, b0, a0, b2, and then the redundancy within the candidates is removed.

[0037]

[64] The motion vectors at the collocated positions in the reference picture can be used for the derivation of temporal candidates. The available reference pictures are selected on a slice basis and indicated in the slice header, and the reference index for the temporal candidates is set to i ref = 0. If the POC distance (td) between the picture of the collocated PU and the reference picture that is the prediction source of the collocated PU is the same as the distance (tb) between the current picture and the reference picture including the collocated PU, the collocated motion vector mv col can be directly used as a temporal candidate. Otherwise, the scaled motion vector tb / td * mv col is used as a temporal candidate. Depending on where the current PU is located, the collocated PU is determined by the sample position at the bottom right or center of the current PU.

[0038]

[65] The maximum number N of merge candidates is explicitly stated in the slice header. If the number of merge candidates is greater than N, only the first N - 1 spatial and temporal candidates are used. Otherwise, if the number of merge candidates is less than N, the set of candidates is filled up to the maximum number N by the candidates generated as a combination of the existing candidates or null candidates. The candidates used in the merge mode can be referred to as "merge candidates" in this application.

[0039]

[66] If CU indicates skip mode, the available index for merge candidates is shown only if the list of merge candidates is greater than 1, and no further information is encoded for CU. In skip mode, the motion vector is applied without residual updates.

[0040]

[67] In AMVP, the video encoder or decoder collects a list of candidates based on motion vectors determined from already encoded blocks. The video encoder then notifies the index in the candidate list to identify the motion vector predictor (MVP) and the motion vector difference (MVD). On the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. The available reference picture index is also explicitly encoded in the PU syntax for AMVP.

[0041]

[68] In AMVP, only two spatial motion candidates are selected. The first spatial motion candidate is selected from the left position {a0, a1}, and the second candidate is selected from the upper position {b0, b1, b2}, but the search order shown in the two sets is maintained. If the number of motion vector candidates is not equal to 2, a time MV candidate may be included. If the set of candidates is still not completely filled, a zero motion vector is used.

[0042]

[69] If the reference picture index of the spatial candidate corresponds to the reference picture index for the current PU (i.e., they use the same reference picture index, or both use a long-term reference picture independently of the reference picture list), the spatial candidate motion vector is used directly. If, otherwise, the reference picture is short-term, the candidate motion vector is scaled according to the distance (tb) between the current picture and the reference picture of the current PU and the distance (td) between the current picture and the reference picture of the spatial candidate. Candidates used in AMVP mode may be referred to in this application as "AMVP candidates".

[0043]

[70] For simplicity of description, blocks tested in "combine" mode on the encoder side or decoded in "combine" mode on the decoder side are referred to as "combine" blocks, and blocks tested in AMVP mode on the encoder side or decoded in AMVP mode on the decoder side are referred to as "AMVP" blocks.

[0044]

[71] Figure 2B shows a typical motion vector representation using AMVP. For the current block 240 being encoded, the motion vector (MV) is estimated by motion estimation. current ) can be obtained. Motion vector (MV) from block 230 on the left. left ) and motion vector (MV) from block 220 above above ) using MV left and MV above From motion vector predictor MVP current It can be selected as such. Then MVD current =MV current -MVP current The difference in motion vectors can be calculated as follows.

[0045]

[72] Motion-compensated prediction may be performed using one or two reference pictures for prediction. In a P slice, only a single prediction reference may be used for interpretation, enabling uniprediction on the prediction block. In a B slice, two reference picture lists are available, and uniprediction or biprediction may be used. In biprediction, one reference picture is used from each of the reference picture lists.

[0046]

[73] In HEVC, the accuracy of motion information for motion compensation is one-quarter of a sample with respect to the luma component (also referred to as 1 / 4 PEL or 1 / 4 PEL) and one-eighth of a sample with respect to the chroma component (also referred to as 1 / 8 PEL) in the case of a 4:2:0 configuration. A 7-tap or 8-tap interpolation filter is used for interpolating the fractional sample positions, i.e., 1 / 4, 1 / 2, and 3 / 4 of the full sample position can be processed with respect to luma in both the horizontal and vertical directions.

[0047]

[74] The predicted residuals are then transformed (125) and quantized (130). The quantized transformed coefficients, as well as the motion vector and other syntactic elements, are entropy coded (145) to output a bitstream. The encoder may skip the transformation and apply quantization directly to the untransformed residual signal on a 4x4 TU base. The encoder may omit both the transformation and quantization, i.e., the residuals are coded directly without the application of a transformation or quantization process. In direct PCM coding, no prediction is applied, and coded unit samples are coded directly into a bitstream.

[0048]

[75] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are inversely quantized (140) and inversely transformed (150) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are joined (155) to reconstruct the image blocks. In-loop filters (165) are applied to the reconstructed picture to perform deblocking / SAO (sample adaptive offset) filtering, for example, to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).

[0049]

[76] Figure 3 shows a block diagram of a typical HEVC video decoder 300. In a typical decoder 300, the bitstream is decoded by the decoder elements as described later. The video decoder 300 generally performs an encoding path and a reciprocal decoding path, as shown in Figure 1, in which video decoding is performed as part of the encoding of the video data.

[0050]

[77] In particular, the input to the decoder includes a bitstream that may be generated by the video encoder 100. The bitstream is first entropy-decoded (330) to obtain transformation coefficients, motion vectors, and other encoded information. The transformation coefficients are inversely quantized (340) and inversely transformed (350) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (355) to reconstruct the image blocks. The predicted blocks may be obtained (370) from intra-predictions (360) or motion-compensated predictions (i.e., inter-predictions) (375). As described above, AMVP and combined-mode techniques may be used to derive motion vectors for motion compensation, which may use interpolation filters to compute interpolated values ​​for sub-integer samples of the reference block. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0051]

[78] As described above, in HEVC, motion-compensated time prediction is used to take advantage of the redundancy present between consecutive pictures of video. For this purpose, a motion vector is associated with each prediction unit (PU). As described above, each CTU is represented by an encoding tree in the compression region. This is a quadtree partition of the CTU, where each leaf is called an encoding unit (CU), as also shown in Figure 4 for CTUs 410 and 420. Each CU is then assigned several intra or interprediction parameters as prediction information. For this purpose, a CU may be spatially partitioned into one or more prediction units (PUs), each PU being assigned some prediction information. The intra or intercoding mode is assigned at the CU level. These concepts are further illustrated in Figure 5 for typical CTUs 500 and CU 510.

[0052]

[79] In HEVC, each PU is assigned one motion vector. This motion vector is used for motion-compensated time prediction of the PU under consideration. Thus, in HEVC, the motion model linking the prediction block and its reference block simply consists of transformations or calculations based on the reference block and the corresponding motion vector.

[0053]

[80] To improve HEVC, the Joint Video Exploration Team (JVET) is developing reference software and / or documented JEM (Joint Exploration Model). In one version of JEM (e.g., “Algorithm Description of Joint Exploration Test Model 5”, document JVET-E1001_v2, Joint Video Exploration Team, ISO / IEC JTC1 / SC29 / WG11, 5th Meeting, January 12-20, 2017, Geneva, Switzerland), several additional motion models are supported to improve time prediction. For this purpose, the PU may be spatially partitioned into sub-PUs, and the model may be used to assign each sub-PU to a dedicated motion vector.

[0054]

[81] In more recent versions of JEM (e.g., “Algorithm Description of Joint Exploration Test Model 2”, document JVET-B1001_v3, ISO / IEC JTC1 / SC29 / WG11 Joint Video Exploration Team, 2nd Meeting, February 20-26, 2016, San Diego, USA), it is not explicitly stated that CUs are to be divided into PUs or TUs. Instead, a more flexible CU size may be used, with some motion data directly assigned to each CU. In this new codec design in newer versions of JEM, CUs may be divided into subCUs, and motion vectors may be calculated for each subCU of the divided CU.

[0055]

[82] One of the new motion models introduced in JEM is the use of an affine model as the motion vector to represent the motion vector in CU. The motion model used is shown in Figure 6 and is represented by Equation 1 as shown below. The affine motion field has the following motion vector component values ​​for each position (x,y) in the block 600 under consideration in Figure 6:

number

[0056]

[83] To reduce complexity, motion vectors are calculated for each of the 4x4 subblocks (subCUs) of the CU700 under consideration, as shown in Figure 7. The affine motion vectors are calculated from the control point motion vectors for each central position of each subblock. The resulting MVs are expressed with an accuracy of 1 / 16 Pell. As a result, the compensation of the coding unit in affine mode lies in the motion-compensated prediction of each subblock by its own motion vector. These motion vectors for each subblock are shown as arrows for each subblock in Figure 7.

[0057]

[84] In JEM, seeds are stored within corresponding 4x4 subblocks, so the affine mode can only be used for CUs having widths and heights up to 4 (so that each seed has its own independent subblock). For example, in a 64x4 CU there is only one left subblock to store the top-left and bottom-left seeds, and in a 4x32 CU there is only one top subblock for the top-left and top-right seeds, and in JEM it is impossible to properly store seeds in such thin CUs. According to our proposal, since seeds are stored individually, it is possible to handle such thin CUs having widths or heights equal to 4.

[0058]

[85] Referring again to the example in Figure 7, the affine CU is defined by an associated affine model consisting of three motion vectors called affine model seeds, as motion vectors from the top-left, top-right, and bottom-left corners of the CU (v0, v1, and v2 in Figure 7). This affine model then allows for the computation of the affine motion vector field within the CU (the black motion vectors in Figure 7), which is performed on a 4x4 subblock basis. In JEM, these seeds are attached to the top-left, top-right, and bottom-left 4x4 subblocks of the CU under consideration. In the proposed solution, the affine model seeds are stored separately as motion information attached to the entire CU (e.g., IC flags). Thus the motion model is decoupled from the motion vectors used for actual motion compensation at the 4x4 block level. This new storage may allow for the storage of the complete motion vector field at the 4x4 subblock level. This also allows for the use of affine motion compensation with respect to blocks with a width or height of size 4.

[0059]

[86] Affine motion compensation can be used in JEM in two ways: affine inter (AF_INTER) mode and affine combined mode. These are described in the following sections.

[0060]

[87] Affine Inter (AF_INTER) Mode: In affine inter mode, CUs in AMVP mode larger than 8x8 can be predicted. This is indicated by a flag in the bitstream. Generating the affine motion field for that interCU involves determining the control point motion vector (CPMV) obtained by the decoder by adding the motion vector difference and the control point motion vector prediction (CPMVP). The CPMVP is a pair of motion vector candidates selected from sets (A, B, C) and (D, E) shown in Figure 8A, respectively, with respect to the current CU800 to be encoded or decoded.

[0061]

[88] Affine merge mode: In affine merge mode, a CU level flag indicates whether the merged CU uses affine motion compensation. If so, the first available neighbor CU encoded in affine mode is selected from an ordered set of candidate positions A, B, C, D, E in Figure 8B with respect to the current CU880 being encoded or decoded, except that this ordered set of candidate positions in JEM is the same as the spatial neighbor candidates in merge mode in HEVC shown in Figure 2A and described above.

[0062]

[89] Once the first adjacent CU in affine mode is obtained, three CPMVs are obtained from the upper left, upper right, and lower left corners of the adjacent affine CU.

number

number

number

number

[0063]

[90] Current CU control point motion vector

number

number

[0064]

[91] Accordingly, a general aspect of at least one embodiment aims to improve the performance of the affine merge mode in JEM so that the compensation performance of the video codec under consideration may be improved. Accordingly, in at least one embodiment, an extended and improved affine motion compensation apparatus and method are presented for, for example, an encoded unit encoded in affine merge mode. The proposed extended and improved affine mode includes evaluating a plurality of predictor candidates in affine merge mode.

[0065]

[92] As described above, in the current JEM, among the surrounding CUs, a first neighboring CU encoded in affine merger mode is selected to predict the affine motion model associated with the current CU being encoded or decoded. That is, a first neighboring CU candidate from the ordered set (A, B, C, D, E) of Figure 8B encoded in affine mode is selected to predict the affine motion model of the current CU.

[0066]

[93] Thus, at least one embodiment selects an affine merger prediction candidate that provides the best coding efficiency when coding the current CU in affine merger mode, rather than using only one of the first in the ordered set as described above. Thus, the improvement of this embodiment is, at a general level, for example • Constructing a set of multiple affine merger predictor candidates that are likely to provide a good candidate set for predicting the affine motion model of the CU (regarding encoders / decoders), • (Regarding encoders / decoders) Select one predictor from the configured set for the current CU's control point motion vector, and / or, • (Regarding encoders / decoders) Notify / decode the index of the current control point motion vector predictor. It is equipped with.

[0067]

[94] Accordingly, Figure 10 shows a typical encoding method 1000 relating to a general aspect of at least one embodiment. In 1010, method 1000 determines a set of predictor candidates having a plurality of predictor candidates with respect to a block to be encoded in a picture. In 1020, method 1000 selects a predictor candidate from the set of predictor candidates. In 1030, method 1000 determines one or more corresponding control point motion vectors with respect to a block with respect to the predictor candidate selected from the set of predictor candidates. In 1040, method 1000 determines a corresponding motion field based on a motion model with respect to the selected predictor candidate, based on one or more corresponding control point motion vectors, with respect to the selected predictor candidate, where the corresponding motion field identifies the motion vectors used for predicting a subblock of the block to be encoded. In 1050, method 1000 encodes a block based on the corresponding motion field with respect to the predictor candidate selected from the set of predictor candidates. In 1060, method 1000 encodes an index relating to a predictor candidate selected from a set of predictor candidates.

[0068]

[95] Figure 11 shows another typical encoding method 1100 relating to a general aspect of at least one embodiment. In 1110, method 1100 determines a set of predictor candidates with respect to the block to be encoded in the picture. In 1120, method 1100 determines one or more corresponding control point motion vectors with respect to the block for each of the multiple predictor candidates in the set of predictor candidates. In 1130, method 1100 determines a corresponding motion field based on a motion model with respect to each of the multiple predictor candidates in the set of predictor candidates, based on one or more corresponding control point motion vectors for each of the multiple predictor candidates. In 1140, method 1100 evaluates multiple predictor candidates according to one or more criteria and based on the corresponding motion fields. In 1150, method 1100 selects a predictor candidate from the multiple predictor candidates based on the evaluation. In 1160, method 1100 encodes an index relating to the predictor candidate selected from the set of predictor candidates.

[0069]

[96] Figure 12 shows a typical decoding method 1200 relating to a general aspect of at least one embodiment. In 1210, method 1200 receives an index corresponding to a specific predictor candidate with respect to a block to be decoded in the picture. In various embodiments, the specific predictor candidate is selected in the encoder, and the index allows for the selection of one of several predictor candidates. In 1220, method 1200 determines one or more corresponding control point motion vectors with respect to the block to be decoded with respect to the specific predictor candidate. In 1230, method 1200 determines a corresponding motion field based on one or more corresponding control point motion vectors with respect to the specific predictor candidate. In various embodiments, the motion field is based on a motion model, and the corresponding motion field identifies motion vectors used for predicting subblocks of the block to be decoded. In 1240, method 1200 decodes the block based on the corresponding motion field.

[0070]

[97] Figure 13 shows another typical decoding method 1300 relating to a general aspect of at least one embodiment. In 1310, method 1300 obtains an index corresponding to a selected predictor candidate with respect to the block to be decoded in the picture. As also shown in 1310, the selected predictor candidate is selected by the encoder by determining a set of predictor candidates with respect to the block to be encoded in the picture; determining one or more corresponding control point motion vectors with respect to the block to be encoded for each of the multiple predictor candidates in the set of predictor candidates; determining a corresponding motion field based on a motion model with respect to each of the multiple predictor candidates in the set of predictor candidates, based on one or more corresponding control point motion vectors for each of the multiple predictor candidates; evaluating the multiple predictor candidates according to one or more criteria and based on the corresponding motion field; selecting a predictor candidate from the multiple predictor candidates based on the evaluation; and encoding an index relating to the selected predictor candidate from the set of predictor candidates. In 1320, method 1300 decodes the block based on the index corresponding to the selected predictor candidate.

[0071]

[98] Figure 14 details an embodiment of process 1400 used to predict the affine motion field of the current CU to be encoded or decoded in an existing affine merge mode in JEM. The input 1401 to process 1400 is the current encoded unit from which the affine motion field of a subblock is to be generated, as shown in Figure 7. In 1410, the affine merge CPMV with respect to the current block is obtained using selected predictor candidates, as described above with respect to Figures 6, 7, 8B, and 9. The derivation of these predictor candidates will be described in more detail later with respect to Figure 15.

[0072] As a result, at 1420, the motion vectors of the upper left and upper right control points are

number

number

[0073]

[0100] In at least one implementation, a residual flag is used. At 1450, a flag is activated indicating that the encoding was performed with residual data (noResidual=0). At 1460, the current CU is fully encoded and reconstructed (with residuals), resulting in a corresponding RD cost. Subsequently, a flag indicating that the encoding was performed without residual data is deactivated (1480, 1485, noResidual=1), and the process returns to 1460, where the CU is encoded (without residuals), resulting in a corresponding RD cost. The lowest RD cost between the previous two (1470, 1475) indicates whether the residuals need to be encoded or not (usually or skipped). Method 1400 terminates at 1499. This best RD cost is then subjected to competition with other encoding modes. Rate distortion determination is described in more detail below.

[0074]

[0101] Figure 15 details an embodiment of process 1500 used to predict one or more control points of the affine motion field of the current CU. This involves searching for CUs that are encoded / decoded in affine mode among the spatial locations (A, B, C, D, E) in Figure 8B (1510, 1520, 1530, 1540, 1550). If none of the searched spatial locations are encoded in affine mode, a variable indicating the number of candidate locations, e.g., numValidMergeCand, is set to 0 (1560). Otherwise, a first location corresponding to a CU in affine mode is selected (1515, 1525, 1535, 1545, 1555). Process 1500 then involves calculating control point motion vectors that will later be used to generate the affine motion field assigned to the current CU, and setting numValidMergeCand to 1 (1580). This control point calculation proceeds as follows: The CU containing the selected location is determined. This is one of the adjacent CUs of the current CU, as mentioned above. Next, with respect to Figure 9, as mentioned above, three CPMVs from the upper left, upper right, and lower left corners within the selected adjacent CU.

number

number

number

number

number

[0075]

[0102] The inventors recognize that one aspect of the existing affine merger process described above involves systematically utilizing one and only motion vector predictor to propagate the affine motion field from surrounding informal (i.e., already encoded or decoded) and adjacent CUs to the current CU. In various situations, the inventors further recognize that this aspect may be disadvantageous, for example, because it does not select an optimal motion vector predictor. Furthermore, this predictor selection consists only of the first informal and adjacent CU encoded in affine mode in an ordered set (A, B, C, D, E), as already described above. In various situations, the inventors further recognize that this limited selection may be disadvantageous, for example, because a better predictor may be available. Therefore, the existing process in the current JEM does not take into account that several possible informal and adjacent CUs around the current CU could also have used affine motion, and that CUs other than the first CU that have been found to have used affine motion may be better predictors for the motion information of the current CU.

[0076]

[0103] Therefore, the inventors recognize the potential advantages in several methods of improving the prediction of current CU affine motion vectors that are not utilized by existing JEM codecs. According to a general aspect of at least one embodiment, such advantages provided in the motion model of the present invention are found, as described below and shown in Figures 16 and 17.

[0077]

[0104] In both Figures 16 and 17, the current CU to be encoded or decoded is the large central one, which is 1610 in Figure 16 and 1710 in Figure 17, respectively. Two potential predictor candidates correspond to positions A and C in Figure 8B and are shown as predictor candidate 1620 in Figure 16 and 1720 in Figure 17, respectively. In particular, Figure 16 shows the potential motion field of the current block 1610 to be encoded or decoded when the selected predictor candidate is in the left position (position A in Figure 8B). Similarly, Figure 17 shows the potential motion field of the current block 1710 to be encoded or decoded when the selected predictor candidate is in the upper right position (i.e., position C in Figure 8B). As shown in the illustrative figures, depending on which affine merger predictor is selected, various sets of motion vectors with respect to the current CU can be generated with respect to the subblock. Therefore, the inventors recognize that optimizing one or more criteria, such as rate distortion (RD), between these two candidates may help improve the coding / decoding performance of current CUs in affine merged mode.

[0078]

[0105] Therefore, one general aspect of at least one embodiment involves selecting a better motion predictor candidate from a set of candidates to derive the CPMV of the current CU to be encoded or decoded. On the encoder side, the candidates used to predict the current CPMV are selected according to a rate-distortion cost criterion, according to one aspect of one typical embodiment. Their index is then encoded in the output bitstream for the decoder, according to other aspects of other typical embodiments.

[0079]

[0106] In other embodiments of typical models, a set of candidates may be configured in the decoder, and predictors may be selected from this set in the same manner as on the encoder side. In such embodiments, it is not necessary for the index to be encoded in the output bitstream. Other embodiments of the decoder avoid configuring a set of candidates, or at least avoid selecting predictors from a set similar to that of the encoder, and simply decode the index corresponding to the selected candidate from the bitstream and derive the corresponding associated data.

[0080]

[0107] According to other embodiments of other typical embodiments, the CPMVs used herein are not limited to the two upper-right and upper-left positions of the current CU being encoded or decoded, as shown in Figure 6. Other embodiments may have, for example, just one vector or more than two vectors, and the positions of these CPMVs may be other corner positions or any position inside or outside the current block, for example, the center of a 4x4 subblock of a corner or the position(s) of an interior corner of a 4x4 subblock of a corner, as long as it is possible to derive a motion field.

[0081]

[0108] In a typical embodiment, the set of potential candidate predictors investigated is identical to the set of locations (A, B, C, D, E) used to obtain CPMV predictors in the existing affine merger mode in JEM, as shown in Figure 8B. Figure 18 details one typical selection process 1800 for selecting the best candidates to predict the affine motion model of the current CU, according to a general aspect of this embodiment. However, other embodiments use a set of predictor locations that may contain fewer or more elements in the set, unlike A, B, C, D, E.

[0082]

[0109] As shown in 1801, the input to this typical embodiment 1800 is also information about the current CU to be encoded or decoded. In 1810, a set of multiple affine merger predictor candidates is constructed according to the algorithm 1500 of Figure 15 described above. The algorithm 1500 of Figure 15 includes collecting all adjacent positions (A, B, C, D, E) shown in Figure 8A, corresponding to the informal CU encoded in affine mode, and forming a set of candidates for predicting the affine motion of the current CU. Thus, the process 1800 does not terminate when an informal affine CU is found, but rather stores all possible candidates for affine motion model propagation from the informal CU to the current CU for all of the multiple motion predictor candidates in the set.

[0083]

[0110] Once the process in Figure 15 is complete, as shown in 1810 of Figure 18, process 1800 of Figure 18 calculates the predicted CPMV of the upper left and upper right corners from each candidate in the set provided in 1810 in 1820. This process in 1820 is further detailed and shown in Figure 19.

[0084]

[0111] Again, Figure 19 shows details of 1820 in Figure 18, including loops over each candidate determined and discovered from the preceding step (1810 in Figure 18). For each affine merger predictor candidate, a CU containing the candidate's spatial position is determined. Then, with respect to each reference list L0 and L1 (at the base of the B slice), control point motion vectors useful for generating the motion field of the current CU are used.

number

number

[0085]

[0112] Once the process in Figure 19 is complete, the process returns to Figure 18, where loop 1830 is performed over each affine merger predictor candidate. This may, for example, select the CPMV candidate that yields the lowest rate distortion cost. Within loop 1830 over each candidate, another loop 1840, similar to the process shown in Figure 14, is used to encode the current CU using each CPMV candidate as described above. The algorithm in Figure 14 terminates when all candidates have been evaluated, and its output may contain an index of the best predictor. As described above, for example, the candidate with the minimum rate distortion cost may be selected as the best predictor. Various embodiments use the best predictor to encode the current CU, and certain embodiments also encode an index of the best predictor.

[0086]

[0113] One example of determining rate distortion costs is well known to those skilled in the art, RD cost =D+λ×R Defined as follows, in the formula, D represents the distortion (generally the L2 distance) between the original block and the reconstructed block obtained by encoding and decoding the current CU using the candidates under consideration, R represents the rate cost, for example, the number of bits generated by encoding the current block using the candidates under consideration, and λ represents the rate target when the video sequence is being encoded.

[0087]

[0114] Other typical embodiments are described below. These typical embodiments aim to further improve the coding performance of the affine merge mode by expanding the set of affine merge candidates compared to existing JEMs. These typical embodiments can be implemented similarly on both the encoder and decoder sides to expand the set of candidates. Thus, in one non-limiting embodiment, several additional predictor candidates may be used to constitute a set of multiple affine merge candidates. The additional candidates may be drawn from additional spatial locations such as A'2110 and B'2120 surrounding the current CU2100, as shown in Figure 21. Other embodiments use yet additional spatial locations along or adjacent to one of the edges of the current CU2100.

[0088]

[0115] Figure 22 shows a typical algorithm 2200 corresponding to an embodiment using the additional spatial locations A'2110 and B'2120 as shown in Figure 21 and described above. For example, in 2210-2230 of Figure 22, algorithm 2200 includes testing a new candidate location A' if location A is not a valid affine merger prediction candidate (e.g., not within a CU encoded in affine mode). Similarly, for example, in 2240-2260 of Figure 22, location B' is also tested if location B does not provide any valid candidate (e.g., not within a CU encoded in affine mode). Other aspects of the typical process 2200 for constructing a set of affine merger candidates are essentially the same as those shown and described earlier in Figure 19.

[0089]

[0116] In another typical embodiment, existing merger candidate positions are considered first before newly added positions are evaluated. The added positions are evaluated only if the set of candidates contains fewer candidates than the maximum number of merger candidates, such as 5 or 7. The maximum number may be predetermined or variable. This typical embodiment is illustrated by a typical algorithm 2300 in Figure 23.

[0090]

[0117] In other typical embodiments, additional candidates, called transient candidates, are added to the set of predictor candidates. These transient candidates may be used, for example, if no spatial candidates are found, as described above, or, in the variation example, if the size of the set of affine merged candidates does not reach its maximum, also as described above. Other embodiments use transient candidates before adding spatial candidates to the set. For example, transient candidates for predicting the control point motion vector of the current CU may be obtained from one or more reference pictures available to or used for the current picture. A transient candidate may be adopted, for example, at a position corresponding to the lower-right adjacent CU of the current CU in each of the reference pictures. This corresponds to candidate position F2410 with respect to the current CU2400 being encoded or decoded, as shown in Figure 24.

[0091]

[0118] In one embodiment, for example, for each reference picture in each reference picture list, an affine flag associated with the block at position F2410 in Figure 24 in the reference picture under consideration is tested. If true, the corresponding CU contained in that reference picture is added to the current set of affine merger candidates.

[0092]

[0119] In a further example of change, the temporary candidate is obtained from a reference picture at a spatial position corresponding to the upper left corner of the current CU2400. This position corresponds to candidate position G2420 in Figure 24.

[0093]

[0120] In a further variation example, a temporary candidate is obtained from a reference picture at the position corresponding to the adjacent CU in the lower right. Then, if the set of candidates contains fewer candidates than the pre-fixed maximum number of merged candidates, such as 5 or 7, a temporary candidate corresponding to the upper left corner G2420 of the current CU is obtained. In other embodiments, temporary candidates are obtained from one or more reference pictures at positions corresponding to different (other than G2420) positions of the current CU2400, or to other (other than F2410) adjacent CUs of the current CU2400.

[0094]

[0121] In addition, a typical derivation process for control point motion vectors based on temporary candidates proceeds as follows: For each temporary candidate included in the configured set, the block (tempCU) containing the temporary candidate in its reference picture is identified. Then, three CPMVs located in the upper left, upper right, and lower left corners of the identified temporary CU are identified.

number

number

number

[0095]

[0122] Other typical embodiments include adding a pair of average control point motion vectors, calculated as a function of the control point motion vectors derived from each candidate. A typical process is detailed here by a typical algorithm 2500 shown in Figure 25. Loop 2510 is used for each affine merger predictor candidate in a set configured with respect to the reference picture list under consideration.

[0096]

[0123] Subsequently, in 2520, for each reference picture list Lx that is successively equal to L0 and then L1 (in the case of a B slice), if the current candidate has a valid CPMV with respect to list Lx, • Pair of motion vectors

number

number

number

number

number

number

number

number

number

number

number

[0097]

[0124] Using algorithm 2500 and / or other embodiments, the set of affine merger candidates is further enhanced to include mean movement information calculated from the CPMV derived for each candidate inserted into the set of candidates according to the embodiments described above, as described in the preceding section.

[0098]

[0125] Since several candidates may yield the same CPMV with respect to the current CU, the average candidate described above may yield a weighted average pair of CPMV motion vectors. In fact, the process described above calculates the average of the CPMVs collected up to that point, regardless of their uniqueness in the complete set of CPMVs. Thus, an example of variation in this embodiment lies in again adding other candidates to the set of CPMV predictor candidates. This lies in adding the average CPMV of a set of its own collected CPMVs (separate from the weighted average CPMV described above). This provides further candidate CPMVs to the set of predictor candidates for generating the affine motion field of the current CU.

[0099]

[0126] For example, consider a situation where the following five spatial candidates (L, T, TR, BL, TL) are all available and affine. However, the three leftmost positions (L, BL, TL) are in the same adjacent CU. For each spatial position, a candidate CPMV can be obtained. Then, the first mean is equal to the sum of these five CPMVs divided by 5 (some may be identical). In the second mean, only different CPMVs are considered, so only the three leftmost positions (L, BL, TL) are considered once, and the second mean is equal to the sum of the three different CPMVs (L, T, TR) divided by 3. In the first mean, extra CPMVs are added three times, and a greater weight is assigned to the extra CPMVs. Using the formula, the mean 1 = (L + T + TR + BL + TL) / 5, and since L = BL = TL, the mean 1 = (3 * (L+T+TL) / 5, and the mean is 2 = (L+T+TL) / 3.

[0100]

[0127] The two candidate averages described above are bidirectional, provided that the candidates under consideration hold motion vectors for a reference image in List 0 and another image in List 1. In other variation examples, it is possible to add a unidirectional average. From the weighted average and unique average, four unidirectional candidates can be constructed by individually taking motion vectors from List 0 and List 1.

[0101]

[0128] One advantage of the typical candidate set expansion method described in this application is the increased diversity in the set of candidate control point motion vectors that can be used to construct an affine motion field associated with a given CU. Thus, embodiments of the present disclosure result in technological advancements in computational techniques for encoding and decoding video content. For example, embodiments of the present disclosure improve the rate distortion performance provided by the affine combined encoding mode in JEM. Thus, the overall rate distortion performance of the video codec under consideration is improved.

[0102]

[0129] Further typical embodiments may be provided to modify the process in Figure 18. These embodiments include a rapid evaluation of the performance of each CPMV candidate by the following approximate strain and rate calculations. Thus, for each candidate in the set of CPMVs, the motion field of the current CU is calculated, and a 4x4 subblock-based time prediction of the current CU is performed. Next, the strain is calculated as SATD between the predicted CU and the original CU. The rate cost is obtained as the nearest order of bits tied to the signaling of the merge index of the candidates under consideration. A rough (approximate) RD cost is then obtained for each candidate. The final selection is, in one embodiment, based on the approximate RD cost. In other embodiments, a subset of candidates is subjected to a full RD search, i.e., the candidate with the lowest approximate RD cost is then subjected to a full RD search. The advantage of these embodiments is that they limit the increase in encoder-side complexity that would result from searching for the best affine merger predictor candidate.

[0103]

[0130] Furthermore, according to other general aspects of at least one embodiment, the affine intermode described above may be improved using all of the current teachings presented in this disclosure by having an expanded list of affine predictor candidates. As described above with respect to Figure 8A, one or more CPMVPs of an affine interCU are derived from adjacent motion vectors, regardless of their coding mode. Thus, it is then possible to average the affine neighbors using those affine models to construct one or more CPMVPs of the current affine interCU, similar to the affine merger mode described above. In this case, the affine candidates considered may be the same list as described above with respect to the affine merger mode (and are not limited to, for example, only spatial candidates).

[0104]

[0131] Therefore, a set of multiple predictor candidates is provided to improve the compression / decompression currently offered by HEVC and JEM by using better predictor candidates. The process becomes more efficient, and coding gains are observed, even if supplemental indices may need to be sent.

[0105]

[0132] According to a general aspect of at least one embodiment, a set of affine merger candidates (having at least seven candidates, similar to the merger mode) is, for example, • Spatial candidates from (A, B, C, D, E) • If there are fewer than 5 candidates in the list, the time candidate in the bottom right juxtaposition position will be... • If there are fewer than 5 candidates in the list, the time candidates in the juxtaposed positions are... ·Weighted average, ·Unique average, • If the weighted average is bidirectional and there are fewer than 7 candidates in the list, a one-way average is calculated from the weighted average. • If the unique mean is bidirectional and there are fewer than 7 candidates in the list, then the one-way mean is derived from the unique mean. It consists of.

[0106]

[0133] Also, in the case of AMVP, the predictor candidates are, for example, • Spatial candidates from the set (A, B, C, D, E) • Candidate supplementary spaces from (A', B'), • Time candidates in the lower right juxtaposed position It may be adopted from there.

[0107]

[0134] Tables 1 and 2 below show improvements over JEM4.0 (parallel) using several typical embodiments of the solution proposed in this disclosure. Each table shows the result of the amount of rate reduction for one of the typical embodiments described above. In particular, Table 1 shows the improvement when the five spatial candidates (A, B, C, D, E) shown in Figure 8B are used as a set of multiple predictor candidates according to the typical embodiments described above. Table 2 shows the improvement for a typical embodiment when predictor candidates are used as described above in the order of first spatial candidates, then temporal candidates if the number of candidates is still less than 5, then averages if the number of candidates is still less than 7, and finally unidirectional averages. For example, Table 2 shows that for this embodiment, the rate reductions for Y, U, and V samples are 0.22%, 0.26%, and 0.12% BD (Bjontegaard-Delta) rate reductions with respect to class D, respectively, with little increase in encoding and decoding execution time (i.e., 100% and 101%, respectively). Therefore, a typical embodiment of this disclosure improves compression / decompression efficiency compared to existing JEM implementations while maintaining computational complexity costs. [Table 1] [Table 2]

[0108]

[0135] Figure 26 shows a block diagram of a typical system 2600 in which various aspects of a typical embodiment can be realized. System 2600 may be embodied as a device including various components described later and configured to perform the processes described above. Examples of such devices include, but are not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. As will be known to those skilled in the art, System 2600 may be communicatively coupled to other similar systems and displays via communication channels as shown in Figure 26 in order to realize all or part of the typical video system described above.

[0109]

[0136] Various embodiments of System 2600 include at least one processor 2610 configured to execute instructions loaded to realize the various processes described above. The processor 2610 may include embedded memory, input / output interfaces, and various other circuits as known in the Art. System 2600 may also include at least one memory 2620 (e.g., volatile memory devices, non-volatile memory devices). System 2600 may further include a storage device 2640 which may include, but not limited to, non-volatile memory such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 2640 may include, in non-limiting examples, an internal storage device, a mounted storage device, and / or a network-accessible storage device. System 2600 may also include an encoder / decoder module 2630 configured to process data to provide encoded and / or decoded video, the encoder / decoder module 2630 which may include its own processor and memory.

[0110]

[0137] The encoder / decoder module 2630 represents a module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, such a device may include either or both encoding and decoding modules. In addition, the encoder / decoder module 2630 may be implemented as a separate element of system 2600, as is known to those skilled in the art, or incorporated within one or more processors 2610 as a combination of hardware and software.

[0111]

[0138] Program code to be loaded into one or more processors 2610 to perform the various processes described above may be stored in a storage device 2640 and then loaded into memory 2620 for execution by the processors 2610. According to a typical embodiment, one or more of the processors 2610, memory 2620, storage device 2640, and encoder / decoder modules 2630 may store one or more of the various things during the performance of the processes described above, including but not limited to input video, decoded video, bitstreams, equations, formulas, metrics, variables, actions, and operational logic.

[0112]

[0139] System 2600 may also include a communication interface 2650 that enables communication with other devices via a communication channel 2660. The communication interface 2650 may include, but is not limited to, a transceiver configured to send and receive data from the communication channel 2660. The communication interface 2650 may include, but is not limited to, a modem or a network card, and the communication channel 2650 may be implemented in a wired and / or wireless medium. Various components of System 2600 may be connected to each other or communicated using a variety of appropriate connections, including, but not limited to, internal buses, wires, and printed circuit boards (not shown in Figure 26).

[0113]

[0140] Typical embodiments may be implemented by computer software implemented by the processor 2610 or hardware, or by a combination of hardware and software. In non-limiting examples, typical embodiments may be implemented by one or more integrated circuits. The memory 2620 may be of any type suitable for the technical environment and, in non-limiting examples, may be implemented using any suitable data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 2610 may be of any type suitable for the technical environment and, in non-limiting examples, may include one or more microprocessors, general-purpose computers, dedicated computers, and processors based on multicore architectures.

[0114]

[0141] The implementations described herein may be realized, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if described only in the context of a single form of implementation (for example, described only as a method), the implementation of the described features may be realized in other forms (for example, apparatus or programs). Apparatus may be realized, for example, in appropriate hardware, software, and firmware. Methods may be realized in apparatus such as processors, which generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.

[0115]

[0142] Furthermore, those skilled in the art will readily understand that the typical HEVC encoder 100 shown in Figure 1 and the typical HEVC decoder shown in Figure 3 may be modified in accordance with the above teachings of this disclosure to implement the disclosed improvements to the existing HEVC standard for better compression / decompression. For example, the entropy coding 145, motion compensation 170, and motion estimation 175 in the typical encoder 100 of Figure 1, and the entropy decoding 330 and motion compensation 375 in the typical decoder of Figure 3 may be modified in accordance with the disclosed teachings to implement one or more typical embodiments of this disclosure, including providing advanced affine merge prediction to the existing JEM.

[0116]

[0143] References to “one embodiment,” “embodiment,” “one implementation,” or “implementation,” as well as other variations thereof, mean that certain features, structures, characteristics, etc., described in relation to an embodiment are included in at least one embodiment. Therefore, appearances of the phrases “in one embodiment,” “in one embodiment,” or “in one implementation,” or “in implementation,” as well as any other variations thereof, found in various places throughout this specification, do not necessarily all refer to the same embodiment.

[0117]

[0144] In addition, the scope of this application or claims may refer to “determining” various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.

[0118]

[0145] Furthermore, the scope of this application or claims may refer to “accessing” various types of information. Accessing information may include, for example, receiving information, retrieving information (for example, from memory), storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0119]

[0146] In addition, the scope of this application or claims may refer to “receiving” various types of information. Receiving is intended to be a broad term, similar to “accessing.” Receiving information may include, for example, accessing information or retrieving information (for example, from memory) one or more of these actions. Also, “receiving” is generally involved in any way during actions such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0120]

[0147] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that can be stored or transmitted, for example. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored in a processor-readable medium.

Claims

1. A method for video encoding, Accessing a set of predictor candidates having multiple predictor candidates with respect to the encoded block in the picture, wherein the predictor candidates correspond to the encoded spatial or temporally adjacent block. Selecting a predictor candidate from the aforementioned set of predictor candidates, Using a plurality of motion vectors associated with the selected predictor candidate from the set of predictor candidates, a set of control point motion vectors for the block is obtained, Obtaining a motion field based on a motion model, based on the set of control point motion vectors, wherein the motion field identifies the motion vectors used for predicting the subblocks of the encoded block. Encoding the block based on the motion field, Encoding an index relating to the selected predictor candidate from the set of predictor candidates. A method for providing this.

2. A method for video decoding, With respect to the decrypted blocks within the picture, accessing an index corresponding to a candidate predictor, wherein the candidate predictor corresponds to a decrypted spatial or temporally adjacent block. Using the multiple motion vectors associated with the predictor candidate, a set of control point motion vectors for the decoded block is obtained, Obtaining a motion field based on a motion model, based on the set of control point motion vectors, wherein the motion field identifies the motion vectors used for predicting the subblocks of the decoded block. Decoding the block based on the aforementioned motion field A method for providing this.

3. A device for video encoding, With respect to the encoded block in the picture, means for accessing a set of predictor candidates having multiple predictor candidates corresponding to the encoded spatial or temporally adjacent block, Means for selecting a predictor candidate from the set of predictor candidates, Means for obtaining a set of control point motion vectors for the block using a plurality of motion vectors associated with the selected predictor candidate from the set of predictor candidates, Means for obtaining a motion field based on a motion model, which identifies motion vectors used for predicting subblocks of the encoded block based on the set of control point motion vectors, Means for encoding the block based on the motion field, Means for encoding an index relating to the selected predictor candidate from the set of predictor candidates, A device equipped with the following features.

4. A device for video decoding, With respect to the decrypted blocks within a picture, a means for accessing an index corresponding to a predictor candidate corresponding to a decrypted spatial or temporally adjacent block, Means for obtaining a set of control point motion vectors for the decoded block using a plurality of motion vectors associated with the predictor candidate, Means for obtaining a motion field based on a motion model, which identifies motion vectors used for predicting subblocks of the decoded block based on the set of control point motion vectors, Means for decoding the block based on the motion field A device equipped with the following features.

5. Evaluating the plurality of predictor candidates according to one or more criteria and based on the motion field for each of the plurality of predictor candidates, Based on the above evaluation, select the predictor candidate from the plurality of predictor candidates. The encoding method according to claim 1, further comprising the following:

6. Means for evaluating the plurality of predictor candidates according to one or more criteria and based on the motion field relating to each of the plurality of predictor candidates, Based on the above evaluation, means for selecting the predictor candidate from the plurality of predictor candidates, The encoding device according to claim 3, further comprising the above.

7. The method or apparatus according to claims 5 to 6, wherein the one or more criteria are based on rate distortion determinations corresponding to one or more of the plurality of predictor candidates in the set of predictor candidates.

8. The method according to any one of claims 1, 2, and 5 to 7, wherein decoding or encoding the block based on the motion field comprises decoding or encoding the predictor indicated by the motion vector based on the predictor relating to the subblock.

9. The method according to any one of claims 1, 2, and 5 to 8, or the apparatus according to any one of claims 3 to 8, wherein the set of predictor candidates comprises spatial and / or temporal candidates for the block to be encoded or decoded.

10. The method according to any one of claims 1, 2, and 5 to 9, or the apparatus according to any one of claims 3 to 9, wherein the motion model is an affine model.

11. The motion field relating to each position (x, y) within the block being encoded or decoded is: [Math 1] Determined by, in the formula, (v 0x ,v 0y ) and (v 1x ,v 1y ) is the control point motion vector used to generate the motion field, and (v 0x ,v 0y ) corresponds to the control point motion vector of the upper left corner of the block being encoded or decoded, and (v 1x ,v 1y The method according to any one of claims 1, 2, and 5 to 10, or the apparatus according to any one of claims 3 to 10, wherein w corresponds to the control point motion vector of the upper right corner of the block to be encoded or decoded, and w is the width of the block to be encoded or decoded.

12. The method according to any one of claims 1, 2, and 5 to 11, or the apparatus according to any one of claims 3 to 11, wherein one or more additional predictor candidates are selected, a set of one or more additional control point motion vectors is obtained corresponding to the one or more additional predictor candidates, and the motion field is further obtained based on the set of one or more additional control point motion vectors.

13. A non-temporary computer-readable medium containing data content generated according to the method of any one of claims 1, 2, and 5 to 12.

14. A signal comprising video data generated according to the method of any one of claims 1, 2, and 5 to 12.

15. A computer program product comprising instructions for performing the method according to any one of claims 1, 2, and 5 to 12, when executed by one or more processors.