Improved predictor candidates for motion compensation

By expanding the selection of predictor candidates in affine merge mode using spatially adjacent blocks and motion models, the method improves video encoding compression efficiency and reduces data requirements.

JP2025093942AActive Publication Date: 2025-06-24INTERDIGITAL VC HOLDINGS INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025025983
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-10-05
Filing Date
2025-02-20
Publication Date
2025-06-24
Estimated Expiration
2038-10-04

Smart Images

  • Figure 2025093942000001_ABST
    Figure 2025093942000001_ABST
Patent Text Reader

Abstract

To provide encoding / decoding methods and devices for determining a set of predictor candidates for an affine merge coding mode.SOLUTION: A method comprises: determining, for a picture block, a set of predictor candidates for an inter coding mode on the basis of spatial neighboring blocks, having control point motion vectors and a reference picture; determining, for each predictor candidate, a motion field on the basis of a motion model and the multiple control point motion vectors of the predictor candidate, where the motion field identifies motion vectors used for prediction of sub-blocks of the block; selecting a single predictor candidate from the set of predictor candidates on the basis of a rate distortion determination between predictions in response to the motion field determined for each predictor candidate; and encoding / decoding the block on the basis of the motion field for the selected predictor candidate.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field [1] At least one of the embodiments generally relates to a method or apparatus for video encoding or decoding, and more particularly, to a method and apparatus for selecting one predictor candidate from a set of multiple predictor candidates for motion compensation in an inter-coding mode (merge mode or AMVP) based on a motion model such as an affine model (affine model) for a video encoder or a video decoder.

Background Art

[0002] Background [2] To achieve high compression efficiency, image and video coding schemes typically utilize prediction including motion vector (motion vector) prediction and perform transformation to exploit spatial and temporal redundancy within video content. Generally, intra or inter prediction is used to exploit intra or inter-frame correlation, and then the difference between the original image and the predicted image, often expressed as a prediction error or prediction residual, is transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction.

[0003] [3] Among the recently added technologies for high compression, those based on affine modeling The use of a motion model is included. Specifically, an affine model is used for motion compensation in the encoding and decoding of video pictures. Generally, affine modeling allows, for example, the derivation of a motion field for an entire block of a picture, such as by using at least two parameters, such as two control point motion vectors (CPMVs) representing the motion at individual corners of a block of the picture, to simulate rotation and similarity (zoom). However, the set of control point motion vectors (CPMVs) that are potentially used as predictors in the merge mode is limited. Therefore, by improving the performance of the motion model used in the affine merge and advanced motion vector prediction (AMVP) modes, a method for improving the overall compression performance of the target high compression technology is desirable. SUMMARY OF THE INVENTION

[0004] Summary [4] An object of the present invention is to overcome at least one of the drawbacks of the prior art. is provided. For this purpose, according to a general aspect of at least one embodiment, a method of video encoding is presented. The method includes determining, for a block to be encoded within a picture, at least one spatially adjacent block; determining, for the block to be encoded, a set of predictor candidates for an inter-coding mode based on the at least one spatially adjacent block, where the predictor candidates have one or more control point motion vectors and one reference picture; determining, for the block to be encoded and for each predictor candidate, a motion field based on a motion model and based on one or more control point motion vectors of the predictor candidate, where the motion field identifies motion vectors used for prediction of sub-blocks of the block to be encoded; selecting, based on rate-distortion determination during prediction in response to the motion fields determined for each predictor candidate, one predictor candidate from the set of predictor candidates; encoding the block based on the motion field of the selected predictor candidate; and encoding an index for the selected predictor candidate from the set of predictor candidates. The one or more control point motion vectors and the reference picture are used for prediction of the block to be encoded based on motion information associated with the block. Based on the rate-distortion determination during prediction in response to the motion fields determined for each predictor candidate, one predictor candidate is selected from the set of predictor candidates; the block is encoded based on the motion field of the selected predictor candidate; and an index for the selected predictor candidate from the set of predictor candidates is encoded. The one or more control point motion vectors and the reference picture are used for prediction of the block to be encoded based on motion information associated with the block.

[0005] [5] According to another general aspect of at least one embodiment, video decoding A method for an ing is provided. The method includes receiving, for a decoded block in a picture, an index corresponding to a specific predictor candidate among a set of predictor candidates for an inter-coding mode; determining, for the decoded block, at least one spatial adjacent block; determining, for the decoded block, a set of predictor candidates for the inter-coding mode based on the at least one spatial adjacent block, wherein the predictor candidates have one or more control point motion vectors and one reference picture; determining, for the specific predictor candidate, one or more corresponding control point motion vectors for the decoded block; determining, for the specific predictor candidate, a corresponding motion field based on the one or more control point motion vectors, wherein the corresponding motion field identifies motion vectors used for prediction of sub-blocks of the decoded block; and decoding the block based on the corresponding motion field.

[0006] [6] According to another general aspect of at least one embodiment, for video coding An apparatus is presented, the apparatus comprising means for determining at least one spatial neighboring block for a block to be encoded within a picture; means for determining a set of predictor candidates for an inter-coding mode based on at least one spatial neighboring block for the block to be encoded, wherein the predictor candidates have one or more control point motion vectors and one reference picture; means for selecting one predictor candidate from the set of predictor candidates; means for determining a motion field for the block to be encoded and for each predictor candidate, based on a motion model and on one or more control point motion vectors of the predictor candidate, wherein the motion field identifies motion vectors used for prediction of sub-blocks of the block to be encoded; means for selecting one predictor candidate from the set of predictor candidates in response to the motion fields determined for each predictor candidate, based on rate distortion determination during prediction; means for encoding the block based on the corresponding motion field of the selected predictor candidate from the set of predictor candidates; and means for encoding an index for the selected predictor candidate from the set of predictor candidates.

[0007] [7] According to another general aspect of at least one embodiment, an apparatus for video decoding is presented, the apparatus comprising means for receiving an index corresponding to a particular predictor candidate from a set of predictor candidates for an inter-coding mode for a block to be decoded within a picture; means for determining at least one spatial neighboring block for the block to be decoded; means for determining a set of predictor candidates for an inter-coding mode based on at least one spatial neighboring block for the block to be decoded, wherein the predictor candidates have one or more control point motion vectors and one reference picture; and means for decoding means for determining, for a block to be decoded, one or more corresponding control point motion vectors from a particular predictor candidate; means for determining, for a block to be decoded, a motion field based on a motion model and based on one or more control point motion vectors for the block to be decoded, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be decoded; and means for decoding a block based on the corresponding motion field.

[0008] [8] According to another general aspect of at least one embodiment, video encoding An apparatus for [purpose not clear from the context] is provided, the apparatus having one or more processors and at least one memory. In this case, the one or more processors determine at least one spatial adjacent block for a block to be encoded in a picture, determine a set of predictor candidates for an inter-coding mode based on the at least one spatial adjacent block for the block to be encoded, where the predictor candidates have one or more control point motion vectors and one reference picture, for the block to be encoded and for each predictor candidate, determine a motion field based on a motion model and based on the one or more control point motion vectors of the predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be encoded, select one predictor candidate from the set of predictor candidates based on rate distortion determination during prediction in response to the motion fields determined for each predictor candidate, encode the block based on the motion field of the selected predictor candidate, and encode an index for the selected predictor candidate from the set of predictor candidates. The at least one memory is for at least temporarily storing the encoded block and / or the encoded index.

[0009] [9] According to another general aspect of at least one embodiment, video decoding An apparatus for use is provided, the apparatus having one or more processors and at least one memory. In this case, the one or more processors receive an index corresponding to a specific predictor candidate among a set of predictor candidates for an inter-coding mode for a block to be decoded in a picture, determine at least one spatial adjacent block for the block to be decoded, determine a set of predictor candidates for the inter-coding mode based on at least one spatial adjacent block for the block to be decoded, the predictor candidates having one or more control point motion vectors and one reference picture, determine one or more corresponding control point motion vectors for the block to be decoded for a specific predictor candidate, determine a motion field based on a motion model based on one or more corresponding control point motion vectors for a specific predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be decoded, and decode the block based on the motion field. The at least one memory is for at least temporarily storing the decoded block.

[0010]

[10] According to another general aspect of at least one embodiment, the at least one spatial adjacent block has spatial adjacent blocks of the block to be encoded or decoded among an adjacent upper left corner block, an adjacent upper right corner block, and an adjacent lower left corner block.

[0011]

[11] According to another general aspect of at least one embodiment, the motion information associated with at least one of the at least one spatial adjacent block has non-affine motion information. The non-affine motion model is a translational motion model, in which case only one motion vector representing translation is coded in the model.

[0012]

[12] According to another general aspect of at least one embodiment, motion information associated with all of at least one spatially adjacent block has affine motion information.

[0013]

[13] According to another general aspect of at least one embodiment, a set of predictor candidates has unidirectional predictor candidates and bidirectional predictor candidates.

[0014]

[14] According to another general aspect of at least one embodiment, the method includes determining an upper left list of spatially adjacent blocks of a block to be encoded or decoded among adjacent upper left corner blocks, an upper right list of spatially adjacent blocks of a block to be encoded or decoded among adjacent upper right corner blocks, and a lower left list of spatially adjacent blocks of a block to be encoded or decoded among adjacent lower left corner blocks, and selecting at least one triplet of spatially adjacent blocks, wherein each spatially adjacent block of the triplet belongs to the upper left list, the upper right list, and the lower left list, respectively, and a reference picture used for prediction of the spatially adjacent blocks of the triplet is the same, and further may include determining one or more control point motion vectors for the upper left corner, upper right corner, and lower left corner of the block based on the motion information respectively associated with each spatially adjacent block of the selected triplet for the block to be encoded or decoded, in which case the predictor candidate has the determined one or more control point motion vectors and the reference picture.

[0015]

[15] According to another general aspect of at least one embodiment, the method can further include evaluating at least one selected triplet of spatially adjacent blocks according to one or more criteria based on one or more control point motion vectors determined for the block to be encoded or decoded, and in this case, the predictor candidates are sorted within a set of predictor candidates for the inter-coding mode based on the evaluation.

[0016]

[16] According to another general aspect of at least one embodiment, the one or more criteria have a validity check according to Equation 3 and a cost according to Equation 4.

[0017]

[17] According to another general aspect of at least one embodiment, the cost of a bidirectional predictor candidate is the average of the cost related to its first reference picture list and the cost related to its second reference picture list.

[0018]

[18] According to another general aspect of at least one embodiment, the method includes determining an upper left list of spatially adjacent blocks of the block to be encoded or decoded among adjacent upper left corner blocks, an upper right list of spatially adjacent blocks of the block to be encoded or decoded among adjacent upper right corner blocks, and selecting at least one pair of spatially adjacent blocks, wherein each of the spatially adjacent blocks of the pair belongs to the upper left list and the upper right list respectively, and the reference pictures used for prediction of each of the spatially adjacent blocks of the pair are the same, and for the block to be encoded or decoded, based on the motion information related to the spatially adjacent blocks of the upper left list, the b Determining a control point motion vector for the upper left corner of a block based on a control point motion vector for the upper left corner of the block and motion information related to spatially adjacent blocks in the upper left list, wherein a predictor candidate has the upper left and upper right control point motion vectors and a reference picture, and can further have the above.

[0019]

[19] According to another general aspect of at least one embodiment, a lower left list is used instead of an upper right list, the lower left list has spatially adjacent blocks of a block to be encoded or decoded among adjacent lower left corner blocks, and a lower left control point motion vector has been determined.

[0020]

[20] According to another general aspect of at least one embodiment, a motion model is an affine model, and a motion field at each position (x, y) inside a block to be encoded or decoded is determined by the following formula. [Number]

[0021]

[21] Here, (v 0x , v 0y ) and (v 2x , v 2y ) are control point motion vectors used to generate a motion field, (v 0x , v 0y ) corresponds to the control point motion vector of the upper left corner of a block to be encoded or decoded, (v 2x , v 2y ) corresponds to the control point motion vector of the lower left corner of a block to be encoded or decoded, and h is the height of a block to be encoded or decoded.

[0022]

[22] According to another general aspect of at least one embodiment, the method may further include encoding or obtaining a notification of a motion model used for a block to be encoded or decoded, where the motion model is based on a control point motion vector of an upper left corner and a control point motion vector of a lower left corner, or the motion model is based on a control point motion vector of an upper left corner and a control point motion vector of an upper right corner.

[0023]

[23] According to another general aspect of at least one embodiment, the motion model used for a block to be encoded or decoded is implicitly derived, where the motion model is based on a control point motion vector of an upper left corner and a control point motion vector of a lower left corner, or the motion model is based on a control point motion vector of an upper left corner and a control point motion vector of an upper right corner.

[0024]

[24] According to another general aspect of at least one embodiment, encoding or decoding a block based on a corresponding motion field respectively includes encoding or decoding based on a predictor for a sub-block, where the predictor is notified by a motion vector.

[0025]

[25] According to another general aspect of at least one embodiment, the number of spatially adjacent blocks is at least five or at least seven.

[0026]

[26] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is presented that includes data content generated according to the method or apparatus of any of the preceding descriptions.

[0027]

[27] According to another general aspect of at least one embodiment, a signal having video data generated according to the method or apparatus of any of the preceding descriptions is provided.

[0028]

[28] Also, one or more of the present embodiments provide a computer-readable storage medium having instructions for encoding or decoding video data according to any of the methods described above stored thereon. Further, the present embodiment provides a computer-readable storage medium having a bitstream generated according to the method described above stored thereon. Further, the present embodiment provides a method and apparatus for transmitting a bitstream generated according to the method described above. Further, the present embodiment provides a computer program product including instructions for executing any of the methods described.

Brief Description of the Drawings

[0029] Brief Description of the Drawings

Figure 1

[29] A block diagram of one embodiment of a HEVC (High Efficiency Video Coding) video encoder is shown.

Figure 2A

[30] A drawing example depicting the generation of HEVC reference samples.

Figure 2B

[31] A drawing example depicting the intra prediction direction in HEVC.

Figure 3

[32] A block diagram of one embodiment of a HEVC video decoder is shown.

Figure 4

[33] An example of the coding tree unit (CTU: Coding Tree Unit) and coding tree (CT: Coding Tree) concepts for representing a compressed HEVC picture is shown.

Figure 5

[34] An example of the division of a Coding Tree Unit (CTU) into a Coding Unit (CU), a Prediction Unit (PU), and a Transform Unit (TU) is shown.

Figure 6

[35] An example of an affine model as a motion model used in the Joint Exploration Model (JEM) is shown.

Figure 7

[36] An example of an affine motion vector field based on a 4×4 sub CU used in the Joint Exploration Model (JEM) is shown.

Figure 8A

[37] An example of a motion vector prediction candidate for an affine inter CU is shown.

Figure 8B

[38] An example of a motion vector prediction candidate in the affine merge mode is shown.

Figure 9

[39] An example of the spatial derivation of an affine control point motion vector in the case of an affine merge mode motion model is shown.

Figure 10

[40] An exemplary encoding method according to a general aspect of at least one embodiment is shown.

Figure 11

[41] Another example of an encoding method according to a general aspect of at least one embodiment is shown.

Figure 12

[42] An example of a motion vector prediction candidate for an affine merge mode CU according to a general aspect of at least one embodiment is shown.

Figure 13

[43] An example of an affine model as a motion model and an affine motion vector field based on its corresponding 4×4 sub CU according to a general aspect of at least one embodiment is shown.

Figure 14

[44] Another example of an affine model as a motion model and an affine motion vector field based on its corresponding 4×4 sub CU according to a general aspect of at least one embodiment is shown.

Figure 15

[45] An example of a known process / syntax for evaluating the affine merge mode of a CU in JEM is shown.

Figure 16

[46] An exemplary decoding method according to a general aspect of at least one embodiment is shown.

Figure 17

[47] Another example of an encoding method according to a general aspect of at least one embodiment is shown.

Figure 18

[48] A block diagram of an exemplary apparatus in which various aspects of an embodiment can be implemented is shown. DETAILED DESCRIPTION

[0030] Detailed Description

[49] It should be understood that the figures and the description are simplified to show elements suitable for a clear understanding of the principles, while removing many other elements found in a normal coding and / or decoding apparatus for the purpose of clarity. The terms "first" and "second" may be used herein to describe various elements, but it should be understood that these elements are not limited by these terms. These terms are only used to distinguish elements from each other. It should be understood that when used herein, these elements are not limited by these terms. These terms are only used to distinguish elements from each other.

[0031]

[50] Various embodiments are described in the context of the HEVC standard. However, the principles are not limited to HEVC and can be applied to other standards, recommendations, and their extensions, including HEVC or HEVC extensions such as, for example, Format Range (RExt), Scalability (SHVC), Multi-View (MV-HEVC) extensions, and H.266. Various embodiments are described in the context of slice encoding / decoding. These can be applied to encode / decode an entire picture or an entire sequence of pictures.

[0032]

[51] Various methods have been described above, and each of the methods has one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or usage of specific steps and / or actions may be changed or combined.

[0033]

[52] FIG. 1 shows an exemplary high efficiency video coding (HEVC) encoder 100. HEVC is a compression standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) (see, for example, "ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265").

[0034]

[53] In HEVC, to encode a video sequence having one or more pictures, the pictures are partitioned into one or more slices, where each slice may include one or more slice segments. The slice segments are organized as coding units, prediction units, and transform units.

[0035]

[54] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, and the term "encode" may be used interchangeably with the term "decode". The terms "encoded" or "coded" may also be used interchangeably, and the terms "picture" and "frame" may also be used interchangeably. Usually, the term "reconstructed" is used on the encoder side, while the term "decoded" is used on the decoder side, but this is not mandatory.

[0036]

[55] The HEVC specification discriminates between "blocks" and "units", where in this case a "block" addresses a specific area (e.g., luma, Y) within the sample array, and a "unit" includes all the encoded color components (Y, Cb, Cr, or monochrome), syntax elements, and the block together with the prediction data (e.g., motion vectors) associated with the block.

[0037]

[56] In the case of coding, a picture is partitioned as a square-shaped coding tree block (CTB) with a configurable size, and a consecutive set of coding tree blocks is grouped as a slice. A coding tree unit (CTU) includes the CTB of the encoded color components. A CTB is the root of a quadtree partitioning into coding blocks (CB), and a coding block may be partitioned as one or more prediction blocks (PB), and a quadtree partitioning into transform blocks (TB). may be partitioned, and a quadtree partitioning into transform blocks (TB). forms the root. By corresponding to coding blocks, prediction blocks, and transform blocks, a coding unit (CU) includes a tree-structured set of a prediction unit (PU) and a transform unit (TU). The PU contains prediction information for all color components, and the TU contains a residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luma component are applied to the corresponding CU, PU, and TU. In this application, the term "block" can be used to mean any of, for example, CTU, CU, PU, TU, CB, PB, and TB. In addition to this, "block" can be used to mean macroblocks and partitions defined in H.264 / AVC or other video coding standards, and more generally, an array of data of various sizes.

[0038]

[57] In the exemplary encoder 100, as will be described later, a picture is encoded by encoder elements. The picture to be encoded is processed within a unit of a CU. Each CU is encoded by using an intra or inter mode. When a CU is encoded in the intra mode, intra prediction (160) is performed. In the inter mode, motion estimation (175) and compensation (170) are performed. The encoder determines one of the intra mode or the inter mode (105) for use in encoding the CU, and notifies the intra / inter decision by a prediction mode flag. The prediction residual is calculated by subtracting the predicted block from the original image block (110).

[0039]

[58] CUs in intra mode are predicted from the reconstructed adjacent samples within the same slice. In HEVC, a set of 35 intra prediction modes including one DC prediction mode, planar prediction mode, and 33 angular prediction modes are available. The intra prediction reference is reconstructed from the rows and columns adjacent to the current block. The reference extends horizontally and vertically over twice the block size by using the available samples from the pre-reconstructed blocks. When an angular prediction mode is used for intra prediction, the reference samples can be copied along the direction notified by the angular prediction mode.

[0040]

[59] The applicable luma intra prediction modes for the current block can be coded by using two different options. If the applicable mode is included in the constructed list of the three most probable modes (MPM: most probable mode), the mode is signaled by an index within the MPM list. Otherwise, the mode is signaled by the fixed-length binarization of the mode index. The three most probable modes are derived from the intra prediction modes of the upper and left adjacent blocks.

[0041]

[60] For inter CUs, the corresponding coding block is further partitioned as one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information regarding the way inter prediction is performed. Motion information (i.e., motion vectors and reference picture indices) can be signaled in two ways, namely, "merge mode" and "advanced motion vector prediction (AMVP)".

[0042]

[61] In the merge mode, the video encoder or decoder edits the candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates in the candidate list. On the decoder side, based on the signaled candidate, the motion vector (MV) and the reference picture index are reconstructed. The reference picture index is reconstructed.

[0043]

[62] The possible candidate sets in the merge mode are composed of spatial neighboring candidates, temporal candidates, and generated candidates. FIG. 2A shows the positions of five spatial candidates {a1, b2, b0, a0, b2} for block 21 at the current time. In this case, a0 and a1 are located on the left side of the current block, and b1, b0, and b2 are located above the current block. For each candidate position, the availability is checked in the order of a1, b1, b0, a0, b2, and then the redundancy in the candidates is removed.

[0044]

[63] For the derivation of temporal candidates, the motion vectors of grouped locations in the reference picture can be used. The applicable reference picture is selected based on the slice and notified in the slice header, and the reference index for temporal candidates is set to i ref =0. If the POC distance (td) between the picture of the grouped PU and the reference picture from which the grouped PU is predicted is the same as the distance (tb) between the current picture and the reference picture including the grouped PU, the grouped motion vector mv col can be directly used as a temporal candidate. Otherwise, the scaled motion vector tb / td*mv col is used as a temporal candidate. Depending on the location where the current PU is placed, the grouped PU is determined by the sample location at the lower right or center of the current PU.

[0045]

[64] The maximum number N of merge candidates is defined in the slice header. If the number of merge candidates exceeds N, only the first N - 1 spatial and temporal candidates are used. Otherwise, if the number of merge candidates is less than N, the set of candidates is filled up to the maximum number N with the generated candidates as the combination of the existing candidates or with null candidates. The candidates used in the merge mode may be referred to as "merge candidates" in this application.

[0046]

[65] When the CU notifies the skip mode, the applicable index for the merge candidates is notified only when the list of merge candidates is more than 1, and further information is not coded for the CU. In the skip mode, the motion vector is applied without residual update.

[0047]

[66] In AMVP, the video encoder or decoder edits the candidate list based on the motion vectors determined from the already-coded blocks. Then, the video encoder signals an index in the candidate list and signals the motion vector difference (MVD) to identify the motion vector predictor (MVP). On the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. Also, the applicable reference picture index is explicitly coded in the PU syntax for AMVP.

[0048] ​

[67] In AMVP, only two spatial motion candidates are selected. The first spatial motion candidate is selected from the left positions {a0, a1} while maintaining the search order notified within two groups, and the second one is selected from the upper positions {b0, b1, b2}. When the number of motion vector candidates is not equal to 2, temporal MV candidates can be included. When the candidate group is still not sufficiently filled, zero motion vectors are used.

[0049]

[68] When the reference picture index of the spatial candidate corresponds to the reference picture index for the current PU (i.e., using the same reference picture index independently of the reference picture list, or using both long-term reference pictures), the spatial candidate motion vector is directly used. Otherwise, when both reference pictures are short-term, the candidate motion vector is scaled according to the distance (tb) between the current picture and the reference picture of the current PU and the distance (td) between the current picture and the reference picture of the spatial candidate. The candidates used in AMVP may be referred to as "AMVP candidates" in this application.

[0050]

[69] For ease of notation, blocks tested in the "merge" mode on the encoder side or decoded in the "merge" mode on the decoder side are denoted as "merge blocks", and blocks tested in the AMVP mode on the encoder side or decoded in the AMVP mode on the decoder side are denoted as "AMVP" blocks.

[0051]

[70] Figure 2B shows an exemplary motion vector representation using AMVP. For the current block 240 to be encoded, the motion vector (MV current ) can be obtained through motion estimation. The motion vector (MV left from the left block 230) and the motion vector (MV above ) from the upper block 220, by using left MV above and MV current , the motion vector predictor can be selected as MVP current . Then, the motion vector difference can be calculated as MVD current = MV current - MVP

[0052]

[71] Motion compensation prediction can be performed by using one or more reference pictures for prediction. In P slices, unidirectional prediction for the prediction block can be enabled by using only a single prediction criterion for inter prediction. In B slices, two reference picture lists are available and unidirectional or bidirectional prediction can be used. In bidirectional prediction, one reference picture from each of the reference picture lists is used.

[0053]

[72] In HEVC, the accuracy of the motion information for motion compensation is 1 / 4 sample for the luma component (also referred to as 1 / 4 pel or 1 / 4 pixel), and 1 / 8 sample for the chroma component for the 4:2:0 configuration (also referred to as 1 / 8 pel). In the case of luma, a 7-tap or 8-tap interpolation filter is used for interpolation of fractional sample positions, i.e., it can address 1 / 4, 1 / 2, and 3 / 4 of the full sample locations in both the horizontal and vertical directions.

[0054]

[73] Next, the prediction residual is transformed (125) and quantized (130). To output a bitstream, not only the quantized transform coefficients but also motion vectors and other syntax elements are entropy coded (145). Also, the encoder can skip the transformation and apply quantization directly to the untransformed residual signal in a 4×4 TU manner. Further, the encoder may bypass both the transformation and quantization, i.e., the residual is directly coded without the application of the transformation or quantization process. In direct PCM coding, prediction is not applied and the coding unit samples are directly coded in the bitstream.

[0055]

[74] The encoder decodes the encoded blocks to provide further prediction references. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The image block is reconstructed by combining the decoded prediction residual and the predicted block (155). For example, an in-loop filter (165) is applied to the reconstructed picture to perform deblocking / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (180).

[0056]

[75] FIG. 3 shows a block diagram of an exemplary HEVC video decoder 300. In the exemplary decoder 300, the bitstream is decoded by decoder elements described hereinafter. The video decoder 300 generally performs a decoding path opposite to the encoding path described in FIG. 1 that performs video decoding as part of the encoding of video data.

[0057]

[76] Specifically, the input to the decoder includes a video bitstream that can be generated by the video encoder 100. The bitstream is first entropy decoded (330) to obtain transform efficiency, motion vectors, and other coded information. To decode the prediction residual, the transform coefficients are inverse quantized (340) and inverse transformed (350). An image block is reconstructed by combining the decoded prediction residual and the predicted block (355). The predicted block can be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375) (370). As described above, the AMVP and merge mode techniques can be used to derive motion vectors for motion compensation that can use an interpolation filter to calculate interpolated values for sub-integer samples of the reference block. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (380).

[0058]

[77] As described above, in HEVC, motion-compensated temporal prediction is used to exploit the redundancy that exists between consecutive pictures of a video. To do this, motion vectors are associated with each prediction unit (PU). As described above, each CTU is represented by a coding tree in the compressed domain. This is a quadtree partitioning of the CTU, in which case each leaf is referred to as a coding unit (CU) and is also shown in FIG. 4 for CTUs 410 and 420. Next, each CU is given several intra or inter prediction parameters as prediction information. To perform this, the CU may be spatially partitioned as one or more prediction units (PUs), in which case some prediction information is assigned to each PU. An intra or inter coding mode is assigned at the CU level. These concepts are further shown in FIG. 5 for exemplary CTU 500 and CU 510.

[0059]

[78] In HEVC, one motion vector is assigned to each PU. This motion vector is used for motion-compensated temporal prediction of the target PU. Thus, in HEVC, the motion model that links the predicted block and its reference block is composed of a translation or calculation based on the reference block and the corresponding motion vector.

[0060]

[79] To implement improvements to HEVC, reference software and / or the document JEM (Joint Exploration Model) has been developed by the Joint Video Exploration Team (JVET). In one JEM version (i.e., “Algorithm Description of Joint Exploration Test Model 5”, Document JVET-E1001_v2, Joint Video Exploration Team of ISO / IEC JTC1 / SC29 / WG11, 5rd meeting, 12 - 20 January 2017, Geneva, CH) , to improve temporal prediction, several additional motion models are supported. To perform this, it is possible to spatially divide the PU into sub-PUs and use a model to assign a dedicated motion vector to each sub-PU.

[0061]

[80] The more recent versions of JEM (e.g., “Algorithm Description of Joint Exploration Test Model 2”, Document JVET-B1001_v3, Joint Video Exploration Team of ISO / IEC JTC1 / SC29 / WG11, 2rd meeting, 20-26 February 2016, San Diego, USA”) In it, the CU is no longer defined as being divided into PUs or TUs. Instead, a relatively flexible CU size may be used, and some motion data is directly assigned to each CU. In this new codec design in the relatively new versions of JEM, the CU can be divided into sub-CUs, and for each sub-CU of the divided CU, a motion vector can be calculated.

[0062]

[81] One of the new motion models introduced in JEM is the use of an affine model as a motion model to represent the motion vectors within a CU. The motion model used is shown in FIG. 6 and is represented by Equation 1 below. The affine motion field has the following motion vector component values at each position (x, y) inside the target block 600 of FIG. 6,

Number

Number

Number

[0063]

[82] To reduce complexity, the motion vectors are calculated for each 4×4 sub-block (sub-CU) of the target CU700 as shown in FIG. 7. The affine motion vectors are calculated from the control point motion vectors for each center position of each sub-block. The obtained MV is represented at a precision of 1 / 16 pel. As a result, the compensation of the coding unit in the affine mode has motion-compensated prediction of each sub-block by its own motion vector. These motion vectors for the sub-blocks are individually shown as arrows for each of the sub-blocks in FIG. 7.

[0064]

[83] In JEM, affine motion compensation can be used in two ways: the affine AMVP (AF_AMVP) mode and the affine merge mode. These will be introduced in the following sections.

[0065]

[84] Affine AMVP Mode: In the Affine AMVP mode, CUs in the AMVP mode with a size greater than 8×8 can be predicted. This is signaled through a flag in the bitstream. The generation of the affine motion field for this AMVP-CU involves determining control point motion vectors (CPMVs), which are obtained by the encoder or decoder through the addition of motion vector differences and control point motion vector prediction (CPMVP). CPMVPs are pairs of motion vector candidates respectively obtained from the sets (A, B, C) and (D, E) shown in FIG. 8A for the current CU800 being encoded or decoded.

[0066]

[85] Affine Merge Mode: In the Affine Merge mode, a CU-level flag indicates whether the merge CU is using affine motion compensation. If it is, the first available adjacent CU coded in the affine mode is selected from the ordered set of candidate positions A, B, C, D, E in FIG. 8B for the current CU880 being encoded or decoded. Note that this ordered set of candidate positions in JEM is the same as the spatial adjacent candidates in the HEVC Merge mode shown in FIG. 2A and described above.

[0067]

[86] Once the first adjacent CU in the affine mode is obtained, three CPMVs from the upper left, upper right, and lower left corners of the adjacent affine CU

Number

Number

[0068]

[87] The control point motion vector of the current CU

Number

Number

[0069]

[88] Accordingly, a general aspect of at least one embodiment aims to improve the performance of the affine merge mode in JEM so that the compression performance of the target video codec can be improved. Accordingly, in at least one embodiment, an improved motion compensation apparatus and method are presented for a coding / decoding unit coded in the affine merge mode. The proposed improved affine mode includes determining a set of predictor candidates in the affine merge mode regardless of whether adjacent CUs are coded in the affine mode.

[0070]

[89] As described above, in the current JEM, to predict the affine motion model associated with the current CU encoded or decoded in the affine merge mode, the first adjacent CU coded in the affine mode among the surrounding CUs is selected. That is, to predict the affine motion model of the current CU in the affine merge mode, the first adjacent CU of the ordered set (A, B, C, D, E) of FIG. 8B coded in the affine mode candidate is selected.

[0071]

[90] Thus, at least one embodiment improves the affine merge prediction candidates, thereby generating new motion model candidates from the motion vectors of adjacent blocks used as CPMVP, and provides the best coding efficiency when coding the current CU in the affine merge mode. Also, when decoding in merge to obtain a predictor from the signaled index, corresponding new motion model candidates generated from the motion vectors of adjacent blocks used as CPMVP are also determined. Thus, at a general level, the improvements of this embodiment are, for example, · constructing a set of uni - or bi - directional affine merge predictor candidates based on adjacent CU motion information, regardless of whether the adjacent CUs are coded (for the encoder / decoder) in the affine mode, · constructing a set of affine merge predictor candidates, for example, (for the encoder / decoder) by selecting control point motion vectors from only two adjacent blocks such as the top - left (TL) block and top - right (TR) or bottom - left (BL), · adding a penalty for uni - directional predictor candidates in comparison with bi - directional predictor candidates (for the encoder / decoder) when evaluating the predictor candidates, · adding adjacent blocks only when the adjacent blocks are coded in affine coding (for the encoder / decoder), · and / or · signaling / decoding the notification of the motion model used for the control point motion vector predictor of the current CU (for the encoder / decoder). It has.

[0072]

[91] Although an encoding / decoding method based on a merge mode is described, this principle also applies to the AMVP (i.e., affine_inter) mode. Advantageously, various embodiments for generating predictor candidates are clearly derivable for affine AMVP.

[0073]

[92] This principle is advantageously implemented in an encoder within the motion estimation module 175 and the motion compensation module 170 of FIG. 1, or in a decoder within the motion compensation module 275 of FIG. 2.

[0074]

[93] Accordingly, FIG. 10 shows an exemplary encoding method 1000 according to a general aspect of at least one embodiment. At 1010, method 1000 determines at least one spatially adjacent block for a block to be encoded within a picture. For example, as shown in FIG. 12, adjacent blocks of the block CU to be encoded among A, B, C, D, E, F, G are determined. At 1020, method 1000 determines a set of predictor candidates for affine merge mode from the spatially adjacent blocks for a block to be encoded within a picture. Further details for such determination are provided below in relation to FIG. 11. In the JEM affine merge mode, the first adjacent block from the ordered list F, D, E, G, A represented in FIG. 12, which is coded in the affine mode, is used as a predictor, while according to a general aspect of this embodiment, a plurality of predictor candidates are evaluated for the affine merge mode, in which case the adjacent blocks can be used as predictor candidates regardless of whether the adjacent blocks are coded in the affine mode. According to a particular aspect of this embodiment, the predictor candidates are one or more corresponding It has a control point motion vector and one reference picture. The reference picture is used for prediction of at least one spatially adjacent block based on motion compensation in relation to one or more corresponding control point motion vectors. The reference picture and one or more corresponding control point motion vectors are derived from motion information associated with at least one of the spatially adjacent blocks. Naturally, the motion compensation prediction can be performed by using one or two reference pictures for prediction according to uni - directional or bi - directional prediction. Thus, two reference picture lists are available to store the motion information from which the reference picture and the motion vector / (control point motion vector in the case of the affine mode) are derived. In the case of non - affine motion information, one or more corresponding control point motion vectors correspond to the motion vectors of the spatially adjacent blocks. In the case of affine motion information, one or more corresponding control point motion vectors are, for example, the control point motion vectors of the spatially adjacent blocks

Number

Number

[0075]

[94] FIG. 11 shows details for illustration of the determination 1020 of a set of predictor candidates for the merge mode of an encoding method 1000 according to one aspect of at least one embodiment, particularly adapted when adjacent blocks have non-affine motion information. However, the method is compatible with adjacent blocks having affine motion information. Those skilled in the art will understand that in the case where adjacent blocks are coded by an affine model, as detailed above with reference to FIG. 9, the motion model of the affine-coded adjacent blocks can be used to determine predictor candidates for the affine mode. At 1021, by considering at least one adjacent block of the blocks to be encoded, three lists are determined based on the spatial position of the blocks in relation to the blocks to be encoded. For example, by considering the blocks A, B, C, D, E, F, G in FIG. 12, a first list, called the upper left list, having adjacent upper left corner blocks A, B, C is generated, a second list, called the upper right list, having adjacent upper right corner blocks D, E is generated, and a third list having adjacent lower left corner blocks F, G is generated. According to a first aspect of at least one embodiment, a triplet of spatially adjacent blocks is selected, where each spatial adjacent block of the triplet individually belongs to the upper left list, the upper right list, and the lower left list, and where the reference picture used for the prediction of each spatial adjacent block of the triplet is the same. Thus, the triplets correspond to the selection of {A, B, C}, {D, E}, {F, G}. At 1022, based on the usage of the same reference picture for each adjacent block of the triplets for motion compensation, a first selection among the 12 possible triplets {A, B, C}, {D, E}, {F, G} is applied. Thus, only the triplets A, D, F are selected, where the same reference picture (identified by its index in the reference picture list) is used for the prediction of the adjacent blocks A, D, and F.This first criterion guarantees coherence in prediction. At 1023, it is one or more control point motion vectors of the current CU.

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

[0076] ​

[95] At 1024, an arbitrarily selected second selection among the candidate CPMVPs is applied based on one or more criteria. According to the first criterion, the candidate CPMVPs are further checked for the block to be encoded with height H and width W for verification using Equation 3, and in this case, X and Y are the horizontal and vertical direction components of the motion vector, respectively.

Number

[0077]

[96] Therefore, in this variant, at 1025, while the set of predictor candidates is not full, these valid CPMVPs are stored within the set of CPMVP candidates for the merge mode. Then, according to the second criterion, the valid candidate CPMVPs are sorted according to the value of the lower left motion vector (obtained from position F or G). The closest

Number

Number

Number

Number

[0078]

[97] In the case of bidirectional prediction, by using Equation 4, the cost is computed for each CPMVP and for each reference picture list L0 and L1. For each predictor bidirectional candidate, to compare the unidirectional predictor candidate with the bidirectional predictor candidate, the cost of the CPMVP is the average of the CPMVP cost related to its list L0 and the CPMVP cost related to its list L1. According to this variation, at 1025, the ordered set of valid CPMVPs is stored within the set of CPMVP candidates for merge mode for each reference picture list L0 and L1 respectively.

[0079]

[98] According to a second aspect of at least one embodiment of the determination of the set of predictor candidates for the merge mode 1020 of the encoding method 1000 shown in FIG. 11, instead of using a triplet of adjacent blocks to determine three control motion point vectors for a predictor candidate, the determination is made using a pair of adjacent vectors from which two control motion point vectors for a predictor candidate are determined, i.e., for a predictor candidate

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Mathematics

Mathematics

[0080]

[99] In addition to this, by considering pairs from the upper left list {A, B, C} and the lower left list {F, G}, the two control point motion vectors of the current CU are

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

[0081]

[0100] As described by Equation 1, and as shown in FIGS. 6 and 7 As such, the standard affine motion model is the upper left and upper right control point motion vectors

Number

Number

Number

Number

Number

Number

Number

Number

[0082]

[0101] In addition to this, based on pairs of adjacent blocks to determine two CPMVs It should be noted that the second aspect, where the third CPMV (by formula 5 or formula 6) has not been determined independently of the first two CPMVs, is not compatible with the evaluation 1024 of predictor candidates based on those individual CPMVs. Therefore, the validity check of formula 3 and the cost function of formula 4 are skipped, and the CPMVPs determined at 1023 are added to the set of predictor candidates for the affine merge mode without any sorting.

[0083]

[0102] In addition to this, based on pairs of adjacent blocks for determining two CPMVs According to a variant of the second aspect, first, the use of bidirectional affine merge candidates is prioritized over unidirectional affine merge candidates by adding bidirectional candidates within the set of predictor candidates. As a result, it is guaranteed to obtain the maximum value of the bidirectional candidates added within the set of predictor candidates for the affine merge mode.

[0084]

[0103] In at least one implementation of the determination 1020 of the set of predictor candidates for the merge mode, in one version of the first and second aspects, only affine adjacent blocks are used to generate new affine motion candidates. In other words, the determination of the upper left CPMV, upper right CPMV, and lower left CPMV is based on the motion information of the individual upper left adjacent block, upper right adjacent block, and lower left adjacent block. In this case, the motion information related to at least one spatial adjacent block has only affine motion information. This variant is different from the JEM affine merge mode in that the motion model of the selected affine adjacent block is extended to the block to be encoded.

[0085]

[0104] In at least one implementation of the determination of the set of predictor candidates for merge mode 1020 In another variation of the first and second aspects of the embodiment, at 1030, the upper left CPMV

Number

Number

Number

Number

Number

Number

Number

[0086]

[0105] This variation is particularly well-suited for embodiments based on pairs of upper left and lower left adjacent blocks, in which case the upper left and lower left CMPV to, and in this case, the upper left and lower left CMPV

Number

Number

[0087]

[0106] In addition to this, this transformation provides an additional CPMVP for the combination of predictor candidates in the affine merge mode, and CPMV CPMV with [Number] and CPMV with [Number] are added to the set of predictor candidates. In 1040, rate distortion competition occurs [Number] between predictor candidates based on CPMV and [Number] between predictor candidates based on CPMV. Different from some aspects of the described embodiments that use vectors only to calculate the cost for CPMV for the derivation of CPMV for the affine merge mode, in some cases, [Number] as CPMV [Number] By using this, relatively good prediction can be obtained. Therefore, such a motion model needs to define the CPMV to be used. In this case, in order to notify which one of

Number

Number

[0088]

Table 1

[0089]

[0107] control_point_horizontal_l0_flag[ x0 ][ y0 ] is for the block to be encoded, that is, the control point used for list L0 of the current prediction block. The array indices x0, y0 define the location (x0, y0) of the upper left luma sample of the target prediction block in relation to the upper left luma sample of the picture. If the flag is equal to 1,

Number

Number

[0090]

[0108] At least one implementation of the determination 1020 of the set of predictor candidates for the merge mode In another variation of the first and second aspects of the embodiment, at 1030, the upper left CPMV

Number

Number

Number

Number

Number

Number

[0091]

[0109] Figure 15 shows details of one embodiment of process syntax 1500 used to predict the affine motion field of the current CU being encoded or decoded in the existing affine merge mode of JEM. The input 1501 to this process / syntax 1500 is the current coding unit for which generation of the affine motion field of the sub-blocks shown in Figure 7 is desired. To select a predictor candidate based on the RD cost, at least one of the above-described embodiments (a bi-directional predictor

Number

Number

Number

Number

[0092]

[0110] In at least one implementation, a residual flag is used. 155 At 0, the flag is activated, thereby indicating that coding is being performed with residual data. At 1560, the current CU is fully coded and reconstructed (with residuals), thereby incurring the corresponding RD cost. Then, the flag is deactivated, thereby indicating that coding is being performed without residual data, and the process returns to 1560, where the CU is coded (without residuals), thereby incurring the corresponding RD cost. The lowest RD cost between the two previous ones indicates whether the residuals must be coded (normal or skip). Then, this best RD cost is exposed to competition with other coding modes. Rate distortion determination will be described in more detail later.

[0093]

[0111] FIG. 16 is an exemplary decoder according to a general aspect of at least one embodiment. Method 1600 for dithering is shown. At 1610, method 1600 receives an index corresponding to a particular predictor candidate among a set of predictor candidates for a decoded block within a picture. In various embodiments, the particular predictor candidate is selected by the encoder from among the set of predictor candidates, and the index allows for one of the predictor candidates to be selected. At 1620, method 1600 determines at least one spatial neighboring block for the decoded block. At 1630, method 1600 determines a set of predictor candidates for the decoded block based on at least one spatial neighboring block. The predictor candidates have one or more corresponding control point motion vectors and one reference picture for the decoded block. Any of the variations described for determining the set of predictor candidates in the context of encoding method 1000 are mirrored on the decoder side as shown, for example, in FIG. 17. At 1640, method 1600 determines one or more control point motion vectors from the particular predictor candidate. At 1650, method 1600 determines a corresponding motion field for the particular predictor candidate based on one or more corresponding control point motion vectors. In various embodiments, the motion field is based on a motion model, in which case the corresponding motion field identifies motion vectors used for prediction of sub-blocks of the decoded block. The motion field is the upper left and upper right control point motion vectors

Number

Number

Number

Number

[0094]

[0112] FIG. 17 shows details for illustration of determination 1020 of a set of predictor candidates for the merge mode of decoding method 16 00. Various embodiments described in relation to FIG. 11 regarding details of the encoding method are repeated here.

[0095]

[0113] The inventors have found that one aspect of the existing affine merge process described above surrounds It has been recognized that one and only one motion vector is systematically utilized to propagate the affine motion field from past and adjacent CUs to the current CU. In various situations, the inventors have further recognized that this approach can be disadvantageous, for example, because it does not select the optimal motion vector predictor. Furthermore, as already mentioned above, the selection of this predictor is composed only of the first past and adjacent CU coded in affine mode in the ordered set (A, B, C, D, E). In various situations, the inventors have further recognized that this restricted selection can be disadvantageous, for example, because better predictors may be available. Therefore, the existing process in the current JEM does not take into account the fact that some of the potential past and adjacent CUs around the current CU may also use affine motion, and that a different CU found to have used affine motion may be a better predictor for the motion information of the current CU.

[0096]

[0114] Therefore, the inventors have recognized potential benefits in several ways to improve the prediction of the current CU affine motion vectors not utilized by the existing JEM codec.

[0097]

[0115] FIG. 18 is an exemplary system 1 in which various aspects of the exemplary embodiment can be implemented. ​A block diagram of 800 is shown. System 1800 may be implemented as a device including various components described below and is configured to execute the processes described above. Examples of such devices include, without limitation, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. As shown in FIG. 18, System 1800 may be communicatively coupled via a communication channel to other similar systems and to a display, and may also be communicatively coupled to implement all or a portion of the exemplary video systems known to those skilled in the art.

[0098]

[0116] Various embodiments of System 1800, as described above, include at least one processor 1810 configured to execute instructions loaded therein to implement various processes. Processor 1810 may include an embedded memory, an input / output interface, and various other circuits known in the art. System 1800 may also include at least one memory 1820 (e.g., a volatile memory device, a non-volatile memory device). In addition, System 1800 may include a storage device 1840, which may include non-volatile memory including, without limitation, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. Storage device 1840 may have, by way of non-limiting example, an internal storage device, a mounted storage device, and / or a network-accessible storage device. System 1800 may also include an encoder / decoder module 1830 configured to process data to provide encoded video and / or decoded video, and encoder / decoder module 1830 may include its own processor and memory.

[0099]

[0117] The encoder / decoder module 1830 performs encoding and / or decoding. 18. The encoder / decoder module 1830 represents one or more modules that may be included within an apparatus to perform a coding function. As is known, such an apparatus may include one or both of an encoding and decoding module. In addition, the encoder / decoder module 1830 may be implemented as a separate element of the system 1800 or may be embedded within one or more processors 1810 as a combination of hardware and software, as is known to those skilled in the art.

[0100]

[0118] On one or more processors 1810 to execute the various processes described above. The loaded program code may be stored in storage device 1840 and subsequently loaded onto memory 1820 for execution by processor 1810. According to an example embodiment, one or more of processor(s) 1810, memory 1820, storage device 1840, and encoder / decoder module 1830 may store one or more of a variety of items including, without limitation, input video, decoded video, bitstream, equations, formulas, matrices, variables, operations, and operation logic during execution of the above-described processes.

[0101]

[0119] The system 1800 may also communicate with other devices via a communication channel 1860. It can also include a communication interface 1850 that enables communication between them. The communication interface 1850 can include a transceiver configured to transmit and receive data to and from a communication channel 1860 without limitation. The communication interface 1850 can include, without limitation, a modem or a network card, and the communication channel 1850 can be implemented in a wired and / or wireless medium. The various components of the system 1800 may be connected or communicatively coupled, without limitation, by using various suitable connections including an internal bus, wires, and printed circuit boards (not shown in FIG. 18).

[0102]

[0120] The exemplary embodiments can be implemented by a computer software implemented by the processor 1810 or hardware, or by a combination of hardware and software. As a non-limiting example, the exemplary embodiments can be implemented by one or more integrated circuits. The memory 1820 can be of any type suitable for the technical environment and can be implemented, as a non-limiting example, by using any suitable data storage technology such as an optical memory device, a magnetic memory device, a semiconductor-based memory device, a fixed memory, and a removable memory. The processor 1810 can be of any type suitable for the technical environment and can include, as a non-limiting example, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0103]

[0121] The implementations described herein can be, for example, methods or processes, devices ​​It can be implemented in a device, software program, data stream, or signal. Also, even if it is described only in the context of a single implementation form (for example, described only as a method), the implementation form of the described features can be implemented in other forms (for example, a device or a program). The device can be implemented, for example, in appropriate hardware, software, and firmware. The method can be implemented, for example, in a device such as a computer, a microprocessor, an integrated circuit, or a programmable logic device, which means a general processing device, for example, a processor. Also, the processor can be, for example, a computer, a cell phone, a portable / personal digital assistant (PDA), and other devices that facilitate the communication of information among end users, etc., including communication devices. 」) and other devices that facilitate the communication of information among end users, etc., including communication devices.

[0104]

[0122] Furthermore, those skilled in the art can easily understand that the exemplary HEVC encoder 100 shown in FIG. 1 and the exemplary HEVC decoder shown in FIG. 3 can be modified according to the above teachings of the present disclosure to implement the disclosed improvements to the existing HEVC standard in order to achieve relatively good compression / decompression. For example, the motion compensation 170 and motion estimation 175 in the exemplary encoder 100 of FIG. 1, and the motion compensation 375 in the exemplary decoder of FIG. 3 can be modified according to the disclosed teachings to implement one or more exemplary aspects of the present disclosure, including providing improved affine merge prediction to the existing JEM.

[0105]

[0123] Only "one embodiment" or "an embodiment" or "one implementation" or "an implementation" Rather, reference to these other variations means that the particular features, structures, characteristics, etc. described in connection with the embodiments are included in at least one embodiment. Thus, the terms "in one embodiment" or "in an embodiment" appearing in various places throughout this specification may be used interchangeably. )" or "in one implementation" or "in one implementation The appearance of the phrase "in an implementation," as well as any other variations, are not necessarily all referring to the same embodiment.

[0106]

[0124] In addition, the application or claims thereto may refer to "determining" various pieces of information. Determining the information may include, for example, one or more of estimating the information, calculating the information, predicting the information, or retrieving the information from a memory.

[0107]

[0125] Moreover, this application and its claims refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or inferring information.

[0108]

[0126] In addition, the application or claims thereto may refer to "receiving" various pieces of information. "Receiving" is intended to be a broad term, just like "accessing." Receiving information can refer to, for example, It may include one or more of doing or obtaining information (e.g., from memory). Further, "receiving" usually involves, in one way or another, operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0109]

[0127] As will be apparent to those skilled in the art, the implementation forms can generate various signals formatted to carry information that can be stored or transmitted, for example. As will be apparent to those skilled in the art, the implementation forms can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described implementation forms. For example, the signal can be formatted to carry a bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the high-frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

Claims

1. 1. A method for video encoding, comprising: determining at least one spatial neighboring block for a block to be encoded in a picture; determining, for the block to be encoded, a set of predictor candidates for an inter-coding mode based on the at least one spatially neighboring block, the predictor candidates comprising one or more control point motion vectors and one reference picture; determining, for the to-be-encoded block and for each predictor candidate, a motion field based on a motion model and based on the one or more control point motion vectors of the predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the to-be-encoded block; selecting a predictor candidate from the set of predictor candidates based on a rate-distortion measurement during prediction in response to the motion field determined for each predictor candidate; encoding the block based on the motion field for the selected predictor candidate; The method according to claim 1,

2. 1. A method for video decoding, comprising: determining at least one spatial neighboring block for a block to be decoded; determining, for the block to be decoded, a set of predictor candidates for an inter-coding mode based on the at least one spatially neighboring block, the predictor candidates comprising one or more control point motion vectors and one reference picture; determining one or more control point motion vectors from a particular predictor candidate for the block being decoded; determining, for the decoded block, a motion field based on a motion model and based on the one or more control point motion vectors for the decoded block, the motion field identifying motion vectors used for prediction of sub-blocks of the decoded block; decoding the block based on the determined motion field; and The method according to claim 1,

3. 1. An apparatus for video encoding, comprising: means for determining at least one spatial neighboring block for a block to be encoded in a picture; means for determining, for the block to be encoded, a set of predictor candidates for an inter-coding mode based on the at least one spatially neighboring block, the predictor candidates comprising one or more control point motion vectors and a reference picture; means for determining, for the block to be encoded and for each predictor candidate, a motion field based on a motion model and based on the one or more control point motion vectors of the predictor candidate, the motion field comprising predictions of sub-blocks of the block to be encoded; a determining means for determining a motion vector to be used for the estimation; means for selecting a predictor candidate from the set of predictor candidates based on a rate-distortion measurement during prediction in response to the motion field determined for each predictor candidate; means for encoding the block based on the motion field for the selected predictor candidate from the set of predictor candidates; An apparatus having

4. 1. An apparatus for video decoding, comprising: means for determining at least one spatial neighboring block for a block being decoded; means for determining, for the block to be decoded, a set of predictor candidates for inter-coding based on the at least one spatial neighboring block, the predictor candidates comprising one or more control point motion vectors and a reference picture; means for determining, for the block being decoded, one or more control point motion vectors from a particular predictor candidate; means for determining, for the decoded block, a motion field based on a motion model and based on the one or more control point motion vectors for the decoded block, the motion field identifying motion vectors used for prediction of sub-blocks of the decoded block; means for decoding the block based on the determined motion field; An apparatus having

5. 5. The method of claim 1 or 2 or the apparatus of claim 3 or 4, wherein the at least one spatial neighboring block comprises a spatial neighboring block of the block to be encoded or decoded among an adjacent upper left corner block, an adjacent upper right corner block, and an adjacent lower left corner block.

6. 6. The method of claim 1, 2 or 5, wherein the motion information relating to the at least one of the spatially neighboring blocks comprises translational motion information.

6. Apparatus according to claim 3, 4 or 5.

7. The method according to claim 1, 2 or 5, wherein the motion information relating to all of the at least one spatially neighboring blocks comprises affine motion information.

6. Apparatus according to claim 3, 4 or 5.

8. Determining the set of predictor candidates for an inter-coding mode for the block to be encoded or decoded includes: determining a top left list of spatial neighboring blocks of the block to be encoded or decoded from among the adjacent top left corner blocks, a top right list of spatial neighboring blocks of the block to be encoded or decoded from among the adjacent top right corner blocks, and a bottom left list of spatial neighboring blocks of the block to be encoded or decoded from among the adjacent bottom left corner blocks; Selecting at least one triplet of spatially adjacent blocks, where each spatially adjacent block of the triplet belongs to the top left list, the top right list, and the bottom left list, respectively, and the reference pictures used for prediction of each spatially adjacent block of the triplet are the same. And, determining, for the encoded and decoded block, one or more control point motion vectors for a top left corner, an top right corner, and a bottom left corner of the block based on motion information associated with each selected triplet of spatially neighboring blocks; and The method of claim 1 , 2, 5, 6, or 7, wherein the predictor candidates include the determined one or more control point motion vectors and the reference picture.

9. Determining the set of predictor candidates for an inter-coding mode for the block to be encoded or decoded includes: evaluating the at least one selected triplet of spatially neighboring blocks according to one or more criteria based on the one or more control point motion vectors determined for the block being encoded or decoded; and The method of claim 8 , wherein the predictor candidates are sorted within the set of predictor candidates for an inter-coding mode based on the evaluating.

10. Determining the set of predictor candidates for an inter-coding mode for the block to be encoded or decoded includes: determining a top-left list of spatial neighboring blocks of the block to be encoded or decoded from among adjacent top-left corner blocks, and a top list of spatial neighboring blocks of the block to be encoded or decoded from among adjacent top-right corner blocks; selecting at least one pair of spatial neighboring blocks, each of which belongs to the top-left list and the top-right list, respectively, and the reference pictures used for prediction of each of the spatial neighboring blocks of the pair are identical; for the block being encoded or decoded, determining a control point motion vector for the top left corner of the block based on motion information associated with spatially neighboring blocks in the top left list, and determining a control point motion vector for the top left corner of the block based on motion information associated with spatially neighboring blocks in the top left list; and The method of claim 1 , 2, 5, 6, or 7, wherein the predictor candidates include the top-left and top-right control point motion vectors and the reference picture.

11. 11. The method of claim 10, wherein a bottom-left list is used instead of the top-right list, the bottom-left list having spatial neighbors of the block being encoded or decoded among adjacent bottom-left corner blocks, and a bottom-left control point motion vector is determined.

12. The motion model is an affine model, and for each position (x,y) inside the block being encoded or decoded the motion field is determined by: [0010] Here, (v 0x , v 0y ) and (v 2x , v 2y ) are the control point motion vectors used to generate the motion field, and (v 0x , v 0y ) corresponds to the control point motion vector of the upper left corner of the block being encoded or decoded, and (v 2x , v 2y 12. The method or apparatus of claim 1, wherein a corresponds to the control point motion vector of the bottom left corner of the block being encoded or decoded, and h is the height of the block being encoded or decoded.

13. encoding or obtaining an indication of the motion model to be used for the encoded or decoded block, the motion model being based on a control point motion vector of the top left corner and the control point motion vector of the bottom left corner, or the motion model being based on a control point motion vector of the top left corner and the control point motion vector of the top right corner; 13. The method of any one of claims 1, 2, 5 to 12, further comprising:

14. 13. A method according to claim 1, 2, 5 to 12, wherein the motion model used for the block being encoded or decoded is derived implicitly and the motion vector is based on a control point motion vector of the upper left corner and the control point motion vector of the lower left corner, or the motion model is based on a control point motion vector of the upper left corner and the control point motion vector of the upper right corner.

15. 15. The method or apparatus of any one of claims 1 to 14, wherein the inter-coding mode is one of a merge mode or an advanced motion vector prediction mode.

16. The method of claim 1 or the apparatus of claim 3, wherein an index for the selected predictor candidate from the set of predictor candidates is encoded.

17. The method of claim 2 or the apparatus of claim 4, wherein an index for the particular predictor candidate from the set of predictor candidates is decoded.

18. A non-transitory computer readable medium is presented that includes data content generated according to the method or apparatus of any of the preceding descriptions.

19. A computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described above.

Citation Information

Patent Citations

  • Image prediction method and related device

    WO2016141609A1

  • Method and apparatus for affine merge mode prediction for video coding system

    WO2017118409A1

  • Image encoding / decoding method and device

    WO2017133243A1

  • Method and apparatus of video coding with affine motion compensation

    WO2017148345A1

  • Affine motion prediction for video coding

    WO2017200771A1