Improved predictor candidates for motion compensation.
By determining predictor candidates based on spatial neighboring blocks and reference pictures, the method addresses the limitation of control point motion vectors in affine merge and AMVP modes, improving video encoding compression performance.
Patent Information
- Application Number
- JP2025025983
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-10-05
- Filing Date
- 2025-02-20
- Publication Date
- 2026-01-28
- Estimated Expiration
- 2038-10-04
AI Technical Summary
The set of control point motion vectors (CPMVs) used as predictors in affine merge and Advanced Motion Vector Prediction (AMVP) modes is limited, limiting the overall compression performance of high-compression video encoding techniques.
A method for video encoding and decoding that determines a set of predictor candidates based on spatial neighboring blocks, using control point motion vectors and reference pictures, and selects the best candidate through rate-distortion determination, encoding an index for the selected predictor.
Improves the compression performance of video encoding by expanding the set of predictor candidates, enhancing the motion compensation process.
Smart Images

Figure 0007808217000096 
Figure 0007808217000097 
Figure 0007808217000098
Abstract
Description
[Technical Field]
[0001] Technical Field [1] At least one of the present embodiments generally relates to, for example, video encoding. The present invention relates to a method or apparatus for motion compensation in an inter-coding mode (merge mode or AMVP), and more particularly to a method and apparatus for a video encoder or video decoder for selecting one predictor candidate from a set of multiple predictor candidates based on a motion model, such as an affine model. [Background technology]
[0002] background [2] To achieve high compression efficiency, image and video coding methods typically To exploit spatial and temporal redundancy in the content, prediction, including motion vector prediction, is used and transformed. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original and predicted image, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction.
[0003] [3] Recent additions to high-compression techniques include affine modeling-based This involves the use of motion models. Specifically, affine models are used for motion compensation for encoding and decoding video pictures. Generally, affine modeling involves the use of at least two parameters, such as two control point motion vectors (CPMVs) representing the motion at the respective corners of a block of a picture, allowing the derivation of a motion field for the entire block of a picture, to simulate, for example, rotation and similarity (zoom). However, the set of control point motion vectors (CPMVs) that can potentially be used as predictors in merge mode is limited. Therefore, in affine merge and Advanced Motion Vector Prediction (AMVP) modes, A method that would improve the overall compression performance of the subject high compression techniques by improving the performance of the motion models used in the compression techniques is desirable. Summary of the Invention
[0004] overview [4] An object of the present invention is to overcome at least one of the drawbacks of the prior art. To this end, according to one general aspect of at least one embodiment, a method of video encoding is presented, the method comprising: determining, for a block to be encoded in a picture, at least one spatial neighboring block; determining, for the block to be encoded, a set of predictor candidates for an inter-coding mode based on the at least one spatial neighboring block, the predictor candidates having one or more control point motion vectors and a reference picture; determining, for the block to be encoded and for each predictor candidate, a motion field based on a motion model and on the one or more control point motion vectors of the predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be encoded; selecting a predictor candidate from a set of predictor candidates based on a rate-distortion determination during prediction in response to the selected motion field, encoding the block based on the motion field of the selected predictor candidate, and encoding an index for the selected predictor candidate from the set of predictor candidates, wherein one or more control point motion vectors and reference pictures are used for predicting the block to be encoded based on motion information associated with the block.
[0005] [5] According to another general aspect of at least one embodiment, a video decoding a motion vector predictor for a block to be decoded based on the one or more control point motion vectors and a reference picture; and a motion vector predictor for the block to be decoded based on the one or more control point motion vectors.
[0006] [6] According to another general aspect of at least one embodiment, a method for video coding An apparatus is presented, the apparatus including: means for determining at least one spatially neighboring block for a block to be encoded in a picture; means for determining a set of predictor candidates for an inter-coding mode based on the at least one spatially neighboring block for the block to be encoded, the predictor candidate having one or more control point motion vectors and one reference picture; means for selecting one predictor candidate from the set of predictor candidates for the block to be encoded, and for each predictor candidate, based on a motion model and on the one or more control point motion vectors of the predictor candidate; the encoding method includes: means for determining a motion field based on a sub-block motion vector of a block being encoded, the motion field identifying a motion vector used for predicting a sub-block of the block being encoded; means for selecting, in response to the motion field determined for each predictor candidate, one predictor candidate from a set of predictor candidates based on a rate-distortion measurement during prediction; means for encoding the block based on a corresponding motion field of the selected predictor candidate from the set of predictor candidates; and means for encoding an index for the selected predictor candidate from the set of predictor candidates.
[0007] [7] According to another general aspect of at least one embodiment, video decoding An apparatus for a picture coding method is presented, the apparatus including: means for receiving, for a block to be decoded in a picture, an index corresponding to a particular predictor candidate from a set of predictor candidates for an inter-coding mode; means for determining at least one spatial neighboring block for the block to be decoded; means for determining, for the block to be decoded, a set of predictor candidates for the inter-coding mode based on the at least one spatial neighboring block, the predictor candidates having one or more control point motion vectors and one reference picture; and means for receiving, for the block to be decoded, an index corresponding to a particular predictor candidate from a set of predictor candidates for an inter-coding mode based on the at least one spatial neighboring block, the predictor candidates having one or more control point motion vectors and one reference picture. means for determining, for a block to be decoded, one or more corresponding control point motion vectors from a particular predictor candidate; means for determining, for a block to be decoded, a motion field based on a motion model and based on the one or more control point motion vectors for the block to be decoded, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be decoded; and means for decoding the block based on the corresponding motion field.
[0008] [8] According to another general aspect of at least one embodiment, a video encoding An apparatus for coding is provided, the apparatus having one or more processors and at least one memory, wherein the one or more processors are configured to: determine at least one spatially neighboring block for a block to be coded in a picture; determine a set of predictor candidates for an inter-coding mode based on the at least one spatially neighboring block for the block to be coded, where the predictor candidate has one or more control point motion vectors and a reference picture; determine a motion field for the block to be coded and for each predictor candidate based on a motion model and on the one or more control point motion vectors of the predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be coded; select, in response to the motion field determined for each predictor candidate, one predictor candidate from the set of predictor candidates based on a rate-distortion determination during prediction; encode the block based on the motion field of the selected predictor candidate; and encode an index for the selected predictor candidate from the set of predictor candidates. The at least one memory is for at least temporarily storing the encoded blocks and / or the encoded indexes.
[0009] [9] According to another general aspect of at least one embodiment, video decoding An apparatus for decoding a picture is provided, the apparatus including one or more processors and at least one memory, wherein the one or more processors are configured to: receive, for a block to be decoded in a picture, an index corresponding to a particular predictor candidate from a set of predictor candidates for an inter-coding mode; determine at least one spatially neighboring block for the block to be decoded; determine, for the block to be decoded, a set of predictor candidates for the inter-coding mode based on the at least one spatially neighboring block, the predictor candidates having one or more control point motion vectors and a reference picture; determine, for the particular predictor candidate, one or more corresponding control point motion vectors for the block to be decoded; determine, for the particular predictor candidate, a motion field based on a motion model based on the one or more corresponding control point motion vectors, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be decoded; and decode the block based on the motion field. The at least one memory is for at least temporarily storing the decoded block.
[0010]
[10] According to another general aspect of at least one embodiment, the at least one spatial neighboring block comprises a spatial neighboring block of the block being encoded or decoded from among an adjacent upper left corner block, an adjacent upper right corner block, and an adjacent lower left corner block.
[0011]
[11] According to another general aspect of at least one embodiment, a spatially adjacent block The motion information associated with at least one of the images comprises non-affine motion information. A non-affine motion model is a translational motion model, in which only one motion vector representing a translation is coded in the model.
[0012]
[12] According to another general aspect of at least one embodiment, the motion information associated with all of the at least one spatially neighboring block comprises affine motion information.
[0013]
[13] According to another general aspect of at least one embodiment, the set of predictor candidates includes unidirectional predictor candidates and bidirectional predictor candidates.
[0014] According to another general aspect of at least one embodiment, the method may further include determining an upper-left list of spatial neighboring blocks of a block to be encoded or decoded from adjacent upper-left corner blocks, an upper-right list of spatial neighboring blocks of a block to be encoded or decoded from adjacent upper-right corner blocks, and a bottom-left list of spatial neighboring blocks of a block to be encoded or decoded from adjacent lower-left corner blocks; selecting at least one triplet of spatial neighboring blocks, wherein each spatial neighboring block of the triplet belongs to the upper-left list, the upper-right list, and the bottom-left list, respectively, and wherein reference pictures used for predicting the spatial neighboring blocks of the triplet are identical; and determining, for the block to be encoded or decoded, one or more control point motion vectors for the upper-left corner, the upper-right corner, and the bottom-left corner of the block based on motion information respectively associated with each spatial neighboring block of the selected triplet, wherein predictor candidates include the determined one or more control point motion vectors and reference pictures.
[0015]
[15] According to another general aspect of at least one embodiment, the method may further include evaluating at least one selected triplet of spatially neighboring blocks according to one or more criteria based on one or more control point motion vectors determined for the block being encoded or decoded, and in which the predictor candidates are sorted within the set of predictor candidates for the inter-coding mode based on the evaluation.
[0016]
[16] According to another general aspect of at least one embodiment, the one or more criteria include a validity check according to Equation 3 and a cost according to Equation 4.
[0017]
[17] According to another general aspect of at least one embodiment, the cost of a bidirectional predictor candidate is the average of the cost associated with its first reference picture list and the cost associated with its second reference picture list.
[0018]
[18] According to another general aspect of at least one embodiment, a method includes determining a top-left list of spatial neighboring blocks of a block to be encoded or decoded from adjacent top-left corner blocks and a top-right list of spatial neighboring blocks of a block to be encoded or decoded from adjacent top-right corner blocks; selecting at least one pair of spatial neighboring blocks, each of which belongs to the top-left list and the top-right list, respectively, and each of which uses the same reference picture for prediction; and selecting a block for the block to be encoded or decoded based on motion information associated with the spatial neighboring blocks in the top-left list. determining a control point motion vector for the top left corner of the block based on a control point motion vector for the top left corner of the block and motion information associated with spatially neighboring blocks in a top left list, wherein predictor candidates include the top left and top right control point motion vectors and reference pictures.
[0019]
[19] According to another general aspect of at least one embodiment, a bottom-left list is used instead of an top-right list, the bottom-left list having spatial neighbors of the block being encoded or decoded among the adjacent bottom-left corner blocks, and a bottom-left control point motion vector is determined.
[0020]
[20] According to another general aspect of at least one embodiment, the motion model is an affine model, and the motion field at each position (x,y) inside the block being encoded or decoded is determined by the following equation:
number
[0021]
[21] where (v 0x ,v 0y ) and (v 2x ,v 2y ) are the control point motion vectors used to generate the motion field, and (v 0x ,v 0y ) corresponds to the control point motion vector of the upper left corner of the block being encoded or decoded, and (v 2x ,v 2y ) corresponds to the control point motion vector of the bottom left corner of the block being encoded or decoded, and h is the height of the block being encoded or decoded.
[0022]
[22] According to another general aspect of at least one embodiment, the method may further include encoding or obtaining indication of a motion model to be used for the block being encoded or decoded, the motion model being based on a control point motion vector of the top left corner and a control point motion vector of the bottom left corner, or the motion model being based on a control point motion vector of the top left corner and a control point motion vector of the top right corner.
[0023]
[23] According to another general aspect of at least one embodiment, the motion model used for a block being encoded or decoded is implicitly derived, the motion model being based on a control point motion vector of the top left corner and a control point motion vector of the bottom left corner, or the motion model being based on a control point motion vector of the top left corner and a control point motion vector of the top right corner.
[0024]
[24] According to another general aspect of at least one embodiment, decoding or encoding a block based on a corresponding motion field includes decoding or encoding, respectively, based on a predictor for the sub-block, the predictor being informed by the motion vector.
[0025]
[25] According to another general aspect of at least one embodiment, the number of spatially adjacent blocks is at least five or at least seven.
[0026]
[26] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is presented that includes data content generated according to the method or apparatus of any of the preceding descriptions.
[0027]
[27] According to another general aspect of at least one embodiment, there is provided a signal having video data generated in accordance with the method or apparatus of any of the preceding descriptions.
[0028]
[28] One or more of the present embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the above-described methods. The present embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the above-described methods. The present embodiments also provide methods and apparatus for transmitting a bitstream generated according to the above-described methods. The present embodiments also provide a computer program product including instructions for performing any of the described methods. [Brief explanation of the drawings]
[0029] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1]
[29] shows a block diagram of one embodiment of a High Efficiency Video Coding (HEVC) video encoder. [Figure 2A]
[30] is a plot example illustrating the generation of HEVC reference samples. [Figure 2B]
[31] A drawing example illustrating the intra prediction direction in HEVC. [Figure 3]
[32] A block diagram of one embodiment of an HEVC video decoder is shown. [Figure 4]
[33] shows an example of the Coding Tree Unit (CTU) and Coding Tree (CT) concepts for representing a compressed HEVC picture. [Figure 5]
[34] shows an example of division of a coding tree unit (CTU) into coding units (CU), prediction units (PU), and transform units (TU). [Figure 6]
[35] shows an example of an affine model as a motion model used in JEM (Joint Exploration Model). [Figure 7]
[36] shows an example of an affine motion vector field based on 4x4 sub-CUs used in JEM (Joint Exploration Model). [Figure 8A]
[37] shows an example of motion vector prediction candidates for affine inter-CU. [Figure 8B]
[38] shows an example of motion vector prediction candidates in affine merge mode. [Figure 9]
[39] present an example of spatial derivation of affine control point motion vectors in the case of the affine merge mode motion model. [Figure 10]
[40] illustrates an exemplary encoding method according to a general aspect of at least one embodiment. [Figure 11]
[41] illustrates another example of an encoding method according to a general aspect of at least one embodiment. [Figure 12] 42 illustrates an example of motion vector prediction candidates for an affine merge mode CU according to a general aspect of at least one embodiment. [Figure 13] 43 illustrates an example of an affine motion vector field based on an affine model as a motion model and its corresponding 4x4 sub-CUs according to a general aspect of at least one embodiment. [Figure 14] 44 illustrates another example of an affine motion vector field based on an affine model as a motion model and its corresponding 4x4 sub-CUs according to a general aspect of at least one embodiment. [Figure 15]
[45] presents an example of a known process / syntax for evaluating affine merge modes of CUs in JEM. [Figure 16]
[46] illustrates an exemplary decoding method according to a general aspect of at least one embodiment. [Figure 17]
[47] illustrates another example of an encoding method according to a general aspect of at least one embodiment. [Figure 18]
[48] Figure 4 shows a block diagram of an example device in which various aspects of the embodiments may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0030] Detailed Description
[49] It should be understood that the figures and descriptions have been simplified for purposes of clarity to show elements suitable for a clear understanding of the present principles, while excluding many other elements found in conventional coding and / or decoding devices. The terms "first" and "second" are used herein to describe various elements. Although the terms "a," "b," "c," "d," "e," "f," "g," "h," "i," "j," "j," "k," "k," "ma," "ma," "ma," "ma," "p," "s," "s," "u ...s," "u," "s," "s," "s," "s," "s," "
[0031]
[50] Various embodiments are described in the context of the HEVC standard. However, the present principles are not limited to HEVC and may be applied to other standards, recommendations, and extensions thereof, including HEVC or HEVC extensions, such as, for example, Format Range (RExt), Scalability (SHVC), Multiview (MV-HEVC) extensions, and H.266. Various embodiments are described in the context of encoding / decoding slices. They may be applied to encode / decode entire pictures or entire sequences of pictures.
[0032]
[51] Various methods are described above, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be varied or combined.
[0033]
[52] Figure 1 shows an example High Efficiency Video Coding (HEVC) encoder 100. HEVC is a standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). It is a compression standard developed for video encoding and decoding (see, for example, "ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265").
[0034]
[53] In HEVC, to encode a video sequence having one or more pictures, the pictures are partitioned into one or more slices, where each slice may contain one or more slice segments, which are organized as coding units, prediction units, and transform units.
[0035]
[54] In this application, the terms "reconstructed" and "decoding" are used interchangeably. The terms "decoded" and "encoded" may be used interchangeably. The terms "encoded" or "coded" may also be used interchangeably, and the terms "picture" and "frame" may also be used interchangeably. Typically, the term "reconstructed" is used on the encoder side, while the term "decoded" is used on the decoder side, but this is not required.
[0036]
[55] The HEVC specification distinguishes between "blocks" and "units," where a "block" addresses a specific area within a sample array (e.g., luma, Y), and a "unit" includes a grouped block of all encoded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data (e.g., motion vectors) associated with the block.
[0037]
[56] For coding, a picture is partitioned into square-shaped coding tree blocks (CTBs) with configurable size, and a contiguous set of coding tree blocks is grouped as a slice. A coding tree unit (CTU) contains the CTBs of the encoded color components. The CTB is the root of a quadtree partitioning into coding blocks (CBs), and a coding block is It may be partitioned into one or more prediction blocks (PB) and quad-tree partitioned into transform blocks (TB). Corresponding to coding blocks, prediction blocks, and transform blocks, a coding unit (CU) includes a tree-structured set of prediction units (PUs) and transform units (TUs), where a PU includes prediction information for all color components and a TU includes the residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luma component apply to the corresponding CU, PU, and TU. In this application, the term "block" may be used to refer to any of, for example, a CTU, CU, PU, TU, CB, PB, and TB. Additionally, "block" may be used to refer to macroblocks and partitions defined in H.264 / AVC or other video coding standards, and more generally to arrays of data of various sizes.
[0038]
[57] In the exemplary encoder 100, pictures are encoded by the encoder elements, as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using either intra or inter mode. When encoding in intra mode, the CU performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder determines (105) one of intra or inter modes to use for encoding the CU, and a prediction mode flag indicates the intra / inter decision. A prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0039]
[58] In intra modes, CUs are predicted from reconstructed neighboring samples within the same slice. HEVC offers a set of 35 intra prediction modes, including one DC prediction mode, a planar prediction mode, and 33 angular prediction modes. The intra prediction reference is reconstructed from rows and columns neighboring the current block. The reference spans twice the block size in the horizontal and vertical directions by using available samples from previously reconstructed blocks. When an angular prediction mode is used for intra prediction, the reference sample can be replicated along the direction indicated by the angular prediction mode.
[0040]
[59] The applicable luma intra-prediction modes for the current block can be coded using two different options: If the applicable mode is included in a constructed list of three most probable modes (MPM), the mode is signaled by an index in the MPM list; otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra-prediction modes of the upper and left neighboring blocks.
[0041]
[60] In the case of an inter CU, the corresponding coding block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how inter prediction is performed. Motion information (i.e., motion vectors and reference picture indexes) can be signaled in two ways: merge mode and advanced motion vector prediction (AMVP).
[0042]
[61] In merge mode, a video encoder or decoder compiles a candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates in the candidate list. At the decoder side, the motion vector (MV) and reference The reference picture index has been reconstructed.
[0043]
[62] The set of possible candidates in merge mode consists of spatially adjacent candidates, temporal candidates, and generated candidates. Figure 2A shows the locations of five spatial candidates {a1, b2, b0, a0, b2} for the current block 21, where a0 and a1 are located on the left side of the current block, and b1, b0, b2 are located on the top of the current block. For each candidate location, availability is checked in the order a1, b1, b0, a0, b2, and then redundancies within the candidates are removed.
[0044]
[63] For the derivation of temporal candidates, motion vectors of clustered locations in reference pictures can be used. Applicable reference pictures are selected on a slice-by-slice basis and signaled in the slice header, and the reference index for the temporal candidates is i ref = 0. If the POC distance (td) between the picture of the grouped PU and the reference picture from which the grouped PU is predicted is the same as the distance (tb) between the current picture and the reference picture containing the grouped PU, then the grouped motion vector mv col can be used directly as a temporal candidate. Otherwise, the scaled motion vector tb / td*mv col are used as temporal candidates. Depending on where the current PU is located, the lumped PU is determined by a sample location at the bottom right or center of the current PU.
[0045]
[64] The maximum number of merge candidates N is specified in the slice header. If the number of merge candidates is greater than N, only the first N-1 spatial and temporal candidates are used. Otherwise, if the number of merge candidates is less than N, the candidate set is filled up to the maximum number N with generated candidates as a combination of already existing candidates or with null candidates. Candidates used in merge mode are sometimes referred to as "merge candidates" in this application.
[0046]
[65] If the CU is signaling skip mode, applicable indices for merge candidates are signaled only if the list of merge candidates is greater than 1, and further information is not coded for the CU. In skip mode, motion vectors are applied without residual updates.
[0047]
[66] In AMVP, a video encoder or decoder compiles a candidate list based on motion vectors determined from previously coded blocks. The video encoder then uses a motion vector predictor (MVP). To identify the best possible reference picture, an index is signaled in the candidate list and the motion vector difference (MVD) is signaled. At the decoder side, the motion vector (MV) is reconstructed as MVP+MVD. Also, the applicable reference picture index is explicitly coded in the PU syntax for AMVP.
[0048]
[67] In AMVP, only two spatial motion candidates are selected. The first spatial motion candidate is selected from the left position {a0, a1}, and the second is selected from the top position {b0, b1, b2}, maintaining the search order signaled within the two sets. If the number of motion vector candidates does not equal two, they may include a temporal MV candidate. If the candidate set is still not fully filled, a zero motion vector is used.
[0049]
[68] If the reference picture index of the spatial candidate corresponds to the reference picture index for the current PU (i.e., they use the same reference picture index, or they use both long-term reference pictures, independent of the reference picture list), the spatial candidate motion vector is used directly. Otherwise, if both reference pictures are short-term, the candidate motion vector is scaled according to the distance (tb) between the current picture and the reference picture of the current PU and the distance (td) between the current picture and the reference picture of the spatial candidate. Candidates used in AMVP are sometimes referred to as "AMVP candidates" in this application.
[0050]
[69] For ease of notation, a block tested by "merge" mode at the encoder side or decoded by "merge" mode at the decoder side is denoted as a "merge block," and a block tested by AMVP mode at the encoder side or decoded by AMVP mode at the decoder side is denoted as an "AMVP" block.
[0051]
[70] Figure 2B shows an example motion vector representation using AMVP. For the current block to be encoded 240, the motion vector (MV current ) can be obtained through motion estimation. The motion vector (MV left) and the motion vector (MV above ) to make MV left and MV above From MVP current Then, the motion vector difference can be calculated as MVD current =MV current -MVP current It can be calculated as:
[0052]
[71] Motion-compensated prediction can be performed by using one or more reference pictures for prediction. In P slices, only a single prediction reference can be used for inter prediction, allowing unidirectional prediction for the predicted block. In B slices, two reference picture lists are available, and unidirectional or bidirectional prediction can be used. In bidirectional prediction, one reference picture from each of the reference picture lists is used.
[0053]
[72] In HEVC, the precision of the motion information for motion compensation is one-quarter sample (also called quarter-pel or 1 / 4-pel) for the luma component and one-eighth sample (also called 1 / 8-pel) for the chroma components for a 4:2:0 configuration. For luma, a 7-tap or 8-tap interpolation filter is used for interpolation of fractional sample positions, i.e., one-quarter, one-half, and three-quarters of a full sample location in both the horizontal and vertical directions can be addressed.
[0054]
[73] The prediction residual is then transformed (125) and quantized (130). To output a bitstream, the quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145). The encoder can also skip the transform and apply quantization directly to the untransformed residual signal using the 4x4 TU method. The encoder can also bypass both the transform and quantization, i.e., the residual is directly coded without applying the transform or quantization process. In direct PCM coding, no prediction is applied and coding unit samples are directly coded into the bitstream.
[0055]
[74] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. An image block is reconstructed by combining (155) the decoded prediction residual and the predicted block. For example, an in-loop filter (165) is applied to the reconstructed picture to perform deblocking / Sample Adaptive Offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).
[0056]
[75] Figure 3 shows a block diagram of an example HEVC video decoder 300. In the example decoder 300, the bitstream is decoded by the decoder elements described below. The video decoder 300 generally performs a decoding path that is the opposite of the encoding path described in Figure 1, performing video decoding as part of the encoding of the video data.
[0057]
[76] Specifically, the decoder's input includes a video bitstream, such as may be generated by video encoder 100. The bitstream is first entropy decoded (330) to obtain transform efficiency, motion vectors, and other coded information. Transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. An image block is reconstructed by combining (355) the decoded prediction residual and the predicted block. The predicted block may be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). As described above, AMVP and merge mode techniques can be used to derive motion vectors for motion compensation, which may use interpolation filters to calculate interpolated values for sub-integer samples of the reference block. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0058]
[77] As mentioned above, in HEVC, motion-compensated temporal prediction is used to exploit the redundancy that exists between successive pictures of a video. To do this, a motion vector is associated with each prediction unit (PU). As mentioned above, each CTU is represented in the compressed domain by a coding tree, which is a quadtree decomposition of the CTU, where each leaf Each CU is referred to as a coding unit (CU) and is also illustrated in FIG. 4 for CTUs 410 and 420. Each CU is then assigned some intra- or inter-prediction parameters as prediction information. To achieve this, a CU may be spatially partitioned into one or more prediction units (PUs), where each PU is assigned some prediction information. An intra- or inter-coding mode is assigned at the CU level. These concepts are further illustrated in FIG. 5 for example CTU 500 and CU 510.
[0059]
[78] In HEVC, one motion vector is assigned to each PU. This motion vector is used for motion-compensated temporal prediction of the target PU. Therefore, in HEVC, the motion model linking a predicted block and its reference block consists of a translation or calculation based on the reference block and the corresponding motion vector.
[0060]
[79] To implement improvements to HEVC, the Joint Exploration Model (JEM), a reference software and / or document, is being developed by the Joint Video Exploration Team (JVET). One JEM version (i.e., “Algorithm Description of Joint Exploration Test Model 5”, Document JVET-E1001_v2, Joint Video Exploration Team of ISO / IEC JTC1 / SC29 / WG11, 5th meeting, 12-20 January 2017, Geneva, CH) states: To improve temporal prediction, several additional motion models are supported. To do this, a PU can be spatially divided into sub-PUs, and the models can be used to assign a dedicated motion vector to each sub-PU.
[0061]
[80] More recent versions of JEM (e.g., “Algorithm Description of Joint Exploration Test Model 2”, Document JVET-B1001_v3, Joint Video Exploration Team of ISO / IEC JTC1 / SC29 / WG11, 2nd meeting, 20-26 February 2016, San Diego, USA) In JEM, CUs are no longer defined as being divided into PUs or TUs. Instead, relatively flexible CU sizes may be used, and some motion data is assigned directly to each CU. In this new codec design in the newer version of JEM, CUs can be divided into sub-CUs, and motion vectors can be calculated for each sub-CU of a divided CU.
[0062]
[81] One of the new motion models introduced in JEM is the use of an affine motion model to represent motion vectors within a CU. The motion model used is shown in Figure 6 and expressed by Equation 1 shown below. The affine motion field has the following motion vector component values at each location (x,y) inside the target block 600 in Figure 6:
number
number
number
[0063]
[82] To reduce complexity, motion vectors are computed for each 4x4 sub-block (sub-CU) of the target CU 700, as shown in Figure 7. An affine motion vector is computed from the control point motion vector for each center position of each sub-block. The resulting motion vectors are expressed at 1 / 16 pel accuracy. As a result, coding unit compensation in affine mode involves motion-compensated prediction of each sub-block with its own motion vector. These motion vectors for the sub-blocks are shown individually as arrows for each of the sub-blocks in Figure 7.
[0064]
[83] Affine motion compensation can be used in JEM in two ways: affine AMVP (AF_AMVP) mode and affine merge mode, which are introduced in the following sections.
[0065]
[84] Affine AMVP mode: In affine AMVP mode, CUs in AMVP mode whose size is greater than 8x8 can be predicted. This is signaled through a flag in the bitstream. The generation of an affine motion field for this AMVP-CU involves determining control point motion vectors (CPMVs), which are obtained by the encoder or decoder through the addition of motion vector differentials and control motion vector predictions (CPMVPs). A CPMVP is a pair of motion vector candidates obtained from the sets (A, B, C) and (D, E) shown in Figure 8A for the current CU 800 being encoded or decoded, respectively.
[0066]
[85] Affine Merge Mode: In affine merge mode, a CU-level flag signals whether the merged CU uses affine motion compensation. If so, the first available neighboring CU coded in affine mode is selected from the ordered set of candidate positions A, B, C, D, E in Figure 8B for the current CU 880 being encoded or decoded. Note that this ordered set of candidate positions in JEM is the same as the spatially neighboring candidates in HEVC merge mode shown in Figure 2A and described above.
[0067]
[86] Once the first neighboring CU in affine mode is obtained, three CPMVs from the top left, top right, and bottom left corners of the neighboring affine CU are
number
number
[0068]
[87] is the control point motion vector of the current CU
number
number
[0069]
[88] Accordingly, one general aspect of at least one embodiment is to improve the performance of affine merge mode in JEM so that the compression performance of a target video codec can be improved. Accordingly, in at least one embodiment, an improved motion compensation apparatus and method is presented for a coding / decoding unit coded in affine merge mode. The proposed improved affine mode includes determining a set of predictor candidates in affine merge mode, regardless of whether neighboring CUs are coded in affine mode.
[0070]
[89] As mentioned above, in the current JEM, the first neighboring CU coded in affine mode among the surrounding CUs is selected to predict the affine motion model associated with the current CU being encoded or decoded in affine merge mode. That is, to predict the affine motion model of the current CU in affine merge mode, the first neighboring CU coded in affine mode among the ordered set (A, B, C, D, E) of FIG. 8B coded in affine mode is selected. Candidate U is selected.
[0071]
[90] Thus, at least one embodiment provides the best coding efficiency when coding the current CU in affine merge mode by improving affine merge prediction candidates, thereby generating new motion model candidates from the motion vectors of neighboring blocks used as CPMVPs. Also, during merging decoding to obtain predictors from the signaled index, the corresponding new motion model candidates generated from the motion vectors of neighboring blocks used as CPMVPs have also been determined. Thus, at a general level, the improvement of this embodiment can be achieved by, for example, Constructing a set of unidirectional or bidirectional affine merge predictor candidates based on neighboring CU motion information, regardless of whether the neighboring CUs are coded in affine mode (for the encoder / decoder); Constructing a set of candidate affine merge predictors by selecting control point motion vectors from only two adjacent blocks, e.g., (for encoder / decoder) a top left (TL) block and a top right (TR) or bottom left (BL) block; When evaluating predictor candidates, adding a penalty for unidirectional predictor candidates compared to bidirectional predictor candidates (for encoder / decoder); (For encoders / decoders) adding neighboring blocks only if they are affine coded, and / or (For encoder / decoder) signaling / decoding notification of the motion model used for the control point motion vector predictor of the current CU; It has.
[0072]
[91] describes an encoding / decoding method based on merge mode, but the principles also apply to AMVP (i.e., affine inter) mode. Advantageously, various embodiments for generating predictor candidates can be derived unambiguously for affine AMVP.
[0073]
[92] The present principles are advantageously implemented in the encoder, in the motion estimation module 175 and motion compensation module 170 of FIG. 1, or in the decoder, in the motion compensation module 275 of FIG.
[0074]
[93] Accordingly, Figure 10 illustrates an exemplary encoding method 1000 according to one general aspect of at least one embodiment. At 1010, the method 1000 determines at least one spatial neighboring block for a block to be encoded in a picture. For example, as shown in Figure 12, neighboring blocks of the block to be encoded CU are determined from A, B, C, D, E, F, and G. At 1020, the method 1000 determines a set of predictor candidates for an affine merge mode from the spatial neighboring blocks for the block to be encoded in the picture. Further details for such determination are provided below in connection with Figure 11. While in JEM affine merge mode, the first neighboring block from the ordered list F, D, E, G, A depicted in Figure 12, coded in affine mode, is used as the predictor, according to one general aspect of this embodiment, multiple predictor candidates are evaluated for affine merge mode, where neighboring blocks can be used as predictor candidates regardless of whether they are coded in affine mode. According to one particular aspect of this embodiment, the predictor candidates are one or more corresponding The motion compensation prediction comprises a control point motion vector and one reference picture. The reference picture is used for the prediction of at least one spatially neighboring block based on motion compensation in relation to one or more corresponding control point motion vectors. The reference picture and one or more corresponding control point motion vectors are derived from motion information associated with at least one of the spatially neighboring blocks. Naturally, the motion compensation prediction can be performed by using one or two reference pictures for prediction according to unidirectional or bidirectional prediction. Thus, two reference picture lists are available to store the reference picture (in the case of non-affine mode, i.e., translation mode with one motion vector) and the motion information from which the motion vectors / control point motion vectors (in the case of affine mode) are derived. In the case of non-affine motion information, the one or more corresponding control point motion vectors correspond to the motion vectors of the spatially neighboring blocks. In the case of affine motion information, the one or more corresponding control point motion vectors are, for example, the control point motion vectors of the spatially neighboring blocks.
number
number
[0075]
[94] Figure 11 shows exemplary details of determining 1020 a set of predictor candidates for merge mode of encoding method 1000 according to an aspect of at least one embodiment, particularly adapted when neighboring blocks have non-affine motion information. However, the method is compatible with neighboring blocks with affine motion information. Those skilled in the art will understand that in cases where neighboring blocks are coded using an affine model, a motion model of the affine-coded neighboring block may be used to determine predictor candidates for affine mode, as detailed above with reference to Figure 9. At 1021, three lists are determined based on the spatial position of the block relative to the block being encoded by considering at least one neighboring block of the block being encoded. For example, considering blocks A, B, C, D, E, F, and G in Figure 12, a first list, called the top-left list, is generated with adjacent top-left corner blocks A, B, and C, a second list, called the top-right list, is generated with adjacent top-right corner blocks D and E, and a third list is generated with adjacent bottom-left corner blocks F and G. According to a first aspect of at least one embodiment, triplets of spatially neighboring blocks are selected, where each spatially neighboring block of the triplet belongs to the top-left list, the top-right list, and the bottom-left list, respectively, and where the reference pictures used for predicting each spatially neighboring block of the triplet are the same. Thus, the triplet corresponds to the selection of {A, B, C}, {D, E}, and {F, G}. At 1022, a first selection of 12 possible triplets {A, B, C}, {D, E}, {F, G} is applied based on the usage of the same reference picture for each neighboring block of the motion compensation triplet. Thus, only triplets A, D, F are selected, where the same reference picture (identified by its index in the reference picture list) is used for the prediction of neighboring blocks A, D, and F.This first criterion ensures coherence in the prediction. At 1023, one or more control point motion vectors of the current CU.
number
number
number
number
number
number
number
number
number
number
[0076] At 1024, an optional second selection among the candidate CPMVPs is applied based on one or more criteria. According to the first criterion, the candidate CPMVPs are further checked for a block to be encoded of height H and width W for verification using Equation 3, where X and Y are the horizontal and vertical components of the motion vector, respectively.
number
[0077]
[96] Therefore, in this variant, in 1025, while the set of predictor candidates is not full, these valid CPMVPs are preserved in the set of CPMVP candidates for merge mode. Then, according to the second criterion, the valid candidate CPMVP is the bottom-left motion vector (obtained from position F or G).
number
number
number
number
[0078]
[97] In the case of bidirectional prediction, the cost is calculated for each CPMVP and each reference picture list L0 and L1 by using Equation 4. For each predictor bidirectional candidate, to compare the unidirectional predictor candidate with the bidirectional predictor candidate, the cost of the CPMVP is the average of the CPMVP cost associated with that list L0 and the CPMVP cost associated with that list L1. According to this variant, at 1025, an ordered set of valid CPMVPs is stored within the set of CPMVP candidates for merge mode for each reference picture list L0 and L1.
[0079] According to a second aspect of at least one embodiment of the determination of a set of predictor candidates for merge mode 1020 of encoding method 1000 shown in FIG. 11, instead of using triplets of adjacent blocks to determine three control motion point vectors for a predictor candidate, the determination uses pairs of adjacent vectors to determine two control motion point vectors, i.e., a pair of adjacent vectors for a predictor candidate.
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0080]
[99] In addition, by considering pairs in the upper left list {A,B,C} and the lower left list {F,G}, the two control point motion vectors of the current CU are
number
number
number
number
number
number
number
number
[0081]
[0100] As described by Equation 1 and shown in FIGS. As such, the standard affine motion model is the top-left and top-right control point motion vectors
number
number
number
number
number
number
number
number
[0082]
[0101] In addition, two CPMVs are determined based on pairs of adjacent blocks. Note that the second aspect is incompatible with evaluating 1024 predictor candidates based on the third CPMV (according to Equation 5 or 6) because it is not determined independently of the first two CPMVs. Therefore, the validity check of Equation 3 and the cost function of Equation 4 are skipped, and the CPMVP determined in 1023 is added to the set of predictor candidates for the affine merge mode without any sorting.
[0083]
[0102] In addition, two CPMVs are determined based on pairs of adjacent blocks. According to one variation of the second aspect, first, the use of bidirectional affine merge candidates is prioritized over unidirectional affine merge candidates by adding bidirectional candidates in the set of predictor candidates, thereby ensuring a maximum number of added bidirectional candidates in the set of predictor candidates for the affine merge mode.
[0084]
[0103] At least one implementation of determining a set of predictor candidates for merge mode 1020 In one version of the first and second embodiments, only affine neighboring blocks are used to generate new affine motion candidates. In other words, the determination of the top left CPMV, top right CPMV, and bottom left CPMV is based on the motion information of the respective top left, top right, and bottom left neighboring blocks. In this case, the motion information associated with at least one spatial neighboring block comprises only affine motion information. This variant differs from the JEM affine merge mode in that the motion model of the selected affine neighboring block is extended to the block to be encoded.
[0085]
[0104] At least one execution of the determination of the set of predictor candidates for merge mode 1020 In another variation of the first and second embodiments, at 1030, a top left CPMV
number
number
number
number
number
number
number
[0086]
[0105] This variant is different from the embodiment based on pairs of upper left and lower left neighboring blocks. In this case, the upper left and lower left CMPV
number
number
number
number
[0087]
[0106] In addition, this transformation also reduces the predictor candidate in affine merge mode. We provide additional CPMVP for complementary sets,
number
number
number
number
number
number
number
number
[0088] [Table 1]
[0089]
[0107] control_point_horizontal_l0_flag[ x0 ][ y0 ] is encoded The control points used for list L0 for the block, i.e., the current predicted block, The array index x0, y0 specifies the location (x0, y0) of the top left luma sample of the current predicted block relative to the top left luma sample of the picture. If the flag is equal to 1,
number
number
[0090]
[0108] At least one implementation of determining a set of predictor candidates for merge mode 1020 In another variation of the first and second embodiments, at 1030, a top left CPMV
number
number
number
number
number
number
[0091]
[0109] Figure 15 shows the encoding in the existing affine merge mode of JEM. 7 shows details of one embodiment of a process syntax 1500 used to predict the affine motion field of the current CU being decoded. The input 1501 to this process / syntax 1500 is the current coding unit for which it is desired to generate an affine motion field for the sub-blocks shown in FIG. 7. To select predictor candidates based on RD cost, the affine motion field of at least one embodiment of the process syntax 1500 (which is a bidirectional predictor) is used.
number
number
number
number
[0092]
[0110] In at least one implementation, a residual flag is used. 155 At 150, a flag is activated, indicating that coding is being performed with residual data. At 1560, the current CU is fully coded (with residual) and reconstructed, thereby giving a corresponding RD cost. The flag is then deactivated, indicating that coding is being performed without residual data, and the process returns to 1560, where the CU is coded (without residual) and thus given a corresponding RD cost. The lowest RD cost between the two previous ones indicates whether the residual should be coded (normal or skip). This best RD cost is then put into competition with other coding modes. Rate-distortion determination will be described in more detail below.
[0093]
[0111] FIG. 16 illustrates an exemplary decoder according to one general aspect of at least one embodiment. 16 illustrates a coding method 1600. At 1610, the method 1600 receives an index corresponding to a particular predictor candidate from a set of predictor candidates for a block to be decoded in a picture. In various embodiments, the particular predictor candidate is selected at an encoder from the set of predictor candidates, and the index allows one of the predictor candidates to be selected. At 1620, the method 1600 determines at least one spatial neighboring block for the block to be decoded. At 1630, the method 1600 determines a set of predictor candidates for the block to be decoded based on the at least one spatial neighboring block. The predictor candidates have one or more corresponding control point motion vectors and one reference picture for the block to be decoded. Any of the variations described for determining the set of predictor candidates in connection with the encoding method 1000 are mirrored at the decoder side, for example, as shown in FIG. 17. At 1640, method 1600 determines one or more control point motion vectors from a particular predictor candidate. At 1650, method 1600 determines a corresponding motion field for a particular predictor candidate based on the one or more corresponding control point motion vectors. In various embodiments, the motion field is based on a motion model, where the corresponding motion field identifies motion vectors used for prediction of sub-blocks of the block being decoded. The motion field is a top-left and top-right control point motion vector.
number
number
number
number
[0094]
[0112] FIG. 17 illustrates a decoding method 16 according to an aspect of at least one embodiment. 10 shows exemplary details of determining 1020 a set of predictor candidates for merge mode 00. The various embodiments described in relation to FIG. 11 for details of the encoding method are repeated here.
[0095]
[0113] The inventors have found that one aspect of the existing affine merging process described above is that The inventors have recognized that the current JEM's existing process systematically utilizes one and only one motion vector to propagate an affine motion field from past and neighboring CUs that are capable of affine motion toward the current CU. In various situations, the inventors have further recognized that this aspect may be disadvantageous, for example, because it does not select the optimal motion vector predictor. Furthermore, this predictor selection consists of only the first past and neighboring CU coded in affine mode in the ordered set (A, B, C, D, E), as already described above. In various situations, the inventors have further recognized that this limited selection may be disadvantageous, for example, because a better predictor may be available. Thus, the current JEM's existing process does not take into account the fact that some potential past and neighboring CUs around the current CU may also use affine motion, and that a different CU than the first one found to have used affine motion may be a better predictor for the current CU's motion information.
[0096]
[0114] Therefore, the inventors have proposed several methods to improve the performance of the existing JEM codec. We have recognized potential benefits for improving CU affine motion vector prediction that are currently not being exploited.
[0097]
[0115] FIG. 18 illustrates an example system 1 in which various aspects of the example embodiments may be implemented. 18 illustrates a block diagram of a video system 1800. System 1800 may be implemented as a device including various components described below and configured to perform the processes described above. Examples of such devices include, without limitation, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected home appliances, and servers. System 1800 may be communicatively coupled to other similar systems and to displays via communication channels, as shown in FIG. 18, and may be communicatively coupled to implement all or a portion of the exemplary video system described above, as known to those skilled in the art.
[0098]
[0116] Various embodiments of the system 1800 may implement various processes, as described above. The system 1800 includes at least one processor 1810 configured to execute instructions loaded therein to perform the functions described above. The processor 1810 may include built-in memory, input / output interfaces, and various other circuits known in the art. The system 1800 may also include at least one memory 1820 (e.g., a volatile memory device, a non-volatile memory device). The system 1800 may additionally include a storage device 1840, which may include non-volatile memory including, without limitation, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 1840 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device. The system 1800 may also include an encoder / decoder module 1830 configured to process data to provide encoded and / or decoded video, and the encoder / decoder module 1830 may include its own processor and memory.
[0099]
[0117] The encoder / decoder module 1830 performs encoding and / or decoding. 18. The encoder / decoder module 1830 represents one or more modules that may be included within a device to perform coding functions. As is known, such a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 1830 may be implemented as a separate element of the system 1800, or may be embedded within one or more processors 1810 as a combination of hardware and software, as is known to those skilled in the art.
[0100]
[0118] on one or more processors 1810 to execute the various processes described above. The loaded program code may be stored in storage device 1840 and subsequently loaded onto memory 1820 for execution by processor(s) 1810. According to an example embodiment, one or more of processor(s) 1810, memory 1820, storage device 1840, and encoder / decoder module(s) 1830 may store one or more of a variety of items including, without limitation, input video, decoded video, bitstream, equations, formulas, matrices, variables, operations, and operation logic during execution of the above-described processes.
[0101]
[0119] The system 1800 also communicates with other devices via a communication channel 1860. The system 1800 may also include a communications interface 1850 that enables communication between the system 1800 and the communication channel 1860. The communications interface 1850 may include, without limitation, a transceiver configured to transmit and receive data to and from a communications channel 1860. The communications interface 1850 may include, without limitation, a modem or a network card, and the communications channel 1850 may be implemented within a wired and / or wireless medium. The various components of the system 1800 may be connected or communicatively coupled using various suitable connections, including, without limitation, an internal bus, a wire, and a printed circuit board (not shown in FIG. 18 ).
[0102]
[0120] The exemplary embodiment is implemented by the processor 1810 or hardware. The memory 1820 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1810 may be of any type suitable for the technical environment and may be implemented using, by way of non-limiting examples, a microphone. The processor may include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0103]
[0121] The implementations described herein may be, for example, methods or processes, devices, or the like. The invention may be implemented in a device, a software program, a data stream, or a signal. Also, even if described in the context of only a single implementation (e.g., described only as a method), the implementation of the described features may be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. A method may be implemented in an apparatus, such as a processor, which refers to, for example, a general processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor may also be implemented in, for example, a computer, a cell phone, a portable / personal digital assistant (PDA), or a mobile device. "), and other devices that facilitate communication of information between end users.
[0104]
[0122] Furthermore, those skilled in the art will appreciate that the exemplary HEVC encoder 100 and It may be readily understood that the example HEVC decoder shown in Figure 1 and Figure 3 may be modified in accordance with the above teachings of this disclosure to implement the disclosed improvements over the existing HEVC standard to achieve better compression / decompression. For example, motion compensation 170 and motion estimation 175 in the example encoder 100 of Figure 1 and motion compensation 375 in the example decoder of Figure 3 may be modified in accordance with the disclosed teachings to implement one or more example aspects of the present disclosure, including providing improved affine merge prediction to the existing JEM.
[0105]
[0123] "One embodiment" or "an embodiment" or "one implementation" or "an implementation" only Rather, reference to these other variations means that the particular features, structures, characteristics, etc. described in connection with the embodiments are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" appearing in various places throughout this specification may be used interchangeably. )" or "in one implementation" or "in one implementation" The appearances of the phrase "in an implementation," as well as any other variations, do not necessarily all refer to the same embodiment.
[0106]
[0124] Additionally, the application or its claims may refer to "determining" various pieces of information. Determining the information may include, for example, one or more of estimating the information, calculating the information, predicting the information, or retrieving the information from a memory.
[0107]
[0125] Furthermore, this application and its claims refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or inferring information.
[0108]
[0126] Additionally, the application or its claims may refer to "receiving" various pieces of information. "Receiving," like "accessing," is intended to be a broad term. Receiving information can refer to, for example, accessing information. "Receiving" may include one or more of receiving or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves, in one manner or another, an action such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0109]
[0127] As will be apparent to those skilled in the art, the implementations may be stored or transmitted, for example. Various signals formatted to carry information can be generated. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a high-frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
1. 1. A method for video encoding, comprising: determining, for a block to be encoded in affine merge mode, a top-left list of spatial neighboring blocks comprising the neighboring blocks of the block's top left corner, a top-right list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's top right corner, and a bottom-left list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's bottom left corner; determining, for the block to be encoded, a set of predictor candidates for an affine merge mode based on the top-left list, the top-right list, and the bottom-left list; the set of predictor candidates for the affine merge mode includes a first predictor candidate if a reference picture of a first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the top-right list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the first predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of the first predictor candidate; The reference picture of the first spatially neighboring block in the top-left list is used as a reference picture of the first predictor candidate; and the set of predictor candidates for the affine merge mode includes a second predictor candidate if the reference picture of the first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as the first control point motion vector of the second predictor candidate; the motion vector of the second spatially neighboring block in the bottom-left list and the motion vector of the first spatially neighboring block in the top-left list are used to derive a second control point motion vector of the second predictor candidate; determining a set of predictor candidates for an affine merge mode, in which the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the second predictor candidate; determining a motion field for the block to be encoded and for each predictor candidate based on a motion model and based on one or more control point motion vectors of the predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be encoded; selecting a predictor candidate from the set of predictor candidates based on a rate-distortion measurement during prediction in response to the motion field determined for each predictor candidate; encoding the block based on the motion field for the selected predictor candidate; A method having the following.
2. The method of claim 1 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises translational motion information.
3. The method of claim 1 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises affine motion information.
4. the set of predictor candidates for the affine merge mode includes a third predictor candidate if the reference picture of a first spatial neighboring block in the top-left list is the same as the reference picture of a second spatial neighboring block in the top-right list and is the same as the reference picture of a third spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the third predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of the third predictor candidate; the motion vector of the third spatially neighboring block in the bottom-left list is used as a third control point motion vector of the third predictor candidate; The method of claim 1 , further characterized in that the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the third predictor candidate.
5. 1. A method for video decoding, comprising: determining, for a block decoded in affine merge mode, a top-left list of spatial neighboring blocks comprising the neighboring blocks of the block's top left corner, a top-right list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's top right corner, and a bottom-left list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's bottom left corner; determining, for the block to be decoded, a set of predictor candidates for an affine merge mode based on the top-left list, the top-right list, and the bottom-left list; the set of predictor candidates for the affine merge mode includes a first predictor candidate if a reference picture of a first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the top-right list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the first predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of the first predictor candidate; The reference picture of the first spatially neighboring block in the top-left list is used as a reference picture of the first predictor candidate; and the set of predictor candidates for the affine merge mode includes a second predictor candidate if the reference picture of the first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as the first control point motion vector of the second predictor candidate; the motion vector of the second spatially neighboring block in the bottom-left list and the motion vector of the first spatially neighboring block in the top-left list are used to derive the second control point motion vector of the second predictor candidate; determining a set of predictor candidates for an affine merge mode, in which the reference picture of the first spatial neighboring block in the top-left list is used as a reference picture for a constructed affine merge predictor candidate; determining one or more control point motion vectors from a particular predictor candidate for the block being decoded; determining a motion field for the decoded block based on a motion model and based on the one or more control point motion vectors for the decoded block, the motion field identifying motion vectors used for prediction of sub-blocks of the decoded block; decoding the block based on the determined motion field; A method having the following.
6. The method of claim 5 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises translational motion information.
7. The method of claim 5 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises affine motion information.
8. the set of predictor candidates for the affine merge mode includes a third predictor candidate if the reference picture of a first spatial neighboring block in the top-left list is the same as the reference picture of a second spatial neighboring block in the top-right list and is the same as the reference picture of a third spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the third predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of a third predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of a third predictor candidate; The method of claim 5 , further characterized in that the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the third predictor candidate.
9. The motion model is an affine model, and for each position (x,y) inside the block being decoded the motion field is determined by: [Equation 1] Here, (v 0x , v 0y ) and (v 1x , v 1y ) are the control point motion vectors used to generate the motion field, and (v 0x , v 0y ) corresponds to the first control point motion vector, and (v 1x , v 1y 3. The method of claim 2, further characterized in that: ∑ a ∑ b ∑ i ...
10. 6. The method of claim 5, further characterized by decoding an index for the particular predictor candidate from the set of predictor candidates.
11. 1. An apparatus for video encoding, comprising at least a memory and one or more processors, the one or more processors comprising: determining, for a block to be encoded in affine merge mode, a top-left list of spatial neighboring blocks comprising the neighboring blocks of the block's top left corner, a top-right list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's top right corner, and a bottom-left list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's bottom left corner; determining, for the block to be encoded, a set of predictor candidates for an affine merge mode based on the top-left list, the top-right list, and the bottom-left list; the set of predictor candidates for the affine merge mode includes a first predictor candidate if a reference picture of a first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the top-right list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the first predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of the first predictor candidate; The reference picture of the first spatially neighboring block in the top-left list is used as a reference picture of the first predictor candidate; and the set of predictor candidates for the affine merge mode includes a second predictor candidate if the reference picture of the first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as the first control point motion vector of the second predictor candidate; the motion vector of the second spatially neighboring block in the bottom-left list and the motion vector of the first spatially neighboring block in the top-left list are used to derive the second control point motion vector of the second predictor candidate; determining a set of predictor candidates for an affine merge mode, in which the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the second predictor candidate; determining a motion field for the block to be encoded and for each predictor candidate based on a motion model and based on one or more control point motion vectors of the predictor candidate, the motion field identifying motion vectors used for prediction of sub-blocks of the block to be encoded; selecting a predictor candidate from the set of predictor candidates based on a rate-distortion measurement during prediction in response to the motion field determined for each predictor candidate; encoding the block based on the motion field for the selected predictor candidate.
12. The apparatus of claim 11 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises translational motion information.
13. The apparatus of claim 11 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises affine motion information.
14. the set of predictor candidates for the affine merge mode includes a third predictor candidate if the reference picture of a first spatial neighboring block in the top-left list is the same as the reference picture of a second spatial neighboring block in the top-right list and is the same as the reference picture of a third spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the third predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of the third predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a third control point motion vector of the third predictor candidate; The apparatus of claim 11 , further characterized in that the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the third predictor candidate.
15. 1. An apparatus for video decoding, comprising at least a memory and one or more processors, the one or more processors comprising: determining, for a block decoded in affine merge mode, a top-left list of spatial neighboring blocks comprising the neighboring blocks of the block's top left corner, a top-right list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's top right corner, and a bottom-left list of spatial neighboring blocks of the block comprising the neighboring blocks of the block's bottom left corner; determining, for the block to be decoded, a set of predictor candidates for an affine merge mode based on the top-left list, the top-right list, and the bottom-left list; the set of predictor candidates for the affine merge mode includes a first predictor candidate if a reference picture of a first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the top-right list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the first predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of the first predictor candidate; The reference picture of the first spatially neighboring block in the top-left list is used as a reference picture of the first predictor candidate; and the set of predictor candidates for the affine merge mode includes a second predictor candidate if the reference picture of the first spatial neighboring block in the top-left list is the same as a reference picture of a second spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the second predictor candidate; the motion vector of the second spatially neighboring block in the bottom-left list and the motion vector of the first spatially neighboring block in the top-left list are used to derive a second control point motion vector of the second predictor candidate; determining a set of predictor candidates for an affine merge mode, in which the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the second predictor candidate; determining one or more control point motion vectors from a particular predictor candidate for the block being decoded; determining a motion field for the decoded block based on a motion model and based on the one or more control point motion vectors for the decoded block, the motion field identifying motion vectors used for prediction of sub-blocks of the decoded block; and decoding the block based on the determined motion field.
16. The apparatus of claim 15 , further characterized in that the motion information associated with at least one of the spatially neighboring blocks comprises translational motion information.
17. The device of claim 15, further characterized in that the motion information associated with at least one of the spatially adjacent blocks comprises affine motion information.
18. the set of predictor candidates for the affine merge mode includes a third predictor candidate if the reference picture of a first spatial neighboring block in the top-left list is the same as the reference picture of a second spatial neighboring block in the top-right list and is the same as the reference picture of a third spatial neighboring block in the bottom-left list, wherein: the motion vector of the first spatially neighboring block in the top-left list is used as a first control point motion vector of the third predictor candidate; the motion vector of the second spatially neighboring block in the top right list is used as a second control point motion vector of a third predictor candidate; the motion vector of the third spatially neighboring block in the bottom-left list is used as a second control point motion vector of the third predictor candidate; The apparatus of claim 15 , further characterized in that the reference picture of the first spatially neighboring block in the top-left list is used as a reference picture for the third predictor candidate.
19. The motion model is an affine model, and for each position (x, y) inside the block to be decoded the motion field is determined by: [Equation 2] Here, (v 0x , v 0y ) and (v 1x , v 1y ) are the control point motion vectors used to generate the motion field, and (v 0x , v 0y ) corresponds to the first control point motion vector, and (v 1x , v 1y 16. The apparatus of claim 15, further characterized in that: ∇ ...
Citation Information
Patent Citations
Image prediction method and related device
WO2016141609A1
Method and apparatus for affine merge mode prediction for video coding system
WO2017118409A1
Image encoding / decoding method and device
WO2017133243A1
Method and apparatus of video coding with affine motion compensation
WO2017148345A1
Affine motion prediction for video coding
WO2017200771A1