Multiple predictor candidates for motion compensation
By constructing and evaluating multiple predictor candidates for affine merge mode in video codecs, the method optimizes coding efficiency and enhances video compression performance.
Patent Information
- Application Number
- JP2024095548
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-03-30
- Filing Date
- 2024-06-13
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2038-06-25
AI Technical Summary
Existing video encoding and decoding technologies face challenges in efficiently utilizing affine motion models for motion compensation, particularly in improving the performance of affine merge modes in video codecs like JEM, as they often rely on limited predictor candidates that do not optimize coding efficiency.
An improved method and apparatus for video encoding and decoding that involves constructing a set of multiple predictor candidates for affine merge mode, evaluating them based on criteria such as rate-distortion, and selecting the best candidate to enhance coding efficiency by using an affine motion model with extended predictor candidates.
This approach enhances the coding efficiency of affine merge modes by selecting the most efficient predictor candidates, leading to improved video compression performance and quality.
Smart Images

Figure 0007804005000034 
Figure 0007804005000035 
Figure 0007804005000036
Abstract
Description
[Technical Field]
[0001] [1] At least one embodiment of the present invention relates generally to a method or apparatus, e.g., for video encoding or decoding, and more particularly to a method or apparatus for selecting a predictor candidate from a set of multiple predictor candidates for motion compensation based on a motion model, e.g., an affine model, for a video encoder or video decoder. [Background technology]
[0002] [2] To achieve high compression efficiency, image and video coding schemes typically use prediction, including motion vector prediction, and transforms to exploit spatial and temporal redundancy in the video content. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and the difference between the original and predicted image, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction.
[0003] [3] A recent addition to high-compression techniques is the use of motion models based on affine modeling. In particular, affine modeling is used in motion compensation for encoding and decoding video pictures. In general, affine modeling is a model that uses at least two parameters, e.g., two control point motion vectors (CPMVs) representing the motion at each corner of a block of a picture, that allow deriving a motion field for an entire block of a picture, e.g., to simulate rotation and approximation (zoom). Summary of the Invention [Means for solving the problem]
[0004] According to a general aspect of at least one embodiment, a method for video coding is presented, comprising: determining, for a block to be coded in a picture, a set of predictor candidates having a plurality of predictor candidates; selecting a predictor candidate from the set of predictor candidates; determining, for the predictor candidate selected from the set of predictor candidates, one or more corresponding control point motion vectors for the block; determining, for the selected predictor candidate, a corresponding motion field based on a motion model for the selected predictor candidate based on the one or more corresponding control point motion vectors, wherein the corresponding motion field identifies a motion vector used for predicting a sub-block of the block to be coded; encoding the block based on the corresponding motion field for the predictor candidate selected from the set of predictor candidates; and encoding an index for the predictor candidate selected from the set of predictor candidates.
[0005] According to another general aspect of at least one embodiment, a method for video decoding is presented, comprising: receiving, for a block to be decoded in a picture, an index corresponding to a particular predictor candidate; determining, for the particular predictor candidate, one or more corresponding control point motion vectors for the block to be decoded; determining, for the particular predictor candidate, a corresponding motion field based on a motion model that identifies motion vectors used for prediction of sub-blocks of the block to be decoded based on the one or more corresponding control point motion vectors; and decoding the block based on the corresponding motion field.
[0006] According to another general aspect of at least one embodiment, an apparatus for video encoding is presented, comprising: means for determining, for a block to be encoded in a picture, a set of predictor candidates having a plurality of predictor candidates; means for selecting a predictor candidate from the set of predictor candidates; means for determining, for the selected predictor candidate, based on one or more corresponding control point motion vectors, a corresponding motion field based on a motion model for the selected predictor candidate, the corresponding motion field identifying a motion vector used for predicting a sub-block of the block to be encoded; means for encoding the block based on the corresponding motion field for the predictor candidate selected from the set of predictor candidates; and means for encoding an index for the predictor candidate selected from the set of predictor candidates.
[0007] According to another general aspect of at least one embodiment, an apparatus for video decoding is presented, comprising: means for receiving, for a block to be decoded in a picture, an index corresponding to a particular predictor candidate; means for determining, for the particular predictor candidate, one or more corresponding control point motion vectors for the block to be decoded; means for determining, for the particular predictor candidate, a corresponding motion field based on a motion model, the corresponding motion field identifying motion vectors used for prediction of sub-blocks of the block to be decoded based on the one or more corresponding control point motion vectors; and means for decoding the block based on the corresponding motion field.
[0008] According to another general aspect of at least one embodiment, an apparatus for video encoding is provided, comprising one or more processors and at least one memory. The one or more processors are configured to: determine, for a block to be encoded in a picture, a set of predictor candidates having a plurality of predictor candidates; select a predictor candidate from the set of predictor candidates; determine, for the predictor candidate selected from the set of predictor candidates, one or more corresponding control point motion vectors for the block; determine, for the selected predictor candidate, a corresponding motion field based on a motion model for the selected predictor candidate, the corresponding motion field identifying a motion vector used for predicting a sub-block of the block to be encoded, based on the one or more corresponding control point motion vectors; encode the block based on the corresponding motion field for the predictor candidate selected from the set of predictor candidates; and encode an index for the predictor candidate selected from the set of predictor candidates. The at least one memory is configured to at least temporarily store the encoded block and / or the encoded index.
[0009] According to another general aspect of at least one embodiment, an apparatus for video decoding is provided, comprising one or more processors and at least one memory. The one or more processors are configured to receive, for a block to be decoded in a picture, an index corresponding to a particular predictor candidate; determine, for the particular predictor candidate, one or more corresponding control point motion vectors for the block to be decoded; determine, for the particular predictor candidate, a corresponding motion field based on a motion model that identifies motion vectors used for prediction of sub-blocks of the block to be decoded based on the one or more corresponding control point motion vectors; and decode the block based on the corresponding motion field. The at least one memory is configured to at least temporarily store the decoded block.
[0010] According to another general aspect of at least one embodiment, a method for video encoding is presented, comprising: determining, for a block to be coded in a picture, a set of predictor candidates; for each of a plurality of predictor candidates in the set of predictor candidates, determining one or more corresponding control point motion vectors for the block; for each of the plurality of predictor candidates, determining a corresponding motion field based on a motion model for each of the plurality of predictor candidates in the set of predictor candidates based on the one or more corresponding control point motion vectors; evaluating the plurality of predictor candidates according to one or more criteria and based on the corresponding motion field; selecting a predictor candidate from the plurality of predictor candidates based on the evaluation; and encoding the block based on the predictor candidate selected from the set of predictor candidates.
[0011] According to another general aspect of at least one embodiment, a method for video decoding is presented, comprising obtaining an index corresponding to a selected predictor candidate for a block to be decoded in a picture. The selected predictor candidate is selected by: determining, at an encoder, a set of predictor candidates for a block to be coded in the picture; determining, for each of a plurality of predictor candidates in the set of predictor candidates, one or more corresponding control point motion vectors for the block to be coded; determining, for each of the plurality of predictor candidates, a corresponding motion field based on a motion model for each of the plurality of predictor candidates in the set of predictor candidates based on the one or more corresponding control point motion vectors; evaluating the plurality of predictor candidates according to one or more criteria and based on the corresponding motion field; selecting a predictor candidate from the plurality of predictor candidates based on the evaluation; and encoding an index for the selected predictor candidate from the set of predictor candidates. The method further comprises decoding the block based on the index corresponding to the selected predictor candidate.
[0012]
[12] According to another general aspect of at least one embodiment, the method may further comprise evaluating the plurality of predictor candidates according to one or more criteria and based on a corresponding motion field for each of the plurality of predictor candidates; and selecting a predictor candidate from the plurality of predictor candidates based on the evaluation.
[0013] According to another general aspect of at least one embodiment, the apparatus may further include means for evaluating the plurality of predictor candidates according to one or more criteria and based on a corresponding motion field for each of the plurality of predictor candidates, and means for selecting a predictor candidate from the plurality of predictor candidates based on the evaluation.
[0014]
[14] According to another general aspect of at least one embodiment, the one or more criteria are based on a rate-distortion decision corresponding to one or more of a plurality of predictor candidates in a set of predictor candidates.
[0015]
[15] According to another general aspect of at least one embodiment, decoding or encoding a block based on a corresponding motion field comprises decoding or encoding, respectively, a predictor indicated by a motion vector based on a predictor for the sub-block.
[0016]
[16] According to another general aspect of at least one embodiment, the set of predictor candidates comprises spatial and / or temporal candidates for the block to be encoded or decoded.
[0017]
[17] According to another general aspect of at least one embodiment, the motion model is an affine model.
[0018] According to another general aspect of at least one embodiment, the corresponding motion field for each position (x,y) in the block being coded or decoded is:
number
[0019]
[19] According to another general aspect of at least one embodiment, the number of spatial candidates is five or greater.
[0020]
[20] According to another general aspect of at least one embodiment, one or more additional control point motion vectors are added to determine a corresponding motion field based on a function of the determined one or more corresponding control point motion vectors.
[0021]
[21] According to another general aspect of at least one embodiment, the function includes one or more of: 1) an average, 2) a weighted average, 3) a unique average, 4) an average, 5) a median, or 6) a unidirectional portion of one of 1) to 6) above, of the determined one or more corresponding control point motion vectors.
[0022]
[22] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is presented that includes data content generated according to any of the methods or apparatus described above.
[0023]
[23] According to another general aspect of at least one embodiment, there is provided a signal comprising video data generated in accordance with any of the methods or apparatus described above.
[0024]
[24] One or more embodiments of the present disclosure also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the above-described methods. Embodiments of the present disclosure also provide a computer-readable storage medium having stored thereon a bitstream generated according to the above-described methods. Embodiments of the present disclosure also provide methods and apparatuses for transmitting a bitstream generated according to the above-described methods. Embodiments of the present disclosure also provide a computer program product including instructions for performing any of the above-described methods. [Brief explanation of the drawings]
[0025] [Figure 1]
[25] shows a block diagram of an embodiment of an HEVC (High Efficiency Video Coding) video encoder. [Figure 2A]
[26] is an example image illustrating HEVC reference sample generation. [Figure 2B]
[27] is an example image showing intra prediction direction in HEVC. [Figure 3]
[28] shows a block diagram of an embodiment of an HEVC video decoder. [Figure 4]
[29] presents examples of the coding tree unit (CTU) and coding tree (CT) concepts for representing compressed HEVC pictures. [Figure 5]
[30] gives an example of splitting a coding tree unit (CTU) into coding units (CUs), prediction units (PUs), and transform units (TUs). [Figure 6]
[31] gives an example of an affine model as a motion model used in the Joint Exploration Model (JEM). [Figure 7]
[32] shows an example of a 4x4 sub-CU-based affine motion vector field used in the Joint Evolution Model (JEM). [Figure 8A]
[33] shows examples of motion vector prediction candidates for affine inter-CUs. [Figure 8B]
[34] shows examples of motion vector prediction candidates in affine merging mode. [Figure 9]
[35] present an example of the spatial derivation of affine control point motion vectors in the case of an affine merged mode motion model. [Figure 10]
[36] An example method according to a general aspect of at least one embodiment is provided. [Figure 11]
[37] Another example method according to a general aspect of at least one embodiment is shown. [Figure 12]
[38] Another example method according to a general aspect of at least one embodiment is shown. [Figure 13]
[39] Another example method according to a general aspect of at least one embodiment is shown. [Figure 14]
[40] presents an example of a known process for evaluating inter-CU affine merging modes in JEM. [Figure 15]
[41] present an example of the process for selecting predictor candidates in the affine merging mode in JEM. [Figure 16]
[42] shows an example of an affine motion field propagated by an affine merging prediction candidate located to the left of the current block being coded or decoded. [Figure 17]
[43] shows an example of an affine motion field propagated by affine merger predictor candidates located above and to the right of the current block being coded or decoded. [Figure 18]
[44] illustrates an example of a predictor candidate selection process according to a general aspect of at least one embodiment. [Figure 19]
[45] illustrates an example process for constructing a set of multiple predictor candidates according to a general aspect of at least one embodiment. [Figure 20]
[46] illustrates an example of the process for deriving the top-left and top-right corner CPMVs for each predictor candidate, according to a general aspect of at least one embodiment. [Figure 21]
[47] illustrates an example of an expanded set of spatial predictor candidates according to a general aspect of at least one embodiment. [Figure 22]
[48] illustrates another example of a process for constructing a set of multiple predictor candidates according to a general aspect of at least one embodiment. [Figure 23]
[49] illustrates another example of a process for constructing a set of multiple predictor candidates according to a general aspect of at least one embodiment. [Figure 24]
[50] An example of how temporary candidates may be used for predictor candidates according to a general aspect of at least one embodiment is shown. [Figure 25]
[51] Figure 5 illustrates an example process for adding an average CPMV motion vector calculated from stored CPMV candidates to a final CPMV candidate set, according to a general aspect of at least one embodiment. [Figure 26]
[52] Figure 5 shows a block diagram of an example device in which various aspects of the embodiments may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0026]
[53] Figure 1 shows a typical High Efficiency Video Coding (HEVC) encoder 100. HEVC is a compression standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) (see, for example, "ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services—Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265").
[0027]
[54] In HEVC, to encode a video sequence having one or more pictures, the pictures are partitioned into one or more slices, each of which may contain one or more slice segments. The slice segments are organized into coding units, prediction units, and transform units.
[0028]
[55] In this application, the terms "reconstructed" and "decoded" are used interchangeably, the terms "encoded" and "coded" are used interchangeably, and the terms "picture" and "frame" may be used interchangeably. Often, but not always, the term "reconstructed" is used on the encoder side and "decoded" is used on the decoder side.
[0029]
[56] The HEVC specification distinguishes between "blocks" and "units," where a "block" refers to a specific area within a sample array (e.g., luma, Y), and a "unit" includes all coded juxtaposed blocks of color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data associated with the block (e.g., motion vectors).
[0030]
[57] For encoding, a picture is partitioned into square coding tree blocks (CTBs) with configurable sizes, and contiguous sets of coding tree blocks are grouped into slices. A coding tree unit (CTU) contains the CTB of a coded color component. The CTB is the root of a quadtree that divides into coding blocks (CBs), which may be partitioned into one or more prediction blocks (PBs), forming the root of a quadtree that divides into transform blocks (TBs). Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) contains a prediction unit (PU) and a set of tree-structured transform units (TUs), where a PU contains prediction information for all color components and a TU contains residual coding syntax structures for each color component. The sizes of the CBs, PBs, and TBs for the luma component are comparable to the corresponding CUs, PUs, and TUs. In this application, the term "block" may be used to refer to, for example, any of the CTUs, CUs, PUs, TUs, CBs, PBs, and TBs. Additionally, "block" may also be used to refer to macroblocks and partitions specified in H.264 / AVC or other video coding standards, and more generally to data arrays of various sizes.
[0031]
[58] In a typical encoder 100, a picture is coded by the following encoder elements: A picture to be coded is processed in units of CUs. Each CU is coded using either intra mode or inter mode. If a CU is coded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to code the CU and indicates the intra / inter decision with a prediction mode flag. A prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0032]
[59] In intra modes, CUs are predicted from neighboring reconstructed samples within the same slice. A set of 35 intra prediction modes is available in HEVC, including DC prediction mode, planar prediction mode, and 33 angular prediction modes. The intra prediction reference is reconstructed from rows and columns adjacent to the current block. The reference spans twice the block size horizontally and vertically using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra prediction, the reference sample may be copied along the direction indicated by the angular prediction mode.
[0033]
[60] The available luma intra-prediction modes for the current block can be coded using two different options: If the applicable mode is included in the three most probable modes (MPM), the mode is signaled by its index in the MPM list; otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra-prediction modes of the neighboring blocks above and to the left.
[0034]
[61] For an inter CU, the corresponding coding block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how the inter prediction was performed. Motion information (i.e., motion vectors and reference picture indices) can be signaled in two ways: "merged mode" and "advanced motion vector prediction (AMVP)."
[0035]
[62] In merged mode, a video encoder or decoder compiles a candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates in the candidate list. At the decoder side, motion vectors (MVs) and reference picture indices are reconstructed based on the signaled candidates.
[0036]
[63] The set of possible candidates in merge mode consists of spatially adjacent candidates, temporal candidates, and generated candidates. Figure 2A shows the locations of five spatial candidates {a1, b1, b0, a0, b2} relative to the current block 210, where a0 and a1 are to the left of the current block and b1, b0, b2 are above the current block. For each candidate location, availability is checked in the order a1, b1, b0, a0, b2, and then redundancies within the candidates are removed.
[0037]
[64] The motion vectors of the collocated positions in the reference pictures can be used to derive the temporal candidates. The available reference pictures are selected on a slice basis and indicated in the slice header, and the reference index for the temporal candidate is i ref = 0. If the POC distance (td) between the picture of the co-located PU and the reference picture from which the co-located PU is predicted is the same as the distance (tb) between the current picture and the reference picture containing the co-located PU, then the co-located motion vector mv col can be used directly as a temporal candidate. Otherwise, the scaled motion vector tb / td * MV col is used as the time candidate. Depending on where the current PU is located, the collocated PU is determined by the sample location at the bottom right or center of the current PU.
[0038]
[65] The maximum number N of merge candidates is specified in the slice header. If the number of merge candidates is greater than N, only the first N-1 spatial and temporal candidates are used. Otherwise, if the number of merge candidates is less than N, the set of candidates is filled up to the maximum number N with candidates generated as a combination of already existing candidates or null candidates. Candidates used in merge mode may be referred to as "merged candidates" in this application.
[0039]
[66] If a CU indicates skip mode, the available indices for merging candidates are indicated only if the list of merging candidates is greater than 1, and no further information is coded for the CU. In skip mode, motion vectors are applied without residual updates.
[0040]
[67] In AMVP, a video encoder or decoder assembles a candidate list based on motion vectors determined from already coded blocks. The video encoder then signals an index in the candidate list to identify a motion vector predictor (MVP) and signals a motion vector differential (MVD). At the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. The available reference picture indexes are also explicitly coded in the PU syntax for AMVP.
[0041]
[68] In AMVP, only two spatial motion candidates are selected. The first spatial motion candidate is selected from the left position {a0, a1}, and the second candidate is selected from the top position {b0, b1, b2}, while maintaining the search order shown in the two sets. If the number of motion vector candidates is not equal to two, a temporal MV candidate may be included. If the set of candidates is still not completely filled, a zero motion vector is used.
[0042]
[69] If the reference picture index of the spatial candidate corresponds to the reference picture index for the current PU (i.e., use the same reference picture index or both use long-term reference pictures regardless of the reference picture list), the spatial candidate motion vector is used directly. Otherwise, if the reference picture is short-term, the candidate motion vector is scaled according to the distance between the current picture and the reference picture of the current PU (tb) and the distance between the current picture and the reference picture of the spatial candidate (td). Candidates used in AMVP mode may be referred to as "AMVP candidates" in this application.
[0043]
[70] For ease of description, blocks tested in "merged" mode at the encoder side or decoded in "merged" mode at the decoder side are referred to as "merged" blocks, and blocks tested in AMVP mode at the encoder side or decoded in AMVP mode at the decoder side are referred to as "AMVP" blocks.
[0044]
[71] Figure 2B shows a typical motion vector representation using AMVP. For the current block to be coded 240, motion estimation determines a motion vector (MV current ) from the left block 230. left ) and the motion vector (MV above ) and MV left and MV above The motion vector predictor is MVP. current Then, MVD current =MV current -MVP current The motion vector difference can be calculated as:
[0045]
[72] Motion-compensated prediction can be performed using one or two reference pictures for prediction. In P slices, only a single prediction reference may be used for inter prediction, allowing uni-prediction for the predicted block. In B slices, two reference picture lists are available, and uni-prediction or bi-prediction may be used. In bi-prediction, one reference picture from each of the reference picture lists is used.
[0046]
[73] In HEVC, the precision of the motion information for motion compensation is one-quarter sample (also called quarter-pel or 1 / 4-pel) for the luma component and one-eighth sample (also called 1 / 8-pel) for the chroma component in the case of a 4:2:0 configuration. 7-tap or 8-tap interpolation filters are used for the interpolation of fractional sample positions, i.e., 1 / 4, 1 / 2, and 3 / 4 of a full sample position in both the horizontal and vertical directions can be processed for luma.
[0047]
[74] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder may also omit both the transform and quantization, i.e., the residual is directly coded without applying a transform or quantization process. In direct PCM coding, no prediction is applied, and coding unit samples are coded directly into a bitstream.
[0048]
[75] The encoder decodes the coded blocks to provide references for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual is combined (155) with the predicted block to reconstruct an image block. An in-loop filter (165) is applied to the reconstructed picture, for example, to perform deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).
[0049]
[76] Figure 3 shows a block diagram of a typical HEVC video decoder 300. In the typical decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs video decoding as part of the encoding of the video data, performing a decoding pass reciprocal to the encoding pass shown in Figure 1.
[0050]
[77] In particular, the decoder input includes a bitstream that may be generated by video encoder 100. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coding information. The transform coefficients are inversely quantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined (355) with a predicted block to reconstruct an image block. The predicted block may result (370) from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). As described above, AMVP and merged mode techniques may be used to derive motion vectors for motion compensation, which may use interpolation filters to calculate interpolated values for sub-integer samples of the reference block. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0051]
[78] As mentioned above, in HEVC, motion-compensated temporal prediction is used to exploit redundancy that exists between consecutive pictures in a video. To that end, a motion vector is associated with each prediction unit (PU). As mentioned above, each CTU is represented in the compressed domain by a coding tree. This is a quadtree division of the CTU, where each leaf is called a coding unit (CU), and is also shown in FIG. 4 for CTUs 410 and 420. Each CU is then given some intra- or inter-prediction parameters as prediction information. To that end, a CU may be spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. Intra- or inter-coding modes are assigned at the CU level. These concepts are further illustrated in FIG. 5 for exemplary CTU 500 and CU 510.
[0052]
[79] In HEVC, each PU is assigned one motion vector. This motion vector is used for motion-compensated temporal prediction of the considered PU. Therefore, in HEVC, the motion model linking a prediction block and its reference block simply consists of a transformation or calculation based on the reference block and the corresponding motion vector.
[0053]
[80] To improve HEVC, reference software and / or a documented JEM (Joint Exploration Model) is being developed by the Joint Video Exploration Team (JVET). In one version of the JEM (e.g., “Algorithm Description of Joint Exploration Test Model 5,” document JVET-E1001_v2, Joint Video Exploration Team of ISO / IEC JTC1 / SC29 / WG11, 5th Meeting, January 12-20, 2017, Geneva, Switzerland), several additional motion models are supported to improve temporal prediction. To do so, a PU may be spatially partitioned into sub-PUs, and the model may be used to assign each sub-PU a dedicated motion vector.
[0054]
[81] In more recent versions of JEM (e.g., “Algorithm Description of Joint Exploration Test Model 2”, document JVET-B1001_v3, Joint Video Exploration Team of ISO / IEC JTC1 / SC29 / WG11, 2nd Meeting, February 20-26, 2016, San Diego, USA), it is not specified that CUs are divided into PUs or TUs. Instead, more flexible CU sizes may be used, and some motion data may be directly assigned to each CU. In this new codec design in newer JEM versions, CUs may be divided into sub-CUs, and motion vectors may be calculated for each sub-CU of a divided CU.
[0055]
[82] One of the new motion models introduced in JEM is the use of an affine-type model for motion vectors to represent the motion vectors in a CU. The motion model used is illustrated in Figure 6 and expressed by Equation 1 as shown below. The affine-type motion field comprises the following motion vector component values for each location (x,y) within the considered block 600 in Figure 6:
number
[0056]
[83] To reduce complexity, a motion vector is calculated for each 4x4 sub-block (sub-CU) of the considered CU 700, as shown in Figure 7. Affine-style motion vectors are calculated from the control point motion vectors for each center position of each sub-block. The resulting MVs are expressed with 1 / 16-pel accuracy. Consequently, compensation of a coding unit in affine mode consists in the motion-compensated prediction of each sub-block by its own motion vector. These motion vectors for the sub-blocks are respectively shown as arrows for each of the sub-blocks in Figure 7.
[0057]
[84] In JEM, seeds are stored in corresponding 4x4 sub-blocks, so affine mode can only be used for CUs with widths and heights up to 4 (so that there are independent sub-blocks for each seed). For example, in a 64x4 CU, there is only one left sub-block for storing the top-left and bottom-left seeds, and in a 4x32 CU, there is only one upper sub-block for the top-left and top-right seeds. In JEM, it is impossible to properly store seeds in such thin CUs. According to our proposal, it is possible to process such thin CUs with widths or heights equal to 4 because the seeds are stored separately.
[0058]
[85] Referring again to the example of Figure 7, an affine CU is defined by an associated affine model consisting of three motion vectors, called affine model seeds, which are motion vectors from the top-left, top-right, and bottom-left corners of the CU (v0, v1, and v2 in Figure 7). This affine model then allows for the calculation of an affine motion vector field within the CU (black motion vectors in Figure 7), performed on a 4x4 sub-block basis. In JEM, these seeds are attached to the top-left, top-right, and bottom-left 4x4 sub-blocks of the considered CU. In the proposed solution, the affine model seeds are stored separately as motion information attached to the entire CU (e.g., IC flags). Thus, the motion model is decoupled from the motion vectors used for the actual motion compensation at the 4x4 block level. This new storage may allow for the storage of the complete motion vector field at the 4x4 sub-block level. This also allows for the use of affine motion compensation for blocks of size 4 in width or height.
[0059]
[86] Affine motion compensation can be used in JEM in two ways: Affine Inter (AF_INTER) mode and Affine Merging mode, which are described in the following sections.
[0060]
[87] Affine Inter (AF_INTER) Mode: In affine inter mode, CUs in AMVP mode with a size greater than 8x8 can be predicted. This is signaled by a flag in the bitstream. Generating an affine motion field for that inter CU involves determining a control point motion vector (CPMV) obtained by the decoder by adding a motion vector differential and a control point motion vector prediction (CPMVP). A CPMVP is a pair of motion vector candidates chosen from sets (A, B, C) and (D, E) shown in FIG. 8A for the current CU 800 being coded or decoded, respectively.
[0061]
[88] Affine Merge Mode: In affine merge mode, a CU-level flag indicates whether the merged CU uses affine motion compensation. If so, the first available neighboring CU coded in affine mode is selected from the ordered set of candidate positions A, B, C, D, E in Figure 8B for the current CU 880 being coded or decoded. However, the ordered set of candidate positions in this JEM is the same as the spatially neighboring candidates in merge mode in HEVC as shown in Figure 2A and described above.
[0062]
[89] Once the first neighboring CU in affine mode is obtained, three CPMVs from the top-left, top-right, and bottom-left corners of the neighboring affine CU are obtained.
number
number
number
number
[0063]
[90] Control point motion vector of the current CU
number
number
[0064]
[91] Therefore, a general aspect of at least one embodiment aims to improve the performance of the affine merge mode in JEM so that the compensation performance of the considered video codec can be improved. Accordingly, in at least one embodiment, an extended and improved affine motion compensation apparatus and method are presented, e.g., for coding units coded in the affine merge mode. The proposed extended and improved affine mode includes evaluating multiple predictor candidates in the affine merge mode.
[0065]
[92] As mentioned above, in the current JEM, the first neighboring CU coded in affine merge mode among the surrounding CUs is selected to predict the affine motion model associated with the current CU to be coded or decoded. That is, the first neighboring CU candidate among the ordered set (A, B, C, D, E) coded in affine mode in Figure 8B is selected to predict the affine motion model of the current CU.
[0066]
[93] Thus, at least one embodiment selects the affine merging prediction candidate that provides the best coding efficiency when encoding the current CU in affine merging mode, rather than using only the first one in the ordered set as described above. Thus, an improvement of this embodiment can be achieved at a general level by, for example, (with respect to the encoder / decoder) constructing a set of multiple affine merger predictor candidates that are likely to provide a good set of candidates for prediction of the affine motion model of the CU; (with respect to the encoder / decoder) selecting one predictor for the control point motion vectors of the current CU from the configured set, and / or (For encoder / decoder) Signal / decode the index of the control point motion vector predictor of the current CU Equipped with.
[0067]
[94] Accordingly, Figure 10 illustrates an exemplary encoding method 1000 according to a general aspect of at least one embodiment. At 1010, the method 1000 determines, for a block to be encoded in a picture, a set of predictor candidates having multiple predictor candidates. At 1020, the method 1000 selects a predictor candidate from the set of predictor candidates. At 1030, the method 1000 determines one or more corresponding control point motion vectors for the block for a predictor candidate selected from the set of predictor candidates. At 1040, the method 1000 determines, for the selected predictor candidate, a corresponding motion field based on a motion model for the selected predictor candidate based on the one or more corresponding control point motion vectors for the selected predictor candidate, where the corresponding motion field identifies a motion vector used for predicting a sub-block of the block to be encoded. At 1050, the method 1000 encodes the block based on the corresponding motion field for the predictor candidate selected from the set of predictor candidates. At 1060, the method 1000 encodes an index for the selected predictor candidate from the set of predictor candidates.
[0068]
[95] Figure 11 illustrates another exemplary encoding method 1100 according to a general aspect of at least one embodiment. At 1110, the method 1100 determines a set of predictor candidates for a block to be coded in a picture. At 1120, the method 1100 determines, for each of a plurality of predictor candidates in the set of predictor candidates, one or more corresponding control point motion vectors for the block. At 1130, the method 1100 determines, for each of the plurality of predictor candidates, a corresponding motion field based on a motion model for each of the plurality of predictor candidates in the set of predictor candidates based on the one or more corresponding control point motion vectors. At 1140, the method 1100 evaluates the plurality of predictor candidates according to one or more criteria and based on the corresponding motion field. At 1150, the method 1100 selects a predictor candidate from the plurality of predictor candidates based on the evaluation. At 1160, the method 1100 encodes an index for the selected predictor candidate from the set of predictor candidates.
[0069]
[96] Figure 12 illustrates an exemplary decoding method 1200 according to a general aspect of at least one embodiment. At 1210, method 1200 receives an index corresponding to a particular predictor candidate for a block to be decoded in a picture. In various embodiments, the particular predictor candidate has been selected in an encoder, and the index allows one of multiple predictor candidates to be selected. At 1220, method 1200 determines one or more corresponding control point motion vectors for the block to be decoded for the particular predictor candidate. At 1230, method 1200 determines a corresponding motion field based on the one or more corresponding control point motion vectors for the particular predictor candidate. In various embodiments, the motion field is based on a motion model, and the corresponding motion field identifies motion vectors used for prediction of sub-blocks of the block to be decoded. At 1240, method 1200 decodes the block based on the corresponding motion field.
[0070]
[97] Figure 13 illustrates another exemplary decoding method 1300 according to a general aspect of at least one embodiment. At 1310, method 1300 obtains an index corresponding to a selected predictor candidate for a block to be decoded in a picture. As also shown at 1310, the selected predictor candidate is selected in an encoder by: determining a set of predictor candidates for a block to be coded in the picture; for each of the plurality of predictor candidates in the set of predictor candidates, determining one or more corresponding control point motion vectors for the block to be coded; for each of the plurality of predictor candidates, determining a corresponding motion field based on a motion model for each of the plurality of predictor candidates in the set of predictor candidates based on the one or more corresponding control point motion vectors; evaluating the plurality of predictor candidates according to one or more criteria and based on the corresponding motion field; selecting a predictor candidate from the plurality of predictor candidates based on the evaluation; and encoding an index for the selected predictor candidate from the set of predictor candidates. At 1320, method 1300 decodes the block based on the index corresponding to the selected predictor candidate.
[0071]
[98] Figure 14 details an embodiment of a process 1400 used to predict the affine motion field of a current CU that is encoded or decoded in the existing affine merging mode in JEM. The input 1401 to this process 1400 is the current coding unit for which it is desired to generate an affine motion field for a sub-block, as shown in Figure 7. At 1410, the affine merging CPMV for the current block is obtained using selected predictor candidates, as described above with respect to Figures 6, 7, 8B, and 9. The derivation of these predictor candidates is described in more detail below with respect to Figure 15.
[0072] As a result, at 1420, the top left and top right control point motion vectors
number
number
[0073]
[0100] In at least one implementation, a residual flag is used. At 1450, a flag indicating that encoding was performed with residual data is activated (noResidual=0). At 1460, the current CU is fully encoded and reconstructed (with residual), resulting in a corresponding RD cost. Then, a flag indicating that encoding was performed without residual data is deactivated (1480, 1485, noResidual=1), and the process returns to 1460, where the CU is encoded (without residual) and a corresponding RD cost is generated. The lowest RD cost between the past two (1470, 1475) indicates whether the residual needs to be coded (normal or skip). Method 1400 ends at 1499. This best RD cost is then put into competition with other coding modes. Rate-distortion determination is described in more detail below.
[0074]
[0101] FIG. 15 shows details of an embodiment of a process 1500 used to predict one or more control points of the affine motion field of a current CU. This consists in searching for CUs coded / decoded in affine mode among the spatial locations (A, B, C, D, E) of FIG. 8B (1510, 1520, 1530, 1540, 1550). If none of the searched spatial locations are coded in affine mode, a variable indicating the number of candidate locations, e.g., numValidMergeCand, is set to 0 (1560). Otherwise, the first location corresponding to a CU in affine mode is selected (1515, 1525, 1535, 1545, 1555). Process 1500 then consists in calculating control point motion vectors that are subsequently used to generate the affine motion field assigned to the current CU, and setting numValidMergeCand to 1 (1580). This control point calculation proceeds as follows: The CU containing the selected location is determined. This is one of the neighboring CUs of the current CU, as described above. Next, three CPMVs from the top left, top right, and bottom left corners within the selected neighboring CU are calculated as described above with respect to FIG.
number
number
number
number
number
[0075]
[0102] The inventors have recognized that one aspect of the existing affine merging process described above is to systematically utilize one and only one motion vector predictor to propagate an affine motion field from surrounding informal (i.e., already coded or decoded) and neighboring CUs toward the current CU. In various circumstances, the inventors have further recognized that this aspect may be disadvantageous, for example, because it does not select the optimal motion vector predictor. Furthermore, as already described above, this predictor selection consists only of the first informal and neighboring CU coded in affine mode in the ordered set (A, B, C, D, E). In various circumstances, the inventors have further recognized that this limited selection may be disadvantageous, for example, because a better predictor may be available. Thus, the existing process in the current JEM does not consider that some possible informal and neighboring CUs surrounding the current CU may also have used affine motion, and that CUs other than the first CU found to have used affine motion may be better predictors for the motion information of the current CU.
[0076]
[0103] Accordingly, the present inventors have recognized potential advantages in several ways to improve current CU-affine motion vector prediction that are not utilized by existing JEM codecs. In accordance with general aspects of at least one embodiment, such advantages are found provided in the motion model of the present invention, as described below, and illustrated in Figures 16 and 17.
[0077]
[0104] In both Figures 16 and 17, the current CU to be coded or decoded is the large one in the center, designated 1610 in Figure 16 and 1710 in Figure 17, respectively. Two potential predictor candidates correspond to positions A and C in Figure 8B and are designated predictor candidates 1620 in Figure 16 and 1720 in Figure 17, respectively. In particular, Figure 16 shows potential motion fields for the current block to be coded or decoded 1610 when the selected predictor candidate is in the left position (position A in Figure 8B). Similarly, Figure 17 shows potential motion fields for the current block to be coded or decoded 1710 when the selected predictor candidate is in the top right position (i.e., position C in Figure 8B). As shown in the exemplary figures, depending on which affine merge predictor is selected, various sets of motion vectors for sub-blocks may be generated for the current CU. Therefore, the inventors recognize that a selection between these two candidates that optimizes one or more criteria, such as rate-distortion (RD), can help improve the encoding / decoding performance of the current CU in affine merging mode.
[0078]
[0105] Thus, one general aspect of at least one embodiment resides in selecting a better motion predictor candidate among a set of multiple candidates for deriving the CPMV of a current CU to be coded or decoded. At the encoder side, the candidate used to predict the current CPMV is selected according to a rate-distortion cost criterion according to one aspect of one exemplary embodiment. Its index is then coded in an output bitstream for a decoder according to another aspect of another exemplary embodiment.
[0079]
[0106] According to another aspect of other exemplary embodiments, at the decoder, a set of candidates may be constructed, and a predictor may be selected from this set in the same way as at the encoder side. In such embodiments, no index needs to be coded in the output bitstream. Other embodiments of the decoder avoid constructing a set of candidates, or at least avoid selecting a predictor from the same set as the encoder, and simply decode the index corresponding to the selected candidate from the bitstream and derive the corresponding associated data.
[0080]
[0107] According to other aspects of other exemplary embodiments, the CPMVs used herein are not limited to the two upper right and upper left positions of the current CU being coded or decoded, as shown in Figure 6. Other embodiments may comprise, for example, only one vector or more than two vectors, and the positions of these CPMVs may be, for example, at other corner positions or at any positions inside or outside the current block, such as, for example, at the center of a corner 4x4 sub-block or at one or more interior corner positions of a corner 4x4 sub-block, as long as it is possible to derive a motion field.
[0081]
[0108] In an exemplary embodiment, the set of potential candidate predictors investigated is the same as the set of positions (A, B, C, D, E) used to obtain the CPMV predictor in the existing affine merging mode in JEM, as shown in Figure 8B. Figure 18 details one exemplary selection process 1800 for selecting the best candidate to predict the affine motion model of the current CU, according to a general aspect of this embodiment. However, other embodiments use a set of predictor positions different from A, B, C, D, E, which may include fewer or more elements in the set.
[0082]
[0109] As shown in 1801, the input to this exemplary embodiment 1800 is also information about the current CU to be encoded or decoded. At 1810, a set of multiple affine merged predictor candidates is constructed according to the algorithm 1500 of FIG. 15 described above. The algorithm 1500 of FIG. 15 involves collecting all neighboring positions (A, B, C, D, E) shown in FIG. 8A that correspond to an informal CU coded in affine mode into a candidate set for predicting the affine motion of the current CU. Thus, the process 1800 does not end once an informal affine CU is found, but instead stores all possible candidates for affine motion model propagation from the informal CU to the current CU for all of the multiple motion predictor candidates in the set.
[0083]
[0110] Upon completion of the process of Figure 15 as shown at 1810 in Figure 18, the process 1800 of Figure 18 calculates at 1820 the CPMVs of the top left and top right corners predicted from each candidate in the set provided at 1810. This process of 1820 is further detailed and shown in Figure 19.
[0084]
[0111] Figure 19 again shows details of 1820 in Figure 18, which includes a loop over each of the discovered candidates determined from the previous step (1810 in Figure 18). For each affine merge predictor candidate, a CU containing the spatial location of that candidate is determined. Then, for each reference list L0 and L1 (at the base of the B slice), a control point motion vector is generated that is useful for generating the motion field of the current CU.
number
number
[0085]
[0112] Upon completion of the process of FIG. 19, the process returns to FIG. 18, where a loop 1830 over each affine merging predictor candidate is performed. This may, for example, select the CPMV candidate that results in the lowest rate-distortion cost. Within the loop 1830 over each candidate, another loop 1840, similar to the process shown in FIG. 14, is used to encode the current CU using each CPMV candidate, as described above. The algorithm of FIG. 14 terminates when all candidates have been evaluated, and its output may comprise an index of the best predictor. As described above, by way of example, the candidate with the lowest rate-distortion cost may be selected as the best predictor. Various embodiments use the best predictor to encode the current CU, and certain embodiments also encode an index for the best predictor.
[0086]
[0113] An example of determining the rate-distortion cost, as known to those skilled in the art, is: RD cost =D+λ×R where D represents the distortion (typically the L2 distance) between the original block and the reconstructed block obtained by encoding and decoding the current CU with the candidate under consideration, R represents the rate cost, e.g., the number of bits generated by encoding the current block with the candidate under consideration, and λ represents the rate target when the video sequence is being coded.
[0087]
[0114] Another exemplary embodiment is described below. This exemplary embodiment aims to further improve the coding performance of the affine merge mode by expanding the set of affine merge candidates compared to the existing JEM. This exemplary embodiment can be implemented on both the encoder side and the decoder side to expand the set of candidates. Thus, in one non-limiting aspect, several additional predictor candidates can be used to construct a set of affine merge candidates. The additional candidates can be taken from additional spatial locations, such as A' 2110 and B' 2120, surrounding the current CU 2100 as shown in FIG. 21 . Another embodiment uses even additional spatial locations along or adjacent to one of the edges of the current CU 2100.
[0088]
[0115] Figure 22 illustrates an exemplary algorithm 2200 corresponding to an embodiment using additional spatial locations A' 2110 and B' 2120, as shown in Figure 21 and described above. For example, algorithm 2200 includes testing a new candidate location A' if location A is not a valid affine merge prediction candidate (e.g., not within a CU coded in affine mode), at 2210-2230 of Figure 22. Similarly, for example, if location B does not provide any valid candidate (e.g., not within a CU coded in affine mode), then location B' is also tested, at 2240-2260 of Figure 22. Other aspects of exemplary process 2200 for constructing a set of affine merge candidates remain essentially unchanged compared to Figure 19, as previously shown and described.
[0089]
[0116] In another exemplary embodiment, existing merge candidate positions are considered first before evaluating the newly added position. The added position is evaluated only if the candidate set contains fewer candidates than a maximum number of merge candidates, for example, 5 or 7. The maximum number may be predetermined or may be variable. This exemplary embodiment is detailed by exemplary algorithm 2300 of FIG. 23.
[0090]
[0117] According to another exemplary embodiment, additional candidates, called temporary candidates, are added to the set of predictor candidates. These temporary candidates may be used, for example, if no spatial candidates are found, as described above, or, in a variant, if the size of the set of affine merge candidates has not reached a maximum, also as described above. Other embodiments use temporary candidates before adding spatial candidates to the set. For example, temporary candidates for predicting the control point motion vector of the current CU may be obtained from one or more of the reference pictures available or used for the current picture. The primary candidate may, for example, be employed at a position corresponding to the lower-right neighboring CU of the current CU in each of the reference pictures. This corresponds to candidate position F2410 for the current CU 2400 being coded or decoded as shown in FIG. 24.
[0091]
[0118] In an embodiment, for example, for each reference picture in each reference picture list, the affine flag associated with the block at position F2410 in Figure 24 in the reference picture under consideration is tested. If true, the corresponding CU contained in that reference picture is added to the current set of affine merging candidates.
[0092]
[0119] In a further variation, the temporary candidate is obtained from the reference picture at a spatial location corresponding to the top left corner of the current CU 2400. This location corresponds to candidate location G2420 in Figure 24.
[0093]
[0120] In a further variation, a temporary candidate is obtained from a reference picture at a position corresponding to the lower-right neighboring CU. Then, if the set of candidates includes fewer candidates than a pre-fixed maximum number of merged candidates, such as 5 or 7, a temporary candidate corresponding to the upper-left corner G2420 of the current CU is obtained. In other embodiments, temporary candidates are obtained from positions in one or more reference pictures that correspond to a different position (other than G2420) of the current CU 2400 or that correspond to another neighboring CU (other than F2410) of the current CU 2400.
[0094]
[0121] In addition, a typical derivation process for control point motion vectors based on temporary candidates proceeds as follows: For each temporary candidate included in the constructed set, a block (tempCU) containing the temporary candidate in the reference picture is identified. Then, three CPMVs located at the top left, top right, and bottom left corners of the identified temporary CU are used.
number
number
number
[0095]
[0122] Another exemplary embodiment includes adding an average control point motion vector pair calculated as a function of the control point motion vectors derived from each candidate. An exemplary process is now detailed by exemplary algorithm 2500 shown in Figure 25. A loop 2510 is performed for each affine merging predictor candidate in the set constructed for the reference picture list being considered.
[0096]
[0123] Then, in 2520, for each reference picture list Lx that is successively equal to L0 and then L1 (if in a B slice), if the current candidate has a valid CPMV for list Lx, Motion vector pairs
number
number
number
number
number
number
number
number
number
number
number
[0097]
[0124] Using algorithm 2500 and / or other embodiments, the set of affine merging candidates is further enriched to include average motion information calculated from the CPMV derived for each candidate inserted into the set of candidates according to the above-described embodiments as described in the preceding section.
[0098]
[0125] Since several candidates may result in the same CPMV for the current CU, the above-mentioned average candidate may result in a weighted average pair of CPMV motion vectors. In fact, the above-mentioned process calculates the average of the CPMVs collected so far, regardless of their uniqueness in the complete set of CPMVs. Therefore, a variation of this embodiment consists in adding another candidate to the set of CPMV prediction candidates. This consists in adding the average CPMV of the set of unique collected CPMVs (apart from the weighted average CPMV as described above). This provides further candidate CPMVs to the set of predictor candidates for generating the affine motion field of the current CU.
[0099]
[0126] For example, consider the following situation where five spatial candidates (L, T, TR, BL, TL) are all available and affine, except that the left three locations (L, BL, TL) are in the same neighboring CU. At each spatial location, a candidate CPMV can be obtained. Then, the first average is equal to the sum of these five CPMVs (some may be identical) divided by 5. In the second average, only the different CPMVs are considered, so only the left three (L, BL, TL) are considered once, and the second average is equal to the three different CPMVs (L, T, TR) divided by 3. In the first average, the extra CPMVs are added three times, giving more weight to the extra CPMVs. Using the formula, this can be written as Average 1 = (L + T + TR + BL + TL) / 5, and L = BL = TL, so Average 1 = (3 * L+T+TL) / 5, and the average 2=(L+T+TL) / 3.
[0100]
[0127] The two candidate averages mentioned above are bidirectional as long as the candidate under consideration has a motion vector relative to a reference image in list 0 and another image in list 1. In another variation, a unidirectional average can be added. Four unidirectional candidates can be constructed by taking motion vectors from list 0 and list 1 separately from the weighted average and unique average.
[0101]
[0128] One advantage of the exemplary candidate set expansion method described in this application is an increase in diversity in the set of candidate control point motion vectors that can be used to construct an affine motion field associated with a given CU. Thus, embodiments of the present disclosure provide a technical advance in computational techniques for encoding and decoding video content. For example, embodiments of the present disclosure improve the rate-distortion performance provided by the affine merge coding mode in JEM. In this way, the overall rate-distortion performance of the considered video codec is improved.
[0102]
[0129] To modify the process of FIG. 18 , a further exemplary embodiment may be provided. This embodiment involves a rapid evaluation of the performance of each CPMV candidate by the following approximate distortion and rate calculations. Thus, for each candidate in the CPMV set, the motion field of the current CU is calculated, and a 4×4 subblock-based temporal prediction of the current CU is performed. Next, the distortion is calculated as the SATD between the predicted CU and the original CU. The rate cost is obtained as the approximate degree of the bits associated with the signaling of the merge index of the considered candidate. A rough (approximate) RD cost is then obtained for each candidate. In one embodiment, the final selection is based on the approximate RD cost. In another embodiment, a subset of candidates is subjected to a full RD search, i.e., the candidate with the lowest approximate RD cost is then subjected to a full RD search. An advantage of these embodiments is that they limit the increase in encoder-side complexity caused by searching for the best affine merge predictor candidate.
[0103]
[0130] Also, according to another general aspect of at least one embodiment, the affine inter mode as described above may be improved using all of the current teachings presented in this disclosure by having an expanded list of affine predictor candidates. As described above with respect to FIG. 8A, one or more CPMVPs of an affine inter CU are derived from neighboring motion vectors, regardless of their coding mode. Thus, similar to the affine merge mode as described above, it is then possible to average the affine neighborhood using those affine models to construct one or more CPMVPs of the current affine inter CU. In this case, the affine candidates considered may be the same list as described above with respect to the affine merge mode (e.g., not limited to only spatial candidates).
[0104]
[0131] Therefore, a set of multiple predictor candidates is provided to improve the compression / decompression provided by current HEVC and JEM by using better predictor candidates, the process becomes more efficient, and coding gains are observed even if a supplemental index may need to be transmitted.
[0105]
[0132] According to a general aspect of at least one embodiment, the set of affine merge candidates (which, like the merge mode, has at least seven candidates) may be, for example, · Spatial candidates from (A, B, C, D, E), - Time candidate in bottom right juxtaposition if there are less than 5 candidates in the list, · Time candidates for juxtaposition, if there are less than 5 candidates in the list; ·Weighted average, ·Unique average, · One-way average from weighted average if weighted average is two-way and there are less than 7 candidates in the list, If the unique mean is bidirectional and there are less than 7 candidates in the list, then the unidirectional mean from the unique mean It consists of:
[0106]
[0133] Also, in the case of AMVP, the predictor candidates are, e.g. · Spatial candidates from the set (A, B, C, D, E), · Complementary space candidates from (A', B'), Time candidate for bottom right juxtaposition It can be adopted from
[0107]
[0134] Tables 1 and 2 below show improvements over JEM4.0 (parallel) using several exemplary embodiments of the solutions proposed in this disclosure. Each table shows the results for the amount of rate reduction for one of the exemplary embodiments as described above. In particular, Table 1 shows the improvement when the five spatial candidates (A, B, C, D, and E) shown in FIG. 8B are used as the set of predictor candidates according to the exemplary embodiment described above. Table 2 shows the improvement for the exemplary embodiment when the following predictor candidate order is used as described above: first the spatial candidates, then the temporal candidates if the number of candidates is still less than 5, then the average, and finally the unidirectional average if the number of candidates is still less than 7. Thus, for example, Table 2 shows that for this embodiment, the rate reductions for Y, U, and V samples are 0.22%, 0.26%, and 0.12% BD (Bjontegaard-Delta) rate reductions for class D, respectively, with little increase in encoding and decoding execution time (i.e., 100% and 101%, respectively). Thus, exemplary embodiments of the present disclosure improve compression / decompression efficiency over existing JEM implementations while maintaining computational complexity costs. [Table 1] [Table 2]
[0108]
[0135] Figure 26 shows a block diagram of an exemplary system 2600 in which various aspects of exemplary embodiments may be implemented. System 2600 may be embodied as a device including various components described below and configured to perform the processes described above. Examples of such devices include, but are not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. System 2600 may be communicatively coupled to other similar systems and displays via communication channels such as those shown in Figure 26 to implement all or part of the exemplary video system described above, as known to those skilled in the art.
[0109]
[0136] Various embodiments of the system 2600 include at least one processor 2610 configured to execute loaded instructions to implement the various processes described above. The processor 2610 may include embedded memory, input / output interfaces, and various other circuits as known in the art. The system 2600 may also include at least one memory 2620 (e.g., a volatile memory device, a non-volatile memory device). The system 2600 may further include a storage device 2640, which may include non-volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, a magnetic disk drive, and / or an optical disk drive. The storage device 2640 may comprise, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device. The system 2600 may also include an encoder / decoder module 2630 configured to process data to provide encoded and / or decoded video, which may include its own processor and memory.
[0110]
[0137] The encoder / decoder module 2630 represents a module or modules that may be included in a device to perform encoding and / or decoding functions. As is known, such a device may include either or both encoding and decoding modules. Additionally, the encoder / decoder module 2630 may be implemented as a separate element of the system 2600 or may be incorporated within one or more processors 2610 as a combination of hardware and software, as is known to those skilled in the art.
[0111]
[0138] Program code that is loaded into one or more processors 2610 to perform the various processes described above may be stored in storage device 2640 and then loaded into memory 2620 for execution by processor 2610. According to an exemplary embodiment, one or more of processor(s) 2610, memory 2620, storage device 2640, and encoder / decoder module 2630 may store one or more of various items during the performance of the processes described above, including, but not limited to, input video, decoded video, bitstream, equations, formulas, metrics, variables, operations, and operating logic.
[0112]
[0139] System 2600 may also include a communication interface 2650 that enables communication with other devices over a communication channel 2660. Communication interface 2650 may include, but is not limited to, a transceiver configured to transmit and receive data from communication channel 2660. Communication interface 2650 may include, but is not limited to, a modem or a network card, and communication channel 2650 may be implemented in a wired and / or wireless medium. The various components of system 2600 may be connected or communicatively coupled to each other using various suitable connections, including, but not limited to, an internal bus, a wire, and a printed circuit board (not shown in FIG. 26).
[0113]
[0140] Exemplary embodiments may be performed by the processor 2610 or computer software implemented by hardware, or by a combination of hardware and software. As a non-limiting example, exemplary embodiments may be implemented by one or more integrated circuits. The memory 2620 may be of any type suitable for the technology environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 2610 may be of any type suitable for the technology environment and may include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0114]
[0141] The implementations described herein may be realized in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if described in the context of only one type of implementation (e.g., described only as a method), implementations of the described features may be realized in other forms (e.g., an apparatus or a program). An apparatus may be realized in, for example, appropriate hardware, software, and firmware. A method may be realized in an apparatus such as, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0115]
[0142] Moreover, those skilled in the art will readily appreciate that the exemplary HEVC encoder 100 shown in Figure 1 and the exemplary HEVC decoder shown in Figure 3 can be modified in accordance with the above teachings of the present disclosure to implement the disclosed improvements to the existing HEVC standard to achieve better compression / decompression. For example, the entropy coding 145, motion compensation 170, and motion estimation 175 in the exemplary encoder 100 of Figure 1, and the entropy decoding 330 and motion compensation 375 in the exemplary decoder of Figure 3, can be modified in accordance with the disclosed teachings to implement one or more exemplary aspects of the present disclosure, including providing advanced affine merging prediction to the existing JEM.
[0116]
[0143] References to "one embodiment" or "embodiment" or "one implementation" or "implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation," as well as any other variations thereof, appearing in various places throughout this specification are not necessarily all referring to the same embodiment.
[0117]
[0144] Additionally, the application or claims may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0118]
[0145] Additionally, the application or claims may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0119]
[0146] Additionally, the application or claims may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Also, "receiving" generally involves some activity, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0120]
[0147] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
Claims
1. 1. A method for video encoding, comprising: accessing a set of predictor candidates for a block to be coded in a picture, the set having a plurality of predictor candidates, the predictor candidates corresponding to coded spatial or temporal neighboring blocks, the block being coded in an affine merging mode; selecting a predictor candidate from the set of predictor candidates; obtaining a set of control point motion vectors for the block using a plurality of motion vectors associated with the selected predictor candidate from the set of predictor candidates; obtaining a motion field based on a motion model based on the set of control point motion vectors, the motion field identifying motion vectors used for prediction of all sub-blocks of the block to be coded, the set of control point motion vectors for the block being stored separately from the motion vectors of the motion field for the block; encoding the block based on the motion field; encoding an index for the selected predictor candidate from the set of predictor candidates; A method for providing the above.
2. The method of claim 1 , wherein the motion model is an affine model.
3. The method of claim 1 , wherein the stored motion vectors of the motion field are for motion compensation of the block.
4. 2. The method of claim 1, further comprising accessing a second set of control point motion vectors for the selected predictor candidate, wherein the second set of control point motion vectors is stored separately from the motion vectors of all sub-blocks of the selected predictor candidate, the set of control point motion vectors for the block is obtained in response to the second set of control point motion vectors for the selected predictor candidate, and the motion vectors of all sub-blocks of the selected predictor candidate are for motion compensation of the selected predictor candidate.
5. 1. A method for video decoding, comprising: accessing an index corresponding to a predictor candidate for a block to be decoded in a picture, the predictor candidate corresponding to a decoded spatial or temporal neighboring block, the block being decoded in an affine merging mode; obtaining a set of control point motion vectors for the block being decoded using a plurality of motion vectors associated with the candidate predictors; obtaining a motion field based on a motion model based on the set of control point motion vectors, the motion field identifying motion vectors used for prediction of all sub-blocks of the block being decoded, the set of control point motion vectors for the block being stored separately from the motion vectors of the motion field for the block; decoding the block based on the motion field; A method for providing the above.
6. The method of claim 5 , wherein the motion model is an affine model.
7. The method of claim 5 , wherein the stored motion vectors of the motion field are for motion compensation of the block.
8. 6. The method of claim 5, further comprising accessing a second set of control point motion vectors for the predictor candidate, wherein the second set of control point motion vectors is stored separately from the motion vectors of all sub-blocks of the predictor candidate, the set of control point motion vectors for the block is obtained in response to the second set of control point motion vectors for the predictor candidate, and the motion vectors of all sub-blocks of the predictor candidate are for motion compensation of the predictor candidate.
9. 1. An apparatus for video encoding, comprising: one or more processors, accessing a set of predictor candidates for a block to be coded in a picture, the set having a plurality of predictor candidates, the predictor candidates corresponding to coded spatial or temporal neighboring blocks, the block being coded in an affine merging mode; selecting a predictor candidate from the set of predictor candidates; obtaining a set of control point motion vectors for the block using a plurality of motion vectors associated with the selected predictor candidate from the set of predictor candidates; obtaining a motion field based on a motion model based on the set of control point motion vectors, the motion field identifying motion vectors used for prediction of all sub-blocks of the block to be coded, the set of control point motion vectors for the block being stored separately from the motion vectors of the motion field for the block; encoding the block based on the motion field; encoding an index for the selected predictor candidate from the set of predictor candidates. Device.
10. The apparatus of claim 9 , wherein the motion model is an affine model.
11. The apparatus of claim 9 , wherein the stored motion vectors of the motion field are for motion compensation of the block.
12. 10. The apparatus of claim 9, wherein the one or more processors are further configured to access a second set of control point motion vectors for the selected predictor candidate, the second set of control point motion vectors being stored separately from motion vectors of all sub-blocks of the selected predictor candidate, the set of control point motion vectors for the block being obtained in response to the second set of control point motion vectors for the selected predictor candidate, and the motion vectors of all sub-blocks of the selected predictor candidate are for motion compensation of the selected predictor candidate.
13. 1. An apparatus for video decoding, comprising: one or more processors, accessing an index corresponding to a predictor candidate for a block to be decoded in a picture, the predictor candidate corresponding to a decoded spatial or temporal neighboring block, the block being decoded in an affine merging mode; obtaining a set of control point motion vectors for the block being decoded using a plurality of motion vectors associated with the candidate predictors; obtaining a motion field based on a motion model based on the set of control point motion vectors, the motion field identifying motion vectors used for prediction of all sub-blocks of the block being decoded, the set of control point motion vectors for the block being stored separately from the motion vectors of the motion field for the block; and decoding the block based on the motion field. Device.
14. The apparatus of claim 13 , wherein the motion model is an affine model.
15. The apparatus of claim 13 , wherein the stored motion vectors of the motion field are for motion compensation of the block.
16. 14. The apparatus of claim 13, wherein the one or more processors are further configured to access a second set of control point motion vectors for the predictor candidate, the second set of control point motion vectors being stored separately from motion vectors of all sub-blocks of the predictor candidate, the set of control point motion vectors for the block being obtained in response to the second set of control point motion vectors for the predictor candidate, and the motion vectors of all sub-blocks of the predictor candidate are for motion compensation of the predictor candidate.
Citation Information
Patent Citations
Image prediction method and related device
JP2018511997A
Prediction image generation device, moving image decoding device, and moving image encoding device
WO2017130696A1
Image encoding / decoding method and device
WO2017133243A1