Motion vector refinement for multi-reference prediction

By estimating and refining motion vectors in multi-reference prediction using template matching and calculation, the method reduces complexity and maintains accuracy in motion estimation, enhancing coding efficiency.

JP2025100555APending Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025043809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-06-30
Filing Date
2025-03-18
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Motion vector refinement in multi-reference prediction is computationally complex and requires significant resources, especially when determining motion vectors in multiple reference pictures.

Method used

The method involves estimating initial motion vectors in two reference pictures and refining one vector through template matching while calculating the other based on the refined vector, reducing the search space and complexity.

Benefits of technology

This approach maintains accuracy in motion estimation with reduced computational complexity and memory requirements, improving coding performance without affecting picture quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100555000001_ABST
    Figure 2025100555000001_ABST
Patent Text Reader

Abstract

To perform motion vector refinement in a search space for multireference interprediction, in which two or more reference pictures are selected and one of them is used for motion vector refinement.SOLUTION: A search space in a reference image is constructed on the basis of the initial estimation of a motion vector to a reference picture for motion vector refinement. Template matching is used to refine a first motion vector. A second motion vector for another reference picture is calculated using the initial estimation of the second motion vector, the initial estimation of the first motion vector, and the refined first motion vector.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video coding, and more particularly to motion vector estimation applicable in multi-reference inter prediction. This application claims the priority of PCT / EP2017 / 066342, filed on June 30, 2017, the content of which is incorporated herein by reference.

Background Art

[0002] Current hybrid video codecs use predictive coding. The pictures of a video sequence are subdivided into pixel blocks, which are then coded. Instead of coding the block pixel by pixel, the entire block is predicted using already coded pixels that are spatially or temporally close to the block. The encoder further processes only the difference between the block and its prediction. Further processing typically includes a conversion of the block pixels to coefficients in the transform domain. The coefficients may then be further compressed by quantization and further compressed by entropy coding to form a bitstream. The bitstream further includes any signaling information that enables the decoder to decode the encoded video. For example, the signaling may include settings regarding encoder settings such as the size of the input picture, the frame rate, quantization step indication, prediction applied to the blocks of the picture, and the like.

[0003] Temporal prediction utilizes the temporal correlation between pictures, also called frames, of a video. Since temporal prediction is a prediction that uses the dependencies (inter) between different video frames, it is also called inter prediction. Thus, an encoded block, also called the current block, is predicted from one or more previously encoded pictures, also called reference pictures. The reference picture does not have to be a picture before the current picture in which the current block is arranged in the display order of the video sequence. The encoder may encode pictures in a coding order different from the display order. As the prediction of the current block, a block at the same position within the reference picture may be determined. A block at the same position is a block located within the reference picture at the same position as the current block within the current picture. Such prediction is accurate for stationary picture regions, i.e., picture regions with no movement from one picture to another.

[0004] To obtain a predictor that takes motion into account, i.e., a motion-compensated predictor, motion estimation is typically used when determining the prediction of the current block. Thus, the current block is predicted by a block within the reference picture that is located at a distance given by the motion vector from the position of the block at the same position. To enable the decoder to determine the same prediction of the current block, the motion vector may be signaled within the bitstream. To further reduce the signaling overhead caused by signaling the motion vector for each block, the motion vector itself may be estimated. Motion vector estimation may be performed based on the motion vectors of adjacent blocks in the spatial and / or temporal domain.

[0005] The prediction of the current block may be calculated using one reference picture or by weighting predictions obtained from two or more reference pictures. Since adjacent pictures are most likely to be similar to the current picture, the reference pictures may be adjacent pictures, i.e., the picture immediately before and / or after in the display order. However, in general, the reference pictures may also be any other picture before the current picture in the bitstream (decoding order), either before or after the current picture in the display order. This may provide an advantage, for example, in the case of occlusions and / or non-linear motion in video content. Thus, the reference pictures may also be signaled within the bitstream.

[0006] A special mode of inter prediction is the so-called dual prediction where two reference pictures are used in generating the prediction of the current block. Specifically, the two predictions determined in each of the two reference pictures are combined into the prediction signal of the current block. Dual prediction may result in a more accurate prediction of the current block than single prediction, i.e., prediction using only a single reference picture. A more accurate prediction leads to a smaller difference (also called "residual") between the pixels of the current block and the prediction, which may be more efficiently coded, i.e., compressed into a shorter bitstream. In general, three or more reference pictures may also be used to predict the current block, and three or more reference blocks may be found respectively, i.e., multi-reference inter prediction may be applied. Thus, the term multi-reference prediction includes dual prediction as well as prediction using three or more reference pictures.

[0007] To provide more accurate motion estimation, the resolution of the reference picture may be enhanced by interpolating samples between pixels. Fractional pixel interpolation may be performed by weighted averaging of the closest pixels. In the case of half-pixel resolution, for example, bilinear interpolation is typically used. Other fractional pixels are calculated as the average of the closest pixels weighted by the reciprocal of the distance between the respective closest pixel and the predicted pixel.

[0008] Motion vector estimation is a computationally complex task in which the similarity is calculated between the current block and the corresponding predicted block pointed to by a candidate motion vector in the reference picture. To reduce complexity, some candidate motion vectors are usually reduced by restricting the candidate motion vectors to a specific search space. The search space may be defined, for example, by the number of pixels and / or positions surrounding the position in the reference picture corresponding to the position of the current block in the current image. On the other hand, the candidate motion vectors may be defined by a list of candidate motion vectors formed by the motion vectors of adjacent blocks.

[0009] Motion vectors are usually at least partially determined on the encoder side and notified to the decoder in the coded bitstream. However, motion vectors may also be derived at the decoder. In such a case, the current block is not available at the decoder and cannot be used to calculate the similarity with the block pointed to by the candidate motion vector in the reference picture. Therefore, instead of the current block, a template composed of the pixels of the already decoded block is used. For example, already decoded pixels adjacent to the current block (in the current picture or the reference picture) may be used. Such motion estimation provides the advantage of reducing signaling, and the motion vectors are derived in the same way at both the encoder and the decoder, and thus no signaling is required. On the other hand, the accuracy of such motion estimation may be lower.

[0010] To provide a trade-off between accuracy and signaling overhead, motion vector estimation may be split into two steps: motion vector derivation and motion vector refinement. For example, motion vector derivation may include selection of a motion vector from a list of candidates. Such selected motion vectors may be further refined, e.g., by searching within a search space. Searching within the search space is based on calculating a cost function for each candidate motion vector, i.e., for each candidate position of the block pointed to by the candidate motion vector.

[0011] Document JVET-D0029: Decoder-Side Motion Vector Refinement Based on Bilateral Template Matching, X. Chen, J. An, J. Zheng (the document can be found at the http: / / phenix.it-sudparis.eu / jvet / site) shows motion vector refinement where a first motion vector at integer pixel resolution is found and further refined by searching at half-pixel resolution in a search space around the first motion vector. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION

[0012] When multi-reference prediction is applied, motion vectors in multiple reference pictures have to be determined. Even when motion vectors are signaled at a first stage so that the decoder does not need to perform any further search, motion vector refinement still requires additional search between the motion vectors of the corresponding search space. This can be a complex task that requires computational resources as well as memory. MEANS FOR SOLVING THE PROBLEM

[0013] The present disclosure provides a technique for determining a first motion vector in a first reference picture and a second motion vector in a second reference picture. Accordingly, the complexity can be reduced. First, the first motion vector and the second motion vector are roughly estimated. Then, the first motion vector is refined by performing a search within a search space given by the rough estimate of the first motion vector. The second motion vector is determined by calculations based on its rough estimate and the refined first motion vector. The first and second motion vectors may be applied in the inter prediction of a current block in the current picture at a decoder on the encoding and / or decoding side.

[0014] According to a first aspect, the present invention relates to an apparatus for determining a first motion vector in a first reference picture and a second motion vector in a second reference picture, the first and second motion vectors being applied in the inter prediction of a picture block in a current picture, the apparatus comprising a motion vector refinement unit and a motion vector calculation unit. The motion vector refinement unit is configured to obtain an estimated value of the first motion vector. A search space is specified based on the estimated value of the first motion vector. Within the search space, the motion vector refinement unit performs a search to determine the first motion vector. The motion vector calculation unit obtains an estimated value of the second motion vector. Based on the estimated value of the second motion vector and the first motion vector, the motion vector calculation unit calculates the second motion vector.

[0015] Accordingly, the determination of the motion vectors can be performed with less complexity while still maintaining the accuracy given by the refinement of the first motion vector and estimating the amount of refinement for the second motion vector based thereon.

[0016] In a possible implementation of the apparatus according to such a first aspect, the second motion vector is calculated by adding to the estimated value of the second motion vector a function of the difference between the first motion vector and the estimated value of the second motion vector. This function may include scaling and / or clipping. When the scaling parameter is determined, its value may depend on the ratio between the respective distances of the first reference picture and the second reference picture to the current picture.

[0017] The calculation of the second motion vector as a refinement function performed on the first motion vector is a low-complexity estimation. Further, by further modifying it (e.g., by scaling according to the distance between the respective reference pictures), the estimation can be made even more accurate.

[0018] The apparatus preferably further includes a reference picture selection unit for obtaining reference pictures and selecting which of them are the first reference picture and the second reference picture. Following this selection, it is determined which of the first reference picture or the second reference picture should be used for motion vector refinement. The list of reference pictures associates the indices to be included in the bitstream with the positions of the reference pictures relative to the current picture. The reference picture selection unit is configured to select the first reference picture and the second reference picture based on whether they are referenced in the bitstream by indices within a predefined list of reference pictures.

[0019] In other words, the reference picture selection unit is configured to select either the first picture or the second picture based on whether each of the first or second pictures is referenced in the bitstream, which also includes the coded picture blocks of the video, by an index associated with a predefined list of reference pictures between at least two lists of reference pictures, and the list of reference pictures associates the index with the position of the reference picture relative to the current picture.

[0020] If two reference pictures are referenced in a bitstream by indices within the same predefined list of reference pictures, the reference picture selection unit may select, as the reference picture to be used for motion vector refinement, the picture having the highest position within said list of reference pictures.

[0021] Alternatively, the reference picture to be used for motion vector refinement may be selected as the picture having the lowest temporal layer among two pictures.

[0022] The reference picture to be used for motion vector refinement may likewise be selected as the picture having the lowest base quantization value, or as the picture having the minimum distance to the current picture.

[0023] As a further approach, the reference picture to be used for motion vector refinement may be selected such that the magnitude of the estimated value of the motion vector pointing to the reference picture to be used for motion vector refinement is smaller than the magnitudes of the estimated values of other motion vectors.

[0024] The apparatus may further comprise a motion vector determination unit for determining an estimated value of a first motion vector and an estimated value of a second motion vector. This determination is performed by selecting them from a set of motion vector candidates based on the similarity between a template and the portions of the respective pictures referenced by the motion vector candidates.

[0025] A video encoder for encoding a plurality of pictures into a bitstream includes an inter prediction unit, a bitstream formatter, and a reconstruction unit. The inter prediction unit includes a device for determining a first motion vector and a second motion vector, and a prediction unit. The prediction unit determines a prediction block according to a portion of a first reference picture referenced by the first motion vector and a portion of a second reference picture referenced by the second motion vector. The bitstream formatter includes the estimated value of the first motion vector and the estimated value of the second motion vector in the bitstream. The reconstruction unit reconstructs a current block according to the prediction block and stores the reconstructed block in a memory.

[0026] A video decoder for decoding a plurality of pictures from a bitstream includes an inter prediction unit, a bitstream parser, and a reconstruction unit. The inter prediction unit includes a device for determining a first motion vector and a second motion vector, and a prediction unit. The prediction unit determines a prediction block according to a portion of a first reference picture referenced by the first motion vector and a portion of a second reference picture referenced by the second motion vector. The bitstream parser obtains the estimated value of the first motion vector and the estimated value of the second motion vector from the bitstream. The reconstruction unit reconstructs a current block according to the prediction block.

[0027] The method includes motion vector refinement and motion vector calculation. An estimated value of a first motion vector is obtained. A search space is specified based on the estimated value of the first motion vector. Within the search space, a search for determining the first motion vector is executed. An estimated value of a second motion vector is obtained. Based on the estimated value of the second motion vector and the first motion vector, the second motion vector is calculated. The second motion vector is calculated by adding a function of a difference between the first motion vector and the estimated value of the first motion vector to the estimated value of the second motion vector. This function includes scaling and / or clipping. The value of the scaling parameter depends on a ratio between respective distances of the first reference picture and the second reference picture to the current picture.

[0028] The method further comprises a reference picture selection for obtaining reference pictures and selecting which of them are the first reference picture and the second reference picture. Following this selection, it is determined which of the first reference picture or the second reference picture should be used for motion vector refinement. The list of reference pictures associates the indices to be included in the bitstream with the positions of the reference pictures relative to the current picture. The reference picture selection is performed to select the first reference picture and the second reference picture based on whether they are referenced in the bitstream by indices within a predefined list of reference pictures. If the two reference pictures are referenced in the bitstream by indices within the same predefined list of reference pictures, the reference picture to be used for motion vector refinement is selected as the picture having the highest position within said list of reference pictures. Alternatively, the reference picture to be used for motion vector refinement can be selected as the picture having the lowest temporal layer among the two pictures. The reference picture to be selected for motion vector refinement can similarly be selected as the picture having the lowest base quantization value or as the picture having the minimum distance to the current picture. As a further approach, the reference picture to be used for motion vector refinement can be selected such that the magnitude of the estimated value of the motion vector pointing to the reference picture to be used for motion vector refinement is smaller than the magnitudes of the estimated values of the other motion vectors.

[0029] The method may further determine an estimated value of a first motion vector and an estimated value of a second motion vector. This determination is performed by selecting them from a set of motion vector candidates based on the similarity between the template and the respective portions of the pictures referenced by the motion vector candidates.

[0030] A video encoding method for encoding a plurality of pictures into a bitstream comprises performing inter prediction, bitstream formation, and block reconstruction. The inter prediction includes determining a first motion vector and a second motion vector, and block prediction. The prediction includes determining a prediction block according to a portion of a first reference picture referenced by the first motion vector and a portion of a second reference picture referenced by the second motion vector. The bitstream formation includes including an estimated value of the first motion vector and an estimated value of the second motion vector in the bitstream. The reconstruction includes reconstructing a current block according to the prediction block and storing the reconstructed block in a memory.

[0031] A video decoding method for decoding a plurality of pictures from a bitstream comprises performing inter prediction, bitstream analysis, and block reconstruction. The inter prediction includes determining a first motion vector and a second motion vector, and block prediction. The prediction includes determining a prediction vector according to a portion of a first reference picture referenced by the first motion vector and a portion of a second reference picture referenced by the second motion vector.

[0032] The bitstream analysis includes obtaining an estimated value of the first motion vector and an estimated value of the second motion vector from the bitstream. The reconstruction includes reconstructing a current block according to the prediction block.

[0033] The present invention can reduce the number of search candidates in the process of motion vector refinement without affecting the coding performance while providing similar picture quality. This is achieved by performing a search for motion vector refinement only on one reference picture related to the current block, and another motion vector related to another reference picture of the same current block is calculated based on the refined motion vector.

[0034] Hereinafter, exemplary embodiments will be described in more detail with reference to the accompanying drawings and figures.

Brief Description of the Drawings

[0035]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0036] The present disclosure relates to the determination of motion vectors for multi-reference prediction. It may be used in motion estimation performed during video encoding and decoding. Hereinafter, exemplary encoders and decoders that may perform motion estimation using the search space structure of the present disclosure will be described below.

[0037] FIG. 1 shows an encoder 100 comprising an input for receiving an input block of a frame or picture of a video stream and an output for generating an encoded video stream. The term "frame" in the present disclosure is used as a synonym for a picture. However, it should be noted that the present disclosure is also applicable to fields where interlacing is applicable. Generally, a picture includes m times n pixels. This corresponds to an image sample and may comprise one or more color components. For simplicity, the following description refers to pixels meaning luminance samples. However, it should be noted that the motion vector search of the present invention can be applied to any color component including chrominance, or components of a search space such as RGB for example. On the other hand, it may be beneficial to perform motion vector search for only one component and apply the determined motion vector to more (or all) components.

[0038] The input blocks to be coded do not necessarily have to be of the same size. One picture may include blocks of different sizes, and the block raster of different pictures may also be different.

[0039] In an illustrative implementation, encoder 100 is configured to apply prediction, transformation, quantization, and entropy coding to the video stream. Transformation, quantization, and entropy coding are respectively performed by a transformation unit 101, a quantization unit 102, and an entropy coding unit 103 so as to generate an encoded video bitstream as an output.

[0040] A video stream may include a plurality of frames, and each frame is divided into blocks of a specific size that are either intra- or inter-coded. For example, the blocks of the first frame of the video stream are intra-coded by the intra prediction unit 109. An intra frame is coded using only the information within the same frame so that it can be decoded independently and provide an entry point within the bitstream for random access. The blocks of other frames of the video stream may be inter-coded by the inter prediction unit 110, and the information from previously coded frames (reference frames) is used to reduce temporal redundancy such that each block of the inter-coded frame is predicted from a block within the reference frame. The mode selection unit 108 is configured to select whether the blocks of a frame should be processed by the intra prediction unit 109 or the inter prediction unit 110. This block also controls the intra and inter prediction parameters. To enable refreshing of the image information, intra-coded blocks may be provided within the inter-coded frames. Further, an intra frame containing only intra-coded blocks may be periodically inserted into the video sequence to provide an entry point for decoding, i.e., a point at which the decoder can start decoding without the information from previously coded frames.

[0041] The intra prediction unit 109 is a block prediction unit. In order to perform spatial or temporal prediction, the coded block may be further processed by the inverse quantization unit 104 and the inverse transform unit 105. After the block is reconstructed, the loop filtering unit 106 is applied to further improve the quality of the decoded image. The filtered block then forms a reference frame that is next stored in the frame buffer 107. Such a decoding loop (decoder) on the encoder side provides the advantage of generating the same reference frame as the reference picture reconstructed on the decoder side. Thus, the encoder side and the decoder side operate in corresponding ways. Here, the term "reconstruction" refers to obtaining a reconstructed block by adding a prediction block to the decoded residual block.

[0042] The inter prediction unit 110 receives as input a block of the current frame or picture to be coded and one or several reference frames or pictures from the frame buffer 107. Motion estimation and motion compensation are applied by the inter prediction unit 110. Motion estimation is used to obtain a motion vector and a reference frame based on a specific cost function. Motion compensation then describes the current block of the current frame with respect to the transformation of the reference block of the reference frame, i.e., by the motion vector. The inter prediction unit 110 outputs a prediction block for the current block, and the prediction block minimizes a cost function. For example, the cost function may be the difference between the current block to be coded and its prediction block, i.e., the cost function minimizes the residual block. The minimization of the residual block is based on, for example, calculating the sum of the absolute differences (SAD) between all pixels (samples) of the current block and all pixels (samples) of the candidate block in the candidate reference picture. However, generally, any other similarity metric such as the mean squared error (MSE) or the structural similarity metric (SSIM) may be used.

[0043] However, the cost function may also be the number of bits necessary to code such an inter-block and / or the distortion resulting from such coding. Thus, the rate-distortion optimization procedure may be used to determine, in motion vector selection and / or, generally, in coding parameters, which of inter-prediction or intra-prediction to use for a block and in which settings.

[0044] The intra prediction unit 109 receives, as inputs, a block of the current frame or picture to be intra-coded and one or several reference samples from the already reconstructed region of the current frame. The intra prediction then describes the pixels of the current block of the current frame with respect to a function of the reference samples of the current frame. The intra prediction unit 109 outputs a prediction block for the current block, and the prediction block advantageously minimizes the difference between the current block to be coded and its prediction block, i.e., minimizes the residual block. The minimization of the residual block may be based, for example, on a rate-distortion optimization procedure. In particular, the prediction block is obtained as a directional interpolation of the reference samples. The direction may be determined by rate-distortion optimization and / or by calculating a similarity measure as described above in relation to inter-prediction.

[0045] The difference between the current block and its prediction, i.e., the residual block, is then transformed by the transformation unit 101. The transformation coefficients are quantized by the quantization unit 102 and entropy-coded by the entropy coding unit 103. The encoded video bitstream thus generated comprises intra-coded blocks and inter-coded blocks, and corresponding signaling (e.g., mode indication, motion vector indication, and / or intra prediction direction). The transformation unit 101 may apply a linear transformation such as a Fourier or discrete cosine transform (DFT / FFT or DCT). Such a transformation into the spatial frequency domain offers the advantage that the resulting coefficients typically have higher values at lower frequencies. Thus, after efficient coefficient scanning (such as zigzag) and quantization, the resulting sequence of values typically has some larger values at the beginning and end, along with a run of zeros. This enables further efficient coding. The quantization unit 102 performs actual irreversible compression by reducing the resolution of the coefficient values. The entropy coding unit 103 then assigns binary codewords to the coefficient values to generate the bitstream. The entropy coding unit 103 also codes signaling information (not shown in FIG. 1).

[0046] FIG. 2 shows a video decoder 200. The video decoder 200 comprises a plurality of reference picture buffers 207 and an intra prediction unit 209 which is a block prediction unit. The reference picture buffers 207 are configured to store at least one reference frame reconstructed from the encoded video bitstream, the reference frame being different from the current frame (the currently decoded frame) of the encoded video bitstream. The intra prediction unit 209 is configured to generate a predicted block which is an estimate of the block to be decoded. The intra prediction unit 209 is configured to generate this prediction based on reference samples obtained from the reference picture buffers 207.

[0047] Decoder 200 is configured to decode the encoded video bitstream generated by video encoder 100. Preferably, both decoder 200 and encoder 100 generate the same prediction for each block to be encoded / decoded. The features of reference picture buffer 207 and intra prediction unit 209 are the same as those of reference picture buffer 107 and intra prediction unit 109 in FIG. 1.

[0048] Video decoder 200 includes an inverse quantization unit 204, an inverse transform unit 205, and a loop filtering unit 206, which respectively correspond to further units existing within video encoder 100, such as inverse quantization unit 104, inverse transform unit 105, and loop filtering unit 106 of video encoder 100.

[0049] Entropy decoding unit 203 is configured to decode the received encoded video bitstream and obtain the quantized residual transform coefficients and the signaling information accordingly. The quantized residual transform coefficients are supplied to inverse quantization unit 204 and inverse transform unit 205 to generate a residual block. The residual block is added to the prediction block, and the added result is supplied to loop filtering unit 206 to obtain the decoded video. The frame of the decoded video is stored in reference picture buffer 207 and can function as a reference frame for inter prediction.

[0050] Generally, intra prediction units 109 and 209 in FIGS. 1 and 2 can use reference samples from already encoded regions to generate prediction signals for blocks that need to be encoded or decoded.

[0051] The entropy decoding unit 203 receives the encoded bitstream as its input. Generally, the bitstream is first parsed, i.e., the signaling parameters and the residuals are extracted from the bitstream. Typically, the syntax and meaning of the bitstream are defined by a standard so that the encoder and decoder may operate in a mutually compatible manner. As described in the background section above, the encoded bitstream contains not only the prediction residuals. In the case of motion-compensated prediction, the motion vector indication is also encoded in the bitstream and parsed therefrom in the decoder. The motion vector indication may be given by the reference picture to which the motion vector is provided and by the motion vector coordinates. So far, coating the complete motion vector has been considered. However, only the difference between the current motion vector and the previous motion vector in the bitstream may also be encoded. This technique makes it possible to utilize the redundancy between the motion vectors of adjacent blocks.

[0052] To efficiently code reference pictures, the H.265 codec (ITU-T, H265, Series H: Audiovisual and multimedia systems: High Efficient Video Coding) provides lists of reference pictures that assign each reference frame to a list index. The reference frames are then signaled in the bitstream by including the corresponding assigned list index within them. Such lists may be defined in the standard or signaled at the start of a video or a set of several frames. Note that in H.265, there are two defined lists of reference pictures called L0 and L1. The reference pictures are then signaled within the bitstream by indicating the list (L0 or L1) and the index within that list associated with the desired reference picture. Providing more than two lists may have advantages for better compression. For example, L0 may be used for both unidirectionally interpredicted slices and bidirectionally interpredicted slices, and L1 may be used only for bidirectionally interpredicted slices. However, in general, the present disclosure is not limited to any content of the L0 and L1 lists.

[0053] The lists L0 and L1 may be defined or fixed in the standard. However, higher flexibility in coding / decoding can be achieved by signaling them at the start of the video sequence. Thus, the encoder may construct the lists L0 and L1 using specific reference pictures ordered according to an index. The L0 and L1 lists may have the same fixed size. In general, there may be three or more lists. Motion vectors may be signaled directly by coordinates within the reference picture. Alternatively, as also specified in H.265, a list of candidate motion vectors may be constructed and the index within the list associated with a particular motion vector may be transmitted.

[0054] The motion vector of the current block usually correlates with the motion vectors of adjacent blocks within the current picture or within previously coded pictures. This is because adjacent blocks are likely to correspond to the same moving object with similar motion, and the motion of the object is less likely to change suddenly over time. Therefore, using the motion vectors within adjacent blocks as predictors reduces the size of the difference in the signaled motion vectors. The MVP is usually derived from already decoded motion vectors from spatially adjacent blocks or from temporally adjacent blocks within the picture at the same position. In H.264 / AVC, this is done by taking the median value for each component of three spatially adjacent motion vectors. Using this technique, the signaling of the predictor is not required. The temporal MVP from pictures at the same position is considered only in the so-called temporal direct mode of H.264 / AVC. The H.264 / AVC direct mode is also used to derive other motion data besides the motion vectors. Therefore, they are related by the concept of block merging in HEVC. In HEVC, the technique of implicitly deriving the MVP has been replaced by a technique known as motion vector competition that explicitly signals which MPV from the list of MVPs is used for motion vector derivation. The variable coding quadtree block structure in HEVC can result in one block having several adjacent blocks with motion vectors as potential MVP candidates. Taking the left neighbor as an example, in the worst case, if a 64×64 luma coding tree block is not further divided and the left block is divided to the maximum depth, a 64×64 luma prediction block can have 16 4×4 luma prediction blocks on the left.

[0055] To modify motion vector competition to account for such flexible block structures, Advanced Motion Vector Prediction (AMVP) was introduced. During the development of HEVC, the initial AMVP design was significantly simplified to provide a good trade-off between coding efficiency and an implementation-friendly design. The initial design of AMVP included five MVPs from three different classes of predictors, namely, three motion vectors from spatially adjacent ones, the median of three spatial predictors, and a scaled motion vector from temporally adjacent blocks at the same position. Further, the list of predictors was modified by sorting it to place the most likely motion predictor in the first position and removing redundant candidates to ensure minimal signaling overhead. The final design of the AMVP candidate list structure consists of the following two MVP candidates: a) up to two spatial candidate MVPs derived from five spatially adjacent blocks, b) one temporal candidate MVP derived from two blocks at the same temporal position if both spatial candidate MVPs are unavailable or they are identical, and c) a zero motion vector if spatial candidates, temporal candidates, or both are unavailable. Details regarding motion vector determination can be found in the book by V. Sze et al. (Ed.), High Efficiency Video Coding (HEVC): Algorithms and Architectures, Springer, 2014, which is incorporated herein by reference, particularly in Chapter 5.

[0056] In order to further improve motion vector estimation without further increasing signaling overhead, it may be beneficial to further refine the motion vectors derived on the encoder side and provided in the bitstream. Motion vector refinement may be performed at the decoder without assistance from the encoder. The encoder in that decoder loop may use the same refinement to obtain the corresponding reference picture. Refinement is performed by determining a template, determining a search space, and finding the portion of the reference picture that best matches the template within the search space. The position of the best matching portion is then used to determine the best motion vector to obtain the predictor for the current block, i.e., the currently reconstructed current block.

[0057] FIG. 3 shows a search space (search region) 310 that includes integer pixel positions (full dots) and fractional pixel positions (hollow dots) of a reference picture. In this example, the fractional pixel positions are half pixel positions. As described above, the fractional pixel positions may be obtained from integer (full pixel) positions by interpolation such as bilinear interpolation.

[0058] In the bi-prediction of the current block, two predicted blocks obtained using the respective first motion vectors of list L0 and the second motion vectors of list L1 are combined into a single prediction signal, and this single prediction signal provides a better fit to the original signal than uni-prediction, resulting in less residual information and possibly more efficient compression. FIG. 3 further shows the current block 320 from the current picture. At the decoder, since the current block is not decoded and thus not available for motion vector refinement purposes, a template constructed based on the portion of the picture that has already been processed (i.e., coded on the encoder side and decoded on the decoder side) is used for the estimation of the current block.

[0059] The template may be constructed, for example, based on samples that have already been decoded, i.e., samples belonging to the current picture that were decoded before the current block. Additionally or alternatively, these samples may belong to any of the previously decoded reference pictures. By way of example, the samples to be used in constructing the template may belong to a reference picture that was decoded before the current picture and that precedes the current picture in display order. Alternatively, the samples may belong to a reference picture that was decoded before the current picture and that follows the current picture in display order. Finally, the template may be constructed based on a combination of samples from two different reference pictures. As will be apparent to those skilled in the art, the template may be obtained using different methods such that the current block can be estimated using the constructed template.

[0060] First, the estimated value MV0 of the first motion vector and the estimated value MV1 of the second motion vector are received as inputs to the decoder 200 as shown in FIG. 3. On the encoder side 100, the motion vector estimated values MV0 and MV1 may be obtained by block matching and / or by searching in a list of candidates (such as a merge list) formed by the motion vectors of blocks adjacent to the current block (within the same picture or in an adjacent picture). MV0 and MV1 are then advantageously signaled to the decoder side within the bitstream. However, it should be noted that generally, the first decision stage in the encoder may also be performed by template matching, which offers the advantage of reducing signaling overhead.

[0061] On the decoder side 200, the motion vectors MV0 and MV1 are advantageously obtained based on the information in the bitstream. MV0 and MV1 are either directly signaled or differentially signaled, and / or an index within a list of motion vectors (merge list) is signaled. However, the present disclosure is not limited to signaling motion vectors in the bitstream. Rather, the motion vectors may already be determined by template matching in a first stage corresponding to the operation of the encoder. The template matching in the first stage (motion vector derivation) may be performed based on a search space different from the search space of the second refinement stage. Specifically, the refinement may be performed in a search space having a higher resolution (i.e., a shorter distance between search positions).

[0062] Indications of two reference pictures pointed to by respective MV0 and MV1 are similarly provided to the decoder. The reference pictures are stored in a reference picture buffer in both the encoder and the decoder as a result of previous processing, i.e., respective encoding and decoding. One of these reference pictures is selected by search for motion vector refinement. A reference picture selection unit of the apparatus for determining a motion vector is configured to select a first reference picture pointed to by MV0 and a second reference picture pointed to by MV1. Following the selection, the reference picture selection unit determines which of the first reference picture or the second reference picture is to be used for execution of motion vector refinement. In FIG. 3, the first reference picture pointed to by motion vector MV0 is selected for the search. To perform motion vector refinement, a search area 310 in the first reference picture is defined around a candidate position pointed to by motion vector MV0. Candidate search space positions within the search area 310 are analyzed to find the block most similar to the template block by performing template matching within the search space and determining a similarity metric such as sum of absolute differences (SAD). As described above, in one implementation, the template is constructed based on a combination of samples from two different reference pictures each having respective motion vectors of MV0 and MV1. Template matching is performed to find a block based on a point within the search area 310 most similar to the template. Alternatively, in another implementation, the template is constructed to find a similarity between a predicted block P0 generated based on MV0 in L0 and a predicted block P1 generated based on MV1 in L1. To perform refinement for MV0, template matching comprises finding a block based on a point within the search area 310 determined by similarity matching (e.g., SAD, etc.) between P0 and P1. The position of the search space 310 indicates a position where the upper left corner of the template 320 coincides. As already described above, the upper left corner is merely a convention and any point of the search space such as the center point 330 can generally be used to indicate the matching position.

[0063] The candidate position having the lowest SAD value is determined as the motion vector MV0". In FIG. 3, the position pointed to by MV0" is a half-pixel position, which is different from the initial estimated MV0 by 1.5 pixel positions in the vertical direction, but remains the same in the horizontal direction.

[0064] According to the present disclosure, for single prediction or multi-reference prediction, at least one motion vector of the current block is refined not by performing template matching, but by calculation based on another refined motion vector of the current block. FIG. 4 shows such refinement. Specifically, the motion vector MV1" is calculated based on the motion vector estimate MV1 and a function of MV0" - MV0 instead of performing a second template matching. In the example of FIG. 4, the determination and refinement of the first motion vector MV0" are performed as described above with reference to FIG. 3. Further, the motion vector MV1" is calculated by subtracting the difference between MV0" and MV0 from the second motion vector estimate MV1".

[0065] This approach utilizes the fact that most of the motion within a video sequence falls within the category of "translational motion". In translational motion, an object moves at a constant velocity (at least between frames that are close to each other in sampling time). This means that within consecutive frames (when the time sampling period does not change over time), the object is displaced by the same pixel distance in the x and y directions. The present invention utilizes the principle of translational motion to some extent.

[0066] In the above example, the first motion vector MV0" is refined by template matching and the second motion vector is refined by calculation. However, according to the present disclosure, a selection process may be further performed to select which motion vectors should be refined by template matching and which should be determined by calculation. FIG. 5 shows a block diagram of an example of a motion vector refiner 500.

[0067] As shown in FIG. 8, the motion vector refiner 500 may be included in an apparatus 810 for determining a motion vector. The apparatus may be included in an inter prediction unit 800 that may replace the inter prediction unit 110 shown in FIG. 1 and / or in the inter prediction unit 210 shown in FIG. 2.

[0068] More specifically, an apparatus 810 is provided for determining a first motion vector in a first reference picture and a second motion vector in a second reference picture. The first and second motion vectors may be applied in the inter prediction of a picture block in the current picture.

[0069] The apparatus 810 includes a motion vector refiner 500. As shown in detail in FIG. 5, the motion vector refiner 500 further includes a motion vector refinement unit 530 configured to obtain an estimated value MV0 of the first motion vector MV0" and determine the first motion vector MV0" by performing a search within a specified search space based on the estimated value MV0. The apparatus further includes a motion vector calculation unit 550 configured to obtain an estimated value MV1 of the second motion vector MV1" and calculate the second motion vector MV1" based on the estimated value MV1 of the second motion vector MV1" and based on the first motion vector MV0".

[0070] In this example, the apparatus includes a first branch including a motion vector calculation unit 530 and a motion vector calculation unit 550, a second branch including a motion vector calculation unit 540 and a motion vector calculation unit 560, and a switch 520 for enabling one of the two branches and disabling the other. The second branch is similar to the first branch and mainly differs from the first branch in that it outputs a first motion vector MV1" and a second motion vector MV0" by processing MV1 as an estimated value of the first motion vector and processing MV0 as an estimated value of the second motion vector.

[0071] More specifically, the motion vector refinement unit 540 is configured to obtain an estimated value MV1 of the first motion vector MV1", and determine the first motion vector MV1" by performing a search within a search space specified based on the estimated value MV1. The apparatus further includes a motion vector calculation unit 560 configured to obtain an estimated value MV0 of the second motion vector MV0", and calculate the second motion vector MV0" based on the estimated value MV0 of the second motion vector MV0" and based on the first motion vector MV1".

[0072] FIG. 5 shows the application of the present invention to dual prediction where there are two motion vectors, namely motion vectors MV0" and MV1", within each of the two determined reference pictures. Thus, the selection of the picture for refinement by template matching is also performed simply by taking one of MV0 and MV1 for template matching and maintaining the other MV1 or MV0 for calculation.

[0073] This process of decoder motion vector refinement (DMVR) is performed by the apparatus 500. Depending on which of the respective motion vector estimated values MV0 and MV1 the template search should be performed on, the motion vector refinement unit 530 or 540 receives the respective motion vector estimated value MV0 or MV1 from the motion vector estimator 820 at the input 505 and sets a search region 310 around MV0 or MV1. The size of the search region in FIGS. 3 and 4 is 3×3 integer pixels, with half pixels interpolated to 7×7, and thus a total of 49 positions. In general, the shape and / or size of the search region may vary, and the present invention functions regardless of the size and shape of the search region. The search region may have a predetermined or predefined size. In other words, the shape and size of the search space may be fixed and specified in a standard. Alternatively, one of several possible shapes and sizes may be selected manually by the user within the encoder settings or automatically based on the video content.

[0074] Some examples of the size and shape of the search space are shown in FIG. 6. The complete triangle indicates the center pixel of the search space, and the complete circle indicates the remaining search space positions. Note that the search space may be further extended by fractional (half pixel, quarter pixel, etc.) interpolation. The present disclosure is generally not limited to any particular pixel pattern.

[0075] For each position or fractional position within the search area, template matching is performed using a template that approximates the current block, providing an SAD value for each search space coordinate. Note that the integer pixel resolution and half pixel resolution herein describe the resolution of the search space, i.e., the displacement of the search position with respect to the non-refined motion vectors input to the process. As a result, the search coordinates do not necessarily coincide with the actual pixel coordinates on the image plane. In other words, the integer pixel (sample) resolution of the search space does not necessarily mean that the search space position is located on the integer pixels of the reference image. The integer position of the search space may coincide with the fractional pixels of the reference image, especially when the initial motion vectors MV0, MV1 point to fractional pixels within the reference image.

[0076] A comparison of the SAD values of the positions within the search area is performed to determine the position with the minimum SAD value. The position with the minimum SAD value is determined as the motion vector MV0". As described in the background section, SAD is merely an example, and any other metric such as MSE, SSIM, correlation coefficient, etc. may generally be used. The determined motion vector MV0" is passed to the motion vector calculation unit 550 together with an estimated value of the second motion vector MV1, where the determination (refinement) of the motion vector MV1" is performed by calculation.

[0077] As already described above with reference to FIG. 4, as a first exemplary approach, the motion vector calculation unit 550 is configured to calculate a second motion vector by adding to the estimated value of the second motion vector, the difference between the first motion vector and the estimated value of the first motion vector, or a function of the difference between the first motion vector and the estimated value of the first motion vector.

[0078] Adding the difference can be calculated as follows: MV1" = MV1+(MV0 - MV0") (Equation 1) This Equation 1 (adding the negative difference MV0" - MV0) functions for the case shown in FIG. 4 when one of the two reference pictures is located before the current picture and the other one is located after the current picture. As can be seen in FIG. 4, for estimating normal motion, the difference between the current motion vector and the first motion vector in the previous picture is projected onto the picture following the current picture with the reversed sign.

[0079] If both reference frames with respect to MV0 and MV1 are located either after or before the current picture, then the difference is added directly without changing the sign, MV1" = MV1+(MV0" - MV0) (Equation 2) resulting in. The above-mentioned before / after positions mean those in the display order. The display order of pictures in a video may be different from the encoding / decoding order, i.e., the order in which the coded pictures are embedded in the bitstream. The display order of pictures may be given by the Picture Order Count (POC). In particular, the POC in H.265 / HEVC is signaled in the slice header of each slice that is a picture or a part thereof.

[0080] The condition used to determine whether the two reference pictures are those following the current picture and those preceding the current picture may be implemented using the parameter POC as follows: (POCi - POC0)*(POCi - POC1) < 0 (Equation 3) Here, POCi is the display order (sequence number) of the current picture, POC0 is the display order of the first reference picture, and POC1 is the display order of the second reference picture. If the condition in Equation 3 is true, then either the first reference picture precedes the current picture and the second picture follows the current picture, or vice versa. On the other hand, if the condition of Equation 3 is not true, then either both reference pictures precede the current picture or both follow the current picture. However, the condition may be implemented in any way that can check whether the signs of the two difference expressions (POCi - POC0) and (POCi - POC1) are the same. The above Equation 3 is just an example using the multiplication "*" for its relatively simple implementation form.

[0081] Adding a difference function can be calculated as follows in the case of bidirectional prediction where one of the reference pictures precedes the current picture and the other follows the current picture (the condition in Equation 3 is true). MV1" = MV1 - f(MV0" - MV0) (Equation 4) Again, if both reference pictures precede the current picture in the display order or both follow the current picture (the condition of Equation 3 is false), then the calculation adds the difference without inverting the sign. MV1" = MV1 + f(MV0" - MV0) (Equation 5) For example, the function may include scaling and / or clipping.

[0082] When the function is scaling, the calculation of the motion vector may be as follows MV1" = MV1 - k*(MV0" - MV0) = MV1 + k*(MV0 - MV0") (Equation 6) Here, "*" represents multiplication (scaling), and k is a scaling parameter. When k = 1, Equation 1 is obtained. Using a fixed (pre-defined) k, Equation 6 is applicable when one of the reference pictures precedes the current picture and the other follows the current picture. If both reference pictures precede the current picture or both follow the current picture, the sign needs to be inverted. MV1" = MV1 + k * (MV0" - MV0) = MV1 - k * (MV0 - MV0") (Equation 7)

[0083] According to one embodiment, the value of the scaling parameter depends on the ratio between the respective distances of the first and second pictures with respect to the current picture. Thus, the value of k is not pre-defined and not fixed, but may vary according to the reference pictures for which the motion vectors are defined. For example, the scaling coefficient k may be k = -(POCi - POC1) / (POCi - POC0) (Equation 8) given by. Note that since the scaling parameter changes the sign according to whether both reference pictures are on the same side (preceding or following) of the current picture in the display order or on different sides of the current picture, the above Equation 8 already takes into account the sign change.

[0084] Even when it may be beneficial to consider the distances between the respective first and second reference pictures with respect to the current image, and even if these distances are different from each other, setting k = 1 as in Equations 1 and 2 may still be applied. It is simpler, and the quality of the refined motion vectors may still be substantially higher than without applying the refinement.

[0085] After the scaling operation, a rounding or clipping operation may be applied. Here, rounding refers to providing an integer or real number with a predefined precision closest to the input value. Clipping refers to removing digits from the input value with a precision higher than the predefined precision. This can be particularly beneficial from the perspective of fixed-point arithmetic applied in typical computing systems.

[0086] Since the motion vector refinement is performed only for one of the two reference pictures, the search space for performing template matching is reduced by 50%.

[0087] After the above-described processing for the current block in the current picture, two reference pictures as well as their respective associated motion vectors MV0" and MV1" are provided at output 580. These motion vectors are used, for example, to determine a predictor for the current block by averaging two respective predictors obtained by taking respective portions of the reference pictures pointed to by the motion vectors MV0" and MV1" that correspond to the current block in size and shape. This is performed by prediction unit 890.

[0088] Generally, prediction unit 890 determines a prediction block by combining a portion of a first reference picture referred to by a first motion vector MV0" and a portion of a second reference picture referred to by a second motion vector MV1".

[0089] The combined prediction signal can provide a better fit to the original signal than single prediction, resulting in less residual information and possibly more efficient compression.

[0090] As described in the previous paragraph, the output motion vectors MV0" and MV1" of apparatus 500 may result in less residual information. Thus, using them may improve the prediction performance as compared to the initial motion vectors MV0 and MV1. It should be noted that apparatus 500 may be used multiple times to further improve the prediction performance. In this case, the output of the first application of apparatus 500 is considered as the input of the second application, and the equalities (Equations 1-8) between the motion vectors are maintained after each application. In this case, since the motion vectors MV0" and MV1" provide a better estimation of the current block after the first application of apparatus 500, the template used in the MV refinement unit 530 is constructed based on the samples pointed to by MV0" or the samples pointed to by MV1" in the second application.

[0091] According to an exemplary embodiment, apparatus 500 further comprises a reference picture selection unit 510 configured to obtain two reference pictures and select which of them should be the first reference picture and the second reference picture.

[0092] In other words, the reference picture selection unit 510 controls, for the current block, which of its motion vectors should be refined by template matching and which should be refined by calculations based on the refinement of another refined motion vector. Some examples regarding how the selection process may be performed by the reference picture selection unit 510 are provided below.

[0093] FIG. 7 shows a schematic diagram illustrating a method 700 for determining a first motion vector in a first reference picture and a second motion vector in a second reference picture according to an embodiment. For example, the digital video encoder 100 or decoder 200 described above, particularly the inter prediction unit 110 or 210, may use process 700 to obtain the first and second motion vectors to be applied in the inter prediction of a picture block in the current picture. The method 700 includes step 701 of obtaining an estimated value of the first motion vector, step 703 of determining the first motion vector by performing a search within a search space specified based on the estimated value of the first motion vector, step 705 of obtaining an estimated value of the second motion vector, and step 707 of calculating the second motion vector based on the first motion vector based on the estimated value of the second motion vector.

[0094] Example 1 In a first example, the reference picture selection unit 510 determines which motion vectors are to be determined by a search within the search space based on a list of reference pictures to which the respective motion vectors belong that have been referenced. Specifically, some codecs signal the reference pictures by including in the bitstream an index associated with a particular reference picture in the list of reference pictures in order to restore the same predictors used by the encoder in the bitstream for use by the decoder. For example, the list of reference pictures (reference picture list) may be a table that is available in both the encoder and the decoder and that associates an index with the relative position of each reference picture with respect to the current picture.

[0095] There may be two or more lists of reference pictures. For example, reference picture list L0 and reference picture list L1 that are commonly used in H.265 / HEVC. To signal a reference picture within a bitstream, first, the list of reference pictures is signaled, and then the index within the signaled reference picture list is signaled.

[0096] The reference picture selection unit 510 is then advantageously configured to select either the first reference picture or the second reference picture based on whether it is referenced within the bitstream by an index within a predefined list of reference pictures. In this context, the term "predefined" means, for example, being fixedly defined in a standard or being defined for the entire video in signaling. Exemplarily, the predefined list may be list L0. Then, if the first reference picture is referenced from reference picture list L0 and the second list is referenced from reference picture list L1, the first motion vector pointing to the first reference picture is refined by template matching as it is referenced from L0, and the second motion vector is calculated as it is not referenced from L0. However, it should be noted that the predefined list is not limited to the L0 list. Any of the reference picture lists used may be predefined instead. Usually, L0 lists reference pictures in a closer neighborhood than L1. Alternatively, L0 may include more reference pictures that precede the current picture in display order, and L1 may include more pictures that follow the current picture in display order. However, the present invention functions regardless of any specific differences between two or more reference picture lists.

[0097] Note that a situation may occur where both the first and second reference pictures pointed to by their respective first and second motion vectors are referenced from the same reference picture list. For example, both the first and second reference pictures may be referenced from the pictures in a pre-defined list L0. Alternatively, when the current coding block applies bi-prediction, one reference picture from list L0 and one reference picture from list L1 must be indicated, and both reference pictures may be included together within one (or both) of the reference lists. The reason is that the reference pictures can exist within both lists (list L0 and list L1).

[0098] When both the first reference picture and the second reference picture are included within a pre-defined list (such as L0), the picture at the highest position within the list (L0) is selected as the reference picture to be used for motion vector refinement by template matching of the corresponding motion vectors that point to them.

[0099] If both pictures are included within a list of reference pictures that is not pre-defined (such as L1 when L0 is pre-defined), the same selection may be performed. In other words, when the reference pictures are referenced from a list of reference pictures other than the pre-defined list of reference pictures, the reference picture having the highest position within the list is selected for template matching-based motion vector refinement.

[0100] In summary, when two reference pictures are referenced within a bitstream by indices within the same pre-defined list of reference pictures, the reference picture selection unit 510 selects the first reference picture as the picture having the highest position within the said list of reference pictures.

[0101] This example provides a simple implementation form without any additional intelligent selection process. Specifically, the reference picture selection unit 510, based on the index value of the reference picture, controls switches 520 and 570 to select the upper branch or the lower branch in the block diagram of FIG. 5 when the analyzed reference picture list is a pre-defined list (such as L0) and both pictures are in the same list.

[0102] Example 2 According to the second example, the reference picture selection unit 510 is configured to select, as the first reference picture (to be refined by template matching), the picture with the lowest temporal layer among two (or more) reference pictures related to the current reference block. In other words, the reference picture selection unit 510 controls switches 520 and 580 to select the upper branch (530, 550) or the lower branch (540, 560) according to the temporal layers of the reference pictures to which the motion vectors MV0 and MV1 are related.

[0103] In FIG. 5, the upper branch and the lower branch do not necessarily need to be implemented in duplicate. Generally, in any of the embodiments and examples of the present disclosure, a single motion vector refinement unit 530 and a single motion vector calculation unit 550 are provided, and only the input to a single branch is switched according to the control of the reference picture selection unit 210.

[0104] Specifically, the temporal layers of two (or more) reference pictures are determined. The temporal layer of a reference picture indicates the number of pictures that must be decoded before the reference picture can be decoded. The temporal layer is typically set in a decoder that encodes video sequences at different temporal layers. It may be included in the bitstream. Thus, the decoder may determine which reference picture belongs to which temporal layer based on the signaling information from the bitstream. Thus, a first reference picture or a second reference picture having a lower temporal layer is then selected as the reference picture to be used for motion vector refinement by template matching. This approach relies on fewer previously decoded pictures and thus can offer the advantage of selecting a reference picture with a lower likelihood of errors and artifacts for template matching. Thus, the motion vector refinement procedure is made more flexible.

[0105] Example 3 In a third example, the reference picture selection unit 510 is configured to select, as the first reference picture (to be refined by template matching), the picture having the lowest base quantization value. In other words, the reference picture selection unit 510 controls switches 520 and 580 to select an upper branch (530, 550) or a lower branch (540, 560) according to the quantization parameters of the reference pictures to which motion vectors MV0 and MV1 are related.

[0106] The quantization value or quantization parameter in this context is the information provided within a bitstream that can determine the quantization step. In well-known codecs such as H.264 / AVC and H.265 / HEVC, the quantization parameter can determine the value by which the coefficients to be quantized should be divided. The larger the quantization value, the coarser the quantization, which typically leads to poor image quality after reconstruction. Thus, a lower quantization value means that higher quality of the reconstructed image can be provided. The selection of a reference picture with a lower quantization parameter means that a reference picture with better quality is used for motion vector refinement, which in turn leads to better refinement results.

[0107] The term "base" quantization value refers to the quantization value that is common to a picture slice and is used as a base for all blocks. Usually, such a value is signaled, for example, in the slice header. Then, typically, the difference from the base value is signaled for each block or processing unit.

[0108] However, the present invention is not limited to any particular signaling or even the existence of such values. The same effect can be achieved by determining the base quantization value for a picture according to the quantization values of the elements within the picture where the quantization value is signaled. In other words, the term "base quantization value" indicates the general quantization value for a picture.

[0109] Example 4 According to the fourth example, the reference picture selection unit 510 is configured to select, as the first reference picture, the picture having the shortest distance to the current picture. In other words, the reference picture selection unit 510 controls the switches 520 and 580 to select the upper branch (530, 550) or the lower branch (540, 560) according to the distances of the reference pictures associated with the respective motion vectors MV0 and MV1 to the current picture.

[0110] For example, the differences between the picture order count (POC) values POC0 and POC1 of the respective reference pictures associated with the respective motion vectors MV0 and MV1 and the POC value POCi of the current picture are determined. The POC value specifies the display order of pictures, rather than coding / decoding. Thus, a picture with POC = 2 is displayed before a picture with POC = 8. However, since the present invention is not limited to applications in well-known codecs such as H.264 / AVC and H.265 / HEVC, the difference between the reference picture and the current picture may be determined in any other way that does not depend on specific POC parameters.

[0111] Since it is expected that the motion vector of the closer reference picture is more accurate and / or the reference block pointed to by the motion vector is more similar to the current block, the first reference picture associated with the motion vector MV0 or the second reference picture associated with the motion vector MV1 having the smallest absolute POC distance (between the reference picture and the current picture) is selected as the reference picture to be used for motion vector refinement. This may lead to better quality of refinement.

[0112] Example 5 According to the fifth example, the reference picture selection unit is configured to select the first reference picture and the second reference picture such that the magnitude of the estimated value of the first vector is smaller than the magnitude of the estimated value of the second motion vector. In other words, the reference picture selection unit 510 controls the switches 520 and 580 to select the upper branches (530, 550) or the lower branches (540, 560) according to the lengths (magnitudes) of the motion vectors MV0 and MV1 associated with the respective reference pictures.

[0113] The absolute magnitudes of the motion vectors MV0 and MV1 that point to the first reference picture and the second reference picture are determined using standard procedures for determining the absolute value of a vector. For example, the squared values of each vector component of the motion vector are summed. Either this sum or its square root may be used as the magnitude of the motion vector, although the square root calculation has a higher computational cost. Taking the motion vector with the smaller magnitude for refinement offers the advantage that it is more likely to be accurately determined assuming that the motion between pictures is typically small.

[0114] Some information regarding the estimated value of the first motion vector MV0, the estimated value of the second motion vector MV1, and the indexes of the reference pictures that MV0 and MV1 refer to may be received at the decoder as input. Motion vector information is typically signaled in block units, and the blocks can have different sizes. The same is true for reference picture indications. A bitstream parser implemented as part of the entropy decoding unit 203 obtains the motion vector information from the bitstream. The motion information may be the coordinates of the motion vector directly (coordinates with respect to the point (0,0) given by the position of the block in the same reference picture as the position of the current block within the current picture). Alternatively, the difference from the motion vector of a block preceding the current block in the decoding order may be signaled. This may advantageously be one of the spatially or temporally adjacent blocks of the current block.

[0115] According to another example, an apparatus for motion vector determination, which also includes a motion vector refiner 500, is configured to determine them by selecting an estimated value of a first motion vector and an estimated value of a second motion vector from a set of motion vector candidates based on the similarity of a template with each portion of the picture referred to by the motion vector candidates. In other words, the motion vector determination (of MV0 and MV1) is not necessarily based on template matching in a search space defined in a reference picture. The search space may be given by a (merge) list that lists indices in relation to the motion vectors of spatially or temporally adjacent blocks, or blocks near the current block. This means that the present invention is not limited by the method by which the estimated motion vectors MV0 and MV1 are derived before being provided for refinement.

[0116] In summary, the dual prediction operation of one coding block, two prediction blocks from the motion vectors (MVs) of list L0 and the MVs of list L1 respectively, is combined into a single prediction signal, which can provide a better fit to the original signal than single prediction, resulting in less residual information and possibly more efficient compression. The dual prediction decoding process for the current block in the current picture includes the following processing steps.

[0117] First, the estimated values of the first motion vector MV0 and the second motion vector MV1 are received as inputs on the decoder side. Since the two reference pictures pointed to by MV0 and MV1 have already been decoded before the processing of the current picture, they are in the decoder's picture buffer. One of these reference pictures is selected for motion vector refinement by template matching the reference picture pointed to by MV0 for the sake of explanation. To perform motion vector refinement, a search region within the reference picture pointed to by the selected MV0 is defined around the candidate point pointed to by MV0. The candidate search space positions within the search region are analyzed by performing template matching with the current block space and determining a similarity measure. The candidate search space position having the lowest dissimilarity value is determined as the motion vector MV0". The motion vector MV1" is calculated based on MV1 and a function of MV0" - MV0 instead of performing a second template matching.

[0118] According to an embodiment of the present invention, the similarity comparison is performed by comparing the sample pointed to by MV0" and the sample pointed to by MV1". According to FIG. 4, any point pointed to by MV0" within the search space 310 has a corresponding motion vector MV1" given by one or more of Equations 1 to 8. To compare the similarity between sample blocks in the template matching process, the sample block pointed to by MV0" and the sample block pointed to by MV1" may be used in a function such as the SAD function. In this case, the template is composed only of samples belonging to the reference picture referred to by MV1, and the motion vector refinement operation is performed on MV0. Further, the template is slightly changed for each point within the search space 310.

[0119] More specifically, according to the embodiment, the following steps are applied.

[0120] Step 1: The similarity between the sample pointed to by the input motion vector MV0 and the sample pointed to by the input motion vector MV1 is calculated. The input motion vector MV0 points to point 330 within the search space.

[0121] Step 2: A second point within the search space 310 different from point 330 is selected. The second point is indicated by MV0".

[0122] Step 3: The motion vector MV1" is calculated using one or more of Equations 1 - 8 based on MV0", MV0, and MV1.

[0123] Step 4: The similarity between the sample pointed to by the input motion vector MV0" and the sample pointed to by the input motion vector MV1" is calculated. If the similarity is higher than the value calculated in Step 1, the pair of MV0" and MV1" is selected as the refined motion vector. Otherwise, the pair of MV0 and MV1 (i.e., the pair of the initial motion vectors) is selected as the refined motion vector.

[0124] Steps 2, 3, and 4 can be repeated to evaluate more candidate points within the search space 310. If there are no search points remaining within the search space, then the refined motion vector is output as the final refined motion vector.

[0125] In Step 4, according to one or more of Equations 1 - 8, it is clear that the similarity metric (e.g., SAD, SSIM, etc.) can result in the highest similarity at the initial point 330 pointed to by the initial motion vector MV0 whose counterpart is the initial motion vector MV1. In this case, the refined motion vector, and thus the output of the motion vector refinement process, is considered to be MV0 and MV1.

[0126] Any pair of motion vectors MV0" and MV1", which is the output of the apparatus 500, must comply with the rules described in one or more of Equations 1 to 8. The specific details of the structure of the template and the similarity metric used in the template matching operation may vary without affecting the invention and its advantage of reducing the search points to be checked by the pairing of two motion vectors.

[0127] According to an embodiment of the present invention, in addition to the dual prediction process performed in the inter prediction described above, other processing steps for encoding and decoding comply with the standard H.265 / HEVC.

[0128] However, in general, the present invention is applicable to any video decoder for decoding a plurality of pictures from a bitstream. Such a decoder may then include an inter prediction unit including an apparatus according to any one of claims 1 to 11, and a prediction unit for determining a prediction block according to a portion of a first reference picture referred to by a first motion vector and a portion of a second reference picture referred to by a second motion vector. The decoder may further include a bitstream parser. The bitstream parser may be implemented, for example, as a part of the entropy decoding unit 203 and configured to obtain an estimated value of the first motion vector and an estimated value of the second motion vector from the bitstream. The video decoder may further include a reconstruction unit 211 configured to reconstruct a current block according to the prediction block.

[0129] On the one hand, a video encoder for encoding a plurality of pictures into a bitstream includes an inter prediction unit including an apparatus according to any one of claims 1 to 12, and a prediction unit for determining a prediction block according to a portion of a first reference picture referred to by a first motion vector and a portion of a second reference picture referred to by a second motion vector, a bitstream formatter implemented as part of an entropy encoding unit 103 and configured to include an estimated value of the first motion vector and an estimated value of the second motion vector in the bitstream, and a reconstruction unit 111 configured to reconstruct a current block according to the prediction block and store the reconstructed block in a memory.

[0130] The inter prediction decoding process described above is not limited to the use of two reference pictures. Alternatively, three or more reference pictures and related motion vectors may be considered. In this case, the reference picture selection unit selects three or more reference pictures, and one of the reference pictures is used for motion vector refinement. The selection of the reference picture used for motion vector refinement uses one of the methods described in equations 1 to 5 discussed above. The remaining motion vectors are adjusted using the estimated value of each motion vector and the motion vector resulting from motion vector refinement. In other words, the present invention described above may also function when multi-reference prediction is performed. For example, if there are three reference pictures and three respective motion vectors, one of the three motion vectors may be determined by refinement by template matching and the other two may be calculated. This provides a reduction in complexity. Alternatively, two of the motion vectors are determined by refinement by template matching and one is calculated based on one or both of the refined motion vectors. As will be apparent to those skilled in the art, the present invention is extensible to any number of reference pictures and corresponding motion vectors used to construct a predictor for a current block.

[0131] The invention has the effect of enabling double prediction to be performed in a decoder with reduced processing load and memory requirements. It can be applied in any decoder and may be included within a coding device and / or a decoding device, i.e., on the encoder side or the decoder side.

[0132] The motion vector refinement described above can be implemented as part of the encoding and / or decoding of a video signal (video). However, motion vector refinement may also be used for other purposes in image processing, such as motion detection, motion analysis, etc.

[0133] Motion vector refinement may be implemented as a device. Such a device may be a combination of software and hardware. For example, motion vector refinement may be executed by a chip such as a general-purpose processor, or a digital signal processor (DSP), or a field programmable gate array (FPGA). However, the present invention is not limited to an implementation form in programmable hardware. It may be implemented on a specific application integrated circuit (ASIC) or by a combination of the above-described hardware components.

[0134] Motion vector refinement may also be implemented by program instructions stored on a computer-readable medium. When the program is executed, it causes the computer to perform steps of obtaining an estimated value of a motion vector, determining a first reference picture and a second reference picture based on the estimated value, performing motion vector refinement of the first motion vector, and calculating a second motion vector based on the estimated value of the motion vector and the refined first motion vector. The computer-readable medium can be any medium on which a program is stored, such as a DVD, CD, USB (flash) drive, hard disk, server storage available via a network, etc.

[0135] The encoder / decoder may be implemented in various devices including a TV set, a set-top box, a PC, a tablet, a smartphone, etc. It may be software, an app that executes method steps.

Explanation of Signs

[0136] 100 Encoder, Video Encoder 101 Conversion Unit 102 Quantization Unit 103 Entropy Encoding Unit, Entropy Coding Unit 104 Inverse Quantization Unit 105 Inverse Conversion Unit 106 Loop Filtering Unit 107 Frame Buffer, Reference Picture Buffer 108 Mode Selection Unit 109 Intra Prediction Unit 110 Inter Prediction Unit 111 Reconstruction Unit 200 Video Decoder, Decoder 203 Entropy Decoding Unit 204 Inverse Quantization Unit 205 Inverse Conversion Unit 206 Loop Filtering Unit 207 Reference Picture Buffer 209 Intra Prediction Unit 210 Inter Prediction Unit, Reference Picture Selection Unit 211 Reconstruction Unit 310 Search Space, Search Region 320 Current Block, Template 330 Center Point, Point, Initial Point 500 Motion Vector Refinement Unit, Device 505 Input 510 Reference Picture Selection Unit 520 Switch 530 Motion Vector Refinement Unit, Motion Vector Calculation Unit, MV Refinement Unit 540 Motion vector calculation unit, motion vector refinement unit 550 Motion vector calculation unit 560 Motion vector calculation unit 570 Switch 580 Output 800 Inter prediction unit 810 Device for determining motion vector, device 820 Motion vector estimator, motion vector estimation unit 890 Prediction unit

Claims

1. An apparatus for determining a first motion vector (MV0") in a first reference picture of a video and a second motion vector (MV1") in a second reference picture of the video, wherein the first and second motion vectors are applied in inter prediction of a picture block of the video in a current picture, and the apparatus comprises: a motion vector refinement unit (530) configured to obtain an estimated value (MV0) of the first motion vector and determine the first motion vector by performing a search within a search space (310) specified based on the estimated value of the first motion vector; a motion vector calculation unit (550) configured to obtain an estimated value (MV1) of the second motion vector and calculate the second motion vector based on the estimated value of the second motion vector and the first motion vector; The apparatus comprising.

2. The apparatus according to claim 1, wherein the motion vector calculation unit (550, 560) is configured to calculate the second motion vector by adding to the estimated value of the second motion vector a difference between the first motion vector and the estimated value of the first motion vector, or a function of the difference between the first motion vector and the estimated value of the first motion vector.

3. The apparatus according to claim 2, wherein the function includes scaling by a scaling factor and / or clipping.

4. The apparatus according to claim 3, wherein the value of the scaling factor depends on a ratio between respective distances of the first reference picture and the second reference picture with respect to the current picture.

5. The apparatus according to any one of claims 1 to 4, further comprising a reference picture selection unit (510) configured to obtain two reference pictures, select the first reference picture from the two reference pictures, and select the second reference picture from the two reference pictures.

6. The reference picture selection unit (510) is configured to select either the first or the second picture based on whether each of the first or second pictures is referenced in a bitstream including the coded picture blocks of the video by an index related to a predefined list of reference pictures among at least two lists of reference pictures, wherein the list of reference pictures associates an index with the position of a reference picture relative to the current picture, the apparatus according to claim 5. **Claim 7** The apparatus according to claim 6, wherein the reference picture selection unit (510) is configured to select, as the first reference picture, the picture having the highest position within the list of reference pictures when the two reference pictures are referenced in the bitstream by indices within the same predefined list of reference pictures. **Claim 8** The apparatus according to claim 5, wherein the reference picture selection unit (510) is configured to select, as the first reference picture, the picture having the lowest temporal layer among the two pictures. **Claim 9** The apparatus according to claim 5, wherein the reference picture selection unit (510) is configured to select, as the first reference picture, the picture having the lowest base quantization value. **Claim 10** The apparatus according to claim 5, wherein the reference picture selection unit (510) is configured to select, as the first reference picture, the picture having the minimum distance to the current picture. **Claim 11** The apparatus according to claim 5, wherein the reference picture selection unit (510) is configured to select the first reference picture and the second reference picture such that the estimated value of the first motion vector is smaller than the estimated value of the second motion vector. **Claim 12** The apparatus according to any one of claims 1 to 11, further comprising a motion vector estimator (820) configured to determine the estimated value of the first motion vector and the estimated value of the second motion vector by selecting them from a set of motion vector candidates based on the similarity between a template and the portions of the pictures referenced by the respective motion vector candidates. **Claim 13** A video decoder (200) for decoding a plurality of pictures from a bitstream, An inter prediction unit (210) including the apparatus according to any one of claims 1 to 12, and a prediction unit configured to determine a prediction block according to a portion of the first reference picture referred to by the first motion vector and a portion of the second reference picture referred to by the second motion vector; A bitstream parser (203) configured to obtain the estimated value of the first motion vector and the estimated value of the second motion vector from the bitstream; A reconstruction unit (211) configured to reconstruct the current block according to the prediction block A video decoder comprising.

14. A video encoder (100) for encoding a plurality of pictures into a bitstream, An inter prediction unit (110) including the apparatus according to any one of claims 1 to 12, and a prediction unit configured to determine a prediction block according to a portion of the first reference picture referred to by the first motion vector and a portion of the second reference picture referred to by the second motion vector; A bitstream formatter (103) configured to include the estimated value of the first motion vector and the estimated value of the second motion vector in the bitstream; A reconstruction unit (111) configured to reconstruct the current block according to the prediction block and store the reconstructed block in a memory A video encoder comprising.

15. A method (700) for determining a first motion vector (MV0") in a first reference picture of a video and a second motion vector (MV1") in a second reference picture of the video, wherein the first and second motion vectors are applied in inter prediction of a picture block in a current picture of the video, and the method includes: Step (701) of obtaining an estimated value (MV0) of the first motion vector; Step (703) of determining the first motion vector by performing a search within a search space (310) specified based on the estimated value of the first motion vector; Step (705) of obtaining an estimated value (MV1) of the second motion vector; Based on the estimated value (MV1) of the second motion vector, calculating the second motion vector (MV1") based on the first motion vector (MV0") (step 707); A method comprising the above.

Citation Information

Patent Citations

  • Techniques for motion estimation

    JP2011147130A

  • Image encoder, image encoding method and image encoding program

    JP2013016934A

  • Motion vector derivation in video coding

    US20160286229A1