Method and apparatus for coding video data by sub-pixel motion vector refinement
By employing a decoder-side motion vector derivation method in video encoding, and utilizing multi-step search and high-resolution sub-pixel position evaluation of motion vectors, the problem of insufficient motion vector prediction accuracy is solved, thereby improving encoding efficiency and encoding quality.
Patent Information
- Application Number
- CN202310685170.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-04
- Filing Date
- 2018-06-27
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2038-06-27
AI Technical Summary
Existing video coding techniques suffer from low coding efficiency in motion vector prediction, especially in FRUC merging mode, where insufficient accuracy of motion vectors leads to low coding efficiency.
By employing a decoder-side motion vector derivation method, a multi-step search process is used to evaluate motion vectors at a resolution of 1/16 sub-pixels or higher. Rhombus and cross-shaped patterns are used to refine the motion vectors at sub-pixel locations. By combining the signal characteristics and matching type within the template, the accuracy of the motion vectors is improved.
It improves the coding efficiency of video encoding, especially in FRUC merging mode, by enhancing the accuracy of motion vectors through higher resolution subpixel position evaluation and pattern selection, thereby improving coding quality and compression performance.
Smart Images

Figure CN116506639B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application 201880037185.X entitled "Method and apparatus for encoding or decoding video data by sub-pixel motion vector refinement" having an application date of 27 June 2018. TECHNICAL FIELD
[0002] The present disclosure relates to a method and apparatus for encoding or decoding video data. More particularly, it relates to encoding according to a specific encoding mode using a decoder-side motion vector derivation mode, referred to as Frame-Rate Up-Conversion mode or FRUC mode. BACKGROUND
[0003] Predictive encoding of video data is based on a frame partitioning into a plurality of pixel blocks. For each pixel block, a prediction block is searched in the available data. The prediction block can be a block in a reference frame different from the current frame in INTER encoding mode, or can be generated from neighboring pixels in the current frame in INTRA encoding mode. Different encoding modes are defined according to different ways of determining the prediction block. The result of the encoding is an indication of the prediction block and of a residual block, which is the difference between the block to be encoded and the prediction block.
[0004] Regarding INTER encoding mode, the indication of the prediction block is a motion vector giving the position of the prediction block in the reference image, relative to the position of the block to be encoded. The motion vector itself is predictively encoded based on a motion vector predictor. The HEVC (High Efficiency Video Coding) standard defines several known encoding modes for the predictive encoding of the motion vector, namely the AMVP (Advanced Motion Vector Prediction) mode, the Merge derivation process. These modes are based on the construction of a candidate list of motion vector predictors and the signaling of an index of the motion vector predictor to be used for the encoding in this list. Usually, a residual motion vector is also signaled.
[0005] Recently, a new encoding mode called FRUC has been introduced regarding the motion vector prediction, which defines a decoder-side derivation process of the motion vector predictor that is not signaled at all. The result of the derivation process will be used as the motion vector predictor without any index or residual motion vector to be transmitted by the decoder.
[0006] In FRUC Merge mode, the derivation process includes a refinement step to improve the precision of the obtained motion vector at the sub-pixel level. This process involves the evaluation of different sub-pixel positions around the obtained motion vector according to different patterns. SUMMARY
[0007] The invention has been conceived to improve the known refinement step. It aims at improving the encoding efficiency by taking into account the characteristics of the signal and / or the matching type inside the template.
[0008] According to a first aspect of the application, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of blocks of pixels, the method comprising, for a block of pixels:
[0009] - deriving a list of motion vectors for a motion vector using a mode that obtains motion information by a decoder-side motion vector derivation method;
[0010] - evaluating the list of motion vectors to select one motion vector;
[0011] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector; wherein
[0012] - at least some of the sub-pixel positions are selected at 1 / 16 sub-pixel or higher resolution.
[0013] In an embodiment, refining the selected motion vector comprises a plurality of search steps; and wherein at least one of the search steps involves sub-pixel positions at 1 / 16 sub-pixel or higher resolution.
[0014] In an embodiment, the plurality of search steps comprises at least three consecutive steps; each of the three consecutive search steps involves sub-pixel positions at a given resolution; and wherein each given resolution associated with the last two search steps is greater than the given resolution of the preceding search step.
[0015] In an embodiment, the plurality of search steps comprises at least one search step based on a diamond pattern at a first sub-pixel resolution, and two search steps based on a cross pattern at a sub-pixel resolution greater than the first sub-pixel resolution.
[0016] In an embodiment, at least some of the search steps are performed at a sub-pixel resolution that depends on a matching type used for encoding the block of pixels.
[0017] In an embodiment, the first search step is a search step based on a diamond pattern at 1 / 4 sub-pixel resolution. The second search step is a search step based on a cross pattern at 1 / 8 sub-pixel resolution; the third search step is a search step based on a cross pattern at 1 / 16 sub-pixel resolution.
[0018] According to another aspect of the application, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of blocks of pixels, the method comprising, for a block of pixels:
[0019] - deriving a list of motion vectors for a motion vector using a mode that obtains motion information by a decoder-side motion vector derivation method;
[0020] - evaluating the list of motion vectors to select one motion vector;
[0021] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector; wherein
[0022] - the sub-pixel positions are evaluated according to at least one pattern selected from a plurality of patterns; and wherein:
[0023] - the plurality of patterns comprises a horizontal pattern and a vertical pattern.
[0024] In one embodiment, the plurality of patterns further comprises at least a diagonal pattern.
[0025] In one embodiment, the patterns of the plurality of patterns are selected based on an edge direction detected in neighboring pixel blocks.
[0026] In one embodiment, at least one pattern of the plurality of patterns is defined at a 1 / 16 sub-pixel resolution or higher.
[0027] According to another aspect of the application, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the method comprising, for one pixel block:
[0028] - deriving a list of motion vectors for a motion vector using a mode that obtains motion information by a decoder-side motion vector derivation method;
[0029] - evaluating the list of motion vectors to select one motion vector;
[0030] - determining to refine the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector based on signal content of at least one neighboring pixel block.
[0031] In one embodiment, the signal content is a frequency in the neighboring pixel block.
[0032] According to another aspect of the application, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the method comprising, for one pixel block:
[0033] - deriving a list of motion vectors for a motion vector using a mode that obtains motion information by a decoder-side motion vector derivation method;
[0034] - evaluating the list of motion vectors to select one motion vector; wherein:
[0035] - the derivation of the list of motion vectors for a motion vector is based on a template defined by a pattern of neighboring pixels of the pixel block.
[0036] - the template for derivation of a motion vector list for a motion vector is determined based on a content signal of the template.
[0037] In one embodiment, the signal content is a frequency in the template.
[0038] According to another aspect of the application, there is provided a computer program product for a programmable device, the computer program product comprising a sequence of instructions for implementing the method according to the application when loaded into and executed by the programmable device.
[0039] According to another aspect of the application, there is provided a computer readable storage medium storing instructions of a computer program for implementing the method according to the application.
[0040] According to another aspect of the application, there is provided a decoding device for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding device comprising a processor configured to decode one pixel block by:
[0041] - deriving a motion vector list for a motion vector using a mode that obtains motion information by a decoder-side motion vector derivation method;
[0042] - evaluating the motion vector list to select one motion vector;
[0043] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector; wherein
[0044] - at least some of the sub-pixel positions are selected with 1 / 16 sub-pixel or higher resolution.
[0045] According to another aspect of the application, there is provided a decoding device for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding device comprising a processor configured to decode one pixel block by:
[0046] - deriving a motion vector list for a motion vector using a mode that obtains motion information by a decoder-side motion vector derivation method;
[0047] - evaluating the motion vector list to select one motion vector;
[0048] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector; wherein
[0049] - the sub-pixel positions are evaluated according to at least one pattern selected from a plurality of patterns; and wherein:
[0050] - These multiple patterns include horizontal and vertical patterns.
[0051] According to another aspect of the present invention, a decoding apparatus is provided for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding apparatus comprising a processor configured to decode a pixel block by means of:
[0052] - A list of motion vectors is derived by using a pattern that obtains motion information through a decoder-side motion vector derivation method;
[0053] - Evaluate the list of motion vectors to select one;
[0054] - The selected motion vector is refined by evaluating the motion vector at sub-pixel locations near the selected motion vector based on the signal content of at least one neighboring pixel block.
[0055] According to another aspect of the present invention, a decoding apparatus is provided for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding apparatus comprising a processor configured to decode a pixel block by means of:
[0056] - A list of motion vectors is derived by using a pattern that obtains motion information through a decoder-side motion vector derivation method;
[0057] - Evaluate the list of motion vectors to select one motion vector; where:
[0058] - The derivation of the list of motion vectors is based on a template defined by the pattern of adjacent pixels of a pixel block;
[0059] - The template for deriving the list of motion vectors for motion vectors is determined based on the content signal of the template.
[0060] At least a portion of the method according to the invention can be implemented by a computer. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which are generally referred to herein as “circuit,” “module,” or “system.” Furthermore, the invention can take the form of a computer program product embodied in any tangible medium, expressed as having computer-usable program code embodied in that medium.
[0061] As the application can be implemented in software, the application can be embodied as computer readable code for provision to a programmable apparatus on any suitable carrier medium. A tangible, non-transitory carrier medium can include a storage medium such as a floppy disk, a CD-ROM, a hard disk drive, a magnetic tape device or a solid state memory device and the like. A transient carrier medium can include a signal such as an electrical signal, an electronic signal, an optical signal, a magnetic signal or an electromagnetic signal and the like. BRIEF DESCRIPTION OF DRAWINGS
[0062] Embodiments of the application will now be described, by way of example only, and with reference to the following drawings in which:
[0063] Figure 1 An HEVC encoder architecture is shown;
[0064] Figure 2 The principle of a decoder is shown;
[0065] Figure 3 Template matching and bilateral matching in FRUC merge mode is shown;
[0066] Figure 4 Decoding of FRUC merge information is shown;
[0067] Figure 5 Encoder evaluation of merge mode and merge FRUC mode is shown;
[0068] Figure 6 Merge FRUC mode derivation at coding unit and sub-coding unit level in JEM is shown;
[0069] Figure 7 Template around the current block for JEM template matching method is shown;
[0070] Figure 8 Motion vector refinement is shown;
[0071] Figure 9 Adaptive sub-pixel resolution for motion vector refinement in an embodiment of the application is shown;
[0072] Figure 10 An example of results obtained with prior art and one embodiment of the application is given;
[0073] Figure 11 Some motion vector refinement search shapes used in one embodiment of the application are shown;
[0074] Figure 12 Adaptive search shapes for motion vector refinement in one embodiment of the application are shown;
[0075] Figure 13Adaptive motion vector refinement of one embodiment of the present invention is shown;
[0076] Figure 14 Template selection for motion vector refinement of one embodiment of the present invention is shown;
[0077] Figure 15 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. DETAILED DESCRIPTION
[0078] Figure 1 An HEVC encoder architecture is shown. In a video encoder, the original sequence 101 is divided into blocks of pixels 102, called coding units. Then, a coding mode will affect each block. Two coding mode families are generally used in HEVC: a mode based on spatial prediction, or INTRA mode 103, and an INTER mode, or a mode based on temporal prediction, based on motion estimation 104 and motion compensation 105. An INTRA coding unit is generally predicted from coding pixels at its causal boundary, through a process called INTRA prediction.
[0079] Temporal prediction first consists in finding, in a previous or future frame called reference frame 116, a reference area that is the closest to the coding unit in a motion estimation step 104. This reference area constitutes a prediction block. Next, in a motion compensation step 105, the coding unit is predicted using the prediction block to compute a residual.
[0080] In both cases, spatial and temporal prediction, the residual is computed by subtracting the coding unit from the original prediction value.
[0081] In INTRA prediction, the prediction direction is coded. In temporal prediction, at least one motion vector is coded. However, to further reduce the bit rate cost related to motion vector coding, the motion vector is not directly coded. Indeed, assuming that the motion is uniform, it is particularly interesting to code the motion vector as the difference between this motion vector and its surrounding motion vectors. For example, in the H.264 / AVC coding standard, the motion vector is coded with respect to a median vector computed between 3 blocks above and left of the current block. Only the difference computed between the median vector and the current block motion vector, also called residual motion vector, is coded in the bitstream. This is processed in the module "Mv prediction and coding" 117. The value of each coded vector is stored in the motion vector field 118. The neighboring motion vectors used for prediction are extracted from the motion vector field 118.
[0082] Then, in module 106, the mode that optimizes the rate-distortion performance is selected. To further reduce redundancy, in module 107 a transform (typically DCT) is applied to the residual block and in module 108 quantization is applied to the coefficients. Then, in module 109 the quantized coefficient block is entropy coded and the result is inserted into the bitstream 110.
[0083] Then, the encoder performs a decoding of the encoded frame in modules 111 to 116 for future motion estimation. These steps allow the encoder and the decoder to have the same reference frame. To reconstruct the encoded frame, inverse quantization is performed in module 111 and inverse transform in module 112 for the residual in order to provide a "reconstructed" residual in the pixel domain. Depending on the coding mode (INTER or INTRA), this residual is added to the INTER prediction value 114 or to the INTRA prediction value 113.
[0084] Then, this first reconstruction is filtered in module 115 by one or several post- filters. These post-filters are integrated in the encoding and decoding loop. This means that they need to be applied on the reconstructed frame on the encoder and decoder side in order to use the same reference frame on the encoder and decoder side. The purpose of this post-filtering is to remove compression artifacts.
[0085] The principle of the decoder has been represented in Figure 2 The video stream 201 is first entropy decoded in module 202. Then, the residual data is inverse quantized in module 203 and inverse transformed in module 204 to obtain pixel values. The mode data is also entropy decoded depending on the mode that performs an INTRA type decoding or an INTER type decoding. In the case of INTRA mode, the INTRA prediction value is determined according to the INTRA prediction mode 205 specified in the bitstream. If the mode is INTER, the motion information is extracted from the bitstream 202. It consists of a reference frame index and a motion vector residual. The motion vector prediction value is added to the motion vector residual to obtain the motion vector 210. Then, the motion vector is used to locate the reference region 206 in the reference frame. Note that the motion vector field data 211 is updated with the decoded motion vector in order to be used for the prediction of the next decoded motion vector. Then, the first reconstruction of the decoded frame is post-filtered (207) using the exact same post-filters as the ones used on the encoder side. The output of the decoder is the decompressed video 209.
[0086] The HEVC standard uses three different INTER modes: Inter mode, Merge mode and Merge Skip mode. The main difference between these modes is the data signalling in the bitstream. For motion vector coding, the current HEVC standard includes a competitive based motion vector prediction scheme compared to its predecessor. This means that at the encoder side several candidates compete according to a rate-distortion criterion in order to find the best motion vector predictor or the best motion information for Inter or Merge mode respectively. The index corresponding to the best candidate or the best predictor of the motion information is inserted into the bitstream. The decoder can derive the same set of predictors or candidates and use the best one according to the decoded index.
[0087] The design of the derivation of the predictors and candidates is very important to achieve the best coding efficiency without significant impact on the complexity. In HEVC, two kinds of motion vector derivation are used: one for Inter mode (Advanced Motion Vector Prediction (AMVP)) and the other one for Merge mode (Merge derivation process).
[0088] As already mentioned, the candidates of the Merge mode ("classic" or "skip") represent all the motion information: direction, list and reference frame index as well as the motion vector. Several candidates are generated by the Merge derivation process described below, each of them having an index. In the current HEVC design, the maximum number of candidates for both Merge modes is equal to 5.
[0089] For the current version of JEM, two types of search can be performed: template matching and bilateral matching. Figure 3 Both approaches are illustrated. The principle of the bilateral matching 301 is to find the best match between two blocks (sometimes also called templates by analogy with the template matching described below) along the motion trajectory of the current coding unit.
[0090] The principle of the template matching 302 is to derive the motion information of the current coding unit by computing the matching cost between the reconstructed pixels around the current block and the neighboring pixels around the block pointed by the evaluated motion vector. This template corresponds to a pattern of neighboring pixels around the current block and to a corresponding pattern of neighboring pixels around the predicted block.
[0091] For both matching types (template matching or bilateral matching), the different matching costs computed are compared in order to find the best matching cost. The motion vector or the motion vector pair that obtained the best match is chosen as the derived motion information. More details can be found in JVET-F1001.
[0092] Both matching methods provide the possibility to derive all motion information, motion vectors, reference frames and prediction types. Motion information derivation at the decoder side (denoted as "FRUC" in JEM) is applied to all HEVC inter modes: AMVP, Merge and Merge Skip.
[0093] For AMVP, all motion information is signaled: uni-prediction or bi-prediction, reference frame index, predictor index motion vector and residual motion vector, the FRUC method is applied to determine a new predictor which is set as the first predictor in the predictor list. Its index is therefore 0.
[0094] For Merge and Merge Skip modes, a FRUC flag is signaled for a CU. When the FRUC flag is false, the merge index is signaled and the regular merge mode is used. When the FRUC flag is true, an additional FRUC mode flag is signaled to indicate which method (bilateral matching or template matching) will be used to derive the motion information of the block. Note that bilateral matching is only applied to B-frames and not to P-frames.
[0095] For Merge and Merge Skip modes, a motion vector field is defined for the current block. This means that for sub-coding units smaller than the current coding unit, a vector is defined. In addition, with respect to classical merge, one motion vector for each list can form the motion information of a block.
[0096] Figure 4 is a flow chart illustrating this signaling of the FRUC flag for the merge mode of a block. According to HEVC terminology, a block can be a coding unit or a prediction unit.
[0097] In a first step 401, the skip flag is decoded to know if the coding unit is coded according to the skip mode. If the flag is false in a test in step 402, then the merge flag is decoded in step 403 and tested in step 405. When the coding unit is coded according to the skip or merge mode, the merge FRUC flag is decoded in step 404. When the coding unit is not coded according to the skip or merge mode, the intra prediction information of the classical AMVP inter mode is decoded in step 406. When the FRUC flag of the current coding unit is true in a test in step 407, and if the current slice is a B slice, the matching mode flag is decoded in step 408. Note that bilateral matching in FRUC is only applied to B slices. If the slice is not a B slice and FRUC is selected, the mode must be template matching and there is no matching mode flag. If the coding unit is not FRUC, then the classical merge index is decoded in step 409.
[0098] The FRUC merge mode is evaluated at the encoder side in competition with the classic merge mode (and other possible merge modes). Figure 5 The current coding mode evaluation method in JEM is explained. First, the classic merge mode of HEVC is evaluated in step 501. The candidate list is first evaluated in step 502 with a simple SAD (sum of absolute difference) between the original block and each candidate of the candidate list. Then, the real rate-distortion (RD) cost of each candidate in the list of constrained candidates is evaluated, as shown by steps 504 to 508. In the evaluation, the rate-distortion with residual (step 505) and without residual (step 506) are evaluated. Finally, the best merge candidate is determined in step 509, which can have or not have residual.
[0099] Then, the FRUC merge mode is evaluated in steps 510 to 516. For each matching method of step 510 (i.e. bilateral and template matching), the motion vector field of the current block is retrieved in step 511 and the full rate-distortion cost evaluation with or without residual is computed in steps 512 and 513. Based on these rate-distortion costs, the best motion vector with or without residual is determined in step 515. Finally, the best mode between the classic merge mode and the FRUC merge mode is determined in step 517 before the possible evaluation of other modes.
[0100] Figure 6 The FRUC merge evaluation method at the encoder side is explained. For each matching type (step 601), i.e. template matching type and bilateral matching type, the coding unit level is first evaluated by module 61 and then the sub-coding unit level is evaluated by module 62. The aim is to find the motion information for each sub-coding unit in the current coding unit 603.
[0101] Module 61 handles the coding unit level evaluation. In step 611, a list of motion information is derived. For each motion information in this list, the distortion cost is computed in step 612 and compared to each other. The best motion vector for the template or the best motion vector for the bilateral pair 613 is the one that minimizes the cost. Then, a motion vector refinement step 614 is applied to improve the accuracy of the obtained motion vector. With the FRUC method, for template matching estimation, a bilinear interpolation is used instead of the classic Discrete Cosine Transform Interpolation Filter (DCTIF) interpolation filter. This reduces the memory access around the block to only one pixel instead of 7 pixels around the block for the traditional DCTIF. Indeed, the bilinear interpolation filter only needs 2 pixels to obtain a sub-pixel value in one direction.
[0102] After the motion vector refinement, a better motion vector of the current coding unit is obtained in step 615. This motion vector will be used for the sub-coding unit level assessment.
[0103] In step 602, the current coding unit is subdivided into several sub-coding units. A sub-coding unit is a square block depending on the partition depth of the coding unit in the quad-tree structure. The minimum size is 4x4.
[0104] For each sub-CU, the sub-CU level assessment module 62 assesses the best motion vector. In step 621, a list of motion vectors is derived, including the best motion vector obtained at the CU level in step 615. For each motion vector, a distortion cost is assessed in step 622. However, this cost also includes a cost representative of the distance between the best motion vector obtained at the coding unit level to avoid diverging motion vector field and the current motion vector. The best motion vector is obtained 623 based on the minimum cost. This vector 623 is then refined by a MV refinement process 624 in the same way as in step 614 at the CU level.
[0105] At the end of the process, the motion information of each sub-CU is obtained for one matching type. At the encoder side, the best RD cost between the two matching types is compared to select the best matching type. At the decoder side, this information is decoded from the bitstream (in step 408 of the method of Figure 4 ).
[0106] For the template FRUC matching mode, the templates 702, 703 include 4 lines above the block and 4 lines to the left of the block 701 used to estimate the rate-distortion cost, as shown in grey in Figure 7 Two different templates are used, namely the left template 702 and the top template 703. The current block or matching block 701 is not used to determine the distortion.
[0107] Figure 8 The motion vector refinement steps 614 and 624 in the method of Figure 6 are shown with an additional search around the identified best predictor (613 or 623).
[0108] The method takes as input the best motion vector predictor 801 identified (612 or 622) in the list.
[0109] In step 802, a diamond search is applied at a resolution corresponding to ¼ pixel positions. This diamond search is based on a diamond pattern as shown in diagram 81, centered on the best motion vector at ¼ pixel resolution. This step results in a new best motion vector at ¼ pixel resolution 803.
[0110] In step 804, the optimal motion vector position 803 obtained by the diamond search becomes the center of the cross search (referring to a cross-shaped pattern) at 1 / 4 pixel resolution. This cross search pattern, illustrated in Figure 82, is at 1 / 4 pixel resolution, centered on the optimal motion vector 803. This step yields a new optimal motion vector 805 at 1 / 4 pixel resolution.
[0111] In step 806, the new optimal motion vector position 805 obtained through search step 804 becomes the center of the crosshair search at 1 / 8 pixel resolution. This step yields a new optimal motion vector 807 at 1 / 8 pixel resolution. Figure 83 illustrates the pattern of these three search steps at 1 / 8 resolution, with all positions tested.
[0112] The present invention has been conceived to improve upon known refinement steps. It aims to improve coding efficiency by taking into account the characteristics of signals and / or matching types within the template.
[0113] In one embodiment of the invention, the thinning accuracy of sub-pixels is improved in each step of the thinning method.
[0114] The refinement process involves evaluating the motion vectors at sub-pixel locations. Sub-pixel locations are determined at a given sub-pixel resolution. The sub-pixel resolution determines the number of sub-pixel locations between two pixels. Higher resolutions result in a greater number of sub-pixel locations between two pixels. For example, a 1 / 8 pixel resolution corresponds to 8 sub-pixel locations between two pixels.
[0115] Figure 9 This embodiment is illustrated. In the first illustration 91, it is shown at a resolution of 1 / 16 pixels. Figure 8 The prior art pattern described herein. As illustrated in the example in Figure 92, the first step maintains a diamond search of 1 / 4 pixel, followed by a cross search of 1 / 8 pixel, and a final cross search of 1 / 16 pixel. This embodiment provides the possibility of obtaining a more accurate position around a good location. Furthermore, the motion vector, on average, is closer to the initial position than with the previous sub-constraints.
[0116] In one additional embodiment, this refinement applies only to template matching and not to bilateral matching.
[0117] Figure 10is a flowchart representing these embodiments. The best motion vector 1001 corresponding to the best motion vector obtained by the previous steps (613 or 623) is set to the center position of the diamond search 1002 at the 1 / 4 pel position. The diamond search 1002 results in a new best motion vector 1003. In step 1004 the match type is tested. If the match type is template matching, the motion vector 1003 obtained by the diamond search is the center of a cross search 1009 at 1 / 8 pel accuracy, resulting in a new best motion vector 1010. A cross search 1011 at 1 / 16 pel accuracy is performed on this new best motion vector 1010 to obtain the final best motion vector 1012.
[0118] If the match type is not template matching, a regular cross search at 1 / 4 pel is done in step 1005 to obtain a new best motion vector 1006, followed by a 1 / 8 pel search step 1007 to obtain the final best motion vector 1008.
[0119] This embodiment improves coding efficiency in particular for single prediction. For double prediction, the average between the two similar blocks is sometimes similar to an increase in sub-pel resolution. For example, if both blocks used for double prediction come from the same reference frame and the difference between their two motion vectors is equal to a lower sub-pel resolution, then in this case double prediction corresponds to an increase in sub-pel resolution. For single prediction, there is no such additional average between the two blocks. Therefore, it is more important to increase the sub-pel resolution for single prediction, in particular when this higher resolution does not need to be signaled in the bitstream.
[0120] In one embodiment, the diamond or cross search pattern is replaced by a horizontal, diagonal or vertical pattern. Figure 11 Various diagrams that can be used in this embodiment are shown. Diagram 1101 shows the positions searched according to a 1 / 8 pel horizontal pattern. Diagram 1102 shows the positions searched according to a 1 / 8 pel vertical pattern. Diagram 1103 shows the positions searched according to a 1 / 8 pel diagonal pattern. Diagram 1104 shows the positions searched according to another 1 / 8 pel diagonal pattern. Diagram 1105 shows the positions searched according to a 1 / 16 pel horizontal pattern.
[0121] A horizontal pattern is a pattern of sub-pel positions aligned horizontally. A vertical pattern is a pattern of sub-pel positions aligned vertically. A diagonal pattern is a pattern of sub-pel positions aligned along a diagonal.
[0122] The advantage of these patterns is mainly interesting when the block contains an edge, as it provides a better refinement of this edge for the prediction. Indeed, in classical motion vector estimation, the refinement of the motion vector gives higher results when additional test positions are chosen in the perpendicular axis of the detected edge.
[0123] Figure 12 An example of a method to select a pattern according to this embodiment is given. This flowchart can be used to change one or more pattern search in modules 802, 804, 806 in Figure 8 When it is tested in step 1201 that the matching type is a template matching, the left template of the current block is extracted in step 1202. It is then determined whether the block contains an edge. If so, the direction of the edge is determined in step 1203. For example, the gradient on the current block can be computed to determine the presence of an edge and its direction. If it is tested in step 1204 that the "directionLeft" is horizontal, then a vertical pattern is selected in step 1205, for example pattern 1102, for the motion vector refinement.
[0124] If there is no edge in the left template, or if the identified direction is not horizontal, then the upper template of the current block is extracted in step 1206. It is determined whether the template contains an edge, and if so, the direction is determined in step 1207. If it is tested in step 1208 that the "directionUp" is vertical, then the pattern selected for the motion vector refinement in step 1209 is a horizontal pattern, for example pattern 1101 or 1105. Otherwise, the initial pattern (81, 82, 83) is selected in step 1210.
[0125] Note that if the left block contains a horizontal edge, then this edge should pass through the current block. Likewise, if the upper block contains a vertical edge, then this edge should pass through the current block. Conversely, if the left block contains a vertical edge, then this edge should not pass through the current block, and so on.
[0126] If it is determined in step 1201 that the matching type is a bilateral matching, then one template block is selected in step 1211. It is then determined in step 1212 whether the template contains a direction and which direction it is. If it is tested in step 1213 that the direction is horizontal, then the vertical pattern is selected in step 1205. If it is tested in step 1214 that the direction is vertical, then the horizontal pattern is selected in step 1209. For bilateral matching, some other linear patterns can be used as diagonal patterns 1103 and 1104. In that case, the linear pattern is selected that is most perpendicular to the direction of the edge determined in 1212. If no edge is detected in step 1212, then the initial pattern is kept.
[0127] This change of pattern for motion vector refinement has the advantage of an increase in coding efficiency, as the pattern is adapted to the signal contained inside the template.
[0128] The pattern can also be changed. For example, diagram 1105 shows a horizontal pattern at 1 / 16 pixel precision, where more positions have been concentrated around the initial motion vector position with high motion vector precision.
[0129] In one embodiment, the signal content of the template is used to determine the number of positions tested for motion vector refinement. For example, the presence of high frequencies in the template around the current block is used to determine the application of the refinement step. The reason is that without sufficiently high frequencies, the sub-pixel refinement of the motion vector is not relevant.
[0130] Figure 13 An example of this embodiment is given.
[0131] If in step 1301 it is tested that the matching pattern is a template match, then in step 1302 the left template is extracted, typically the 4 lines adjacent to the current block. Then in step 1303 it is determined whether the block contains high frequencies. This determination can be obtained for example by comparing the sum of gradients to a threshold. If in step 1304 it is tested that the left template contains high frequencies, then the motion vector refinement step 1305 is applied, which corresponds to Figure 6 steps 614 and 624 of figure 6.
[0132] If the left template does not contain sufficiently high frequencies, then in step 1306 the upper template is extracted. In step 1307 it is determined whether the extracted template contains high frequencies. If in step 1308 it is tested that this is the case, then in step 1305 the motion vector refinement step is applied. Otherwise, in step 1309 the motion vector refinement step is skipped.
[0133] If in step 1301 it is determined that the matching type is a bilateral match, then in step 1310 one template is extracted. Then in step 1314 it is determined whether the extracted template contains high frequencies. If in step 1313 it is tested that the extracted template does not contain high frequencies, then in step 1309 the motion vector refinement is not applied, otherwise in step 1305 the motion vector refinement is applied.
[0134] In yet another embodiment, the signal content of the template is used to determine the template to be used for the distortion estimation, which is used to determine the motion vector for the template matching type FRUC evaluation.
[0135] Figure 14 This embodiment is illustrated.
[0136] The process starts in step 1401. The left and top templates are extracted in steps 1402 and 1403. For each extracted template, it is determined in steps 1404 and 1405 if the extracted template contains high frequencies. If the tested template contains high frequencies in steps 1406 and 1407, they are used in the distortion estimation in the FRUC evaluation in steps 1408 and 1409. When both templates contain high frequencies, they are both used in the FRUC evaluation. When no template contains high frequencies, the top template is selected in step 1411 for the distortion estimation.
[0137] All these embodiments can be combined.
[0138] Figure 15 is a schematic block diagram of a computing device 1500 for implementing one or more embodiments of the application. The computing device 1500 can be a device such as a microcomputer, a workstation or a light portable device. The computing device 1500 comprises a communication bus connected to:
[0139] - a central processing unit 1501, for example a microprocessor, denoted CPU;
[0140] - a random access memory 1502, denoted RAM, for storing the executable code of the method of the embodiments of the application, as well as registers suitable for recording the variables and parameters necessary for implementing the method for encoding or decoding at least one portion of an image according to an embodiment of the application, the storage capacity of which can be extended by optional RAMs connected for example to expansion ports;
[0141] - a read-only memory 1503, denoted ROM, for storing the computer program for implementing the embodiments of the application;
[0142] - a network interface 1504, generally connected to a communication network through which digital data to be processed are transmitted or received. The network interface 1504 can be a single network interface or consist of a set of different network interfaces (for example wired and wireless interfaces, or different kinds of wired or wireless interfaces). Under the control of a software application running in the CPU 1501, data packets are written to the network interface for transmission or read from the network interface for reception.
[0143] - a user interface 1505 can be used to receive input from a user or to display information to the user;
[0144] - a hard disk 1506, denoted HD, can be provided as a mass storage device;
[0145] - an I / O module 1507 can be used to receive data from / to external devices such as a video source or a display.
[0146] The executable code can be stored in the read-only memory 1503, on the hard disk 1506 or on a removable digital medium such as a disk. According to a variant, the executable code of a program can be received via the network interface 1504 over a communication network, in order to be stored in one of the storage means of the communication device 1500, for example in the hard disk 1506, before being executed.
[0147] The central processing unit 1501 is adapted to control and direct the execution of the instructions or portions of software code of one or more programs according to an embodiment of the application, these instructions being stored in one of the storage means mentioned above. After power-up, the CPU 1501 is able to execute the instructions relating to a software application, for example from the main RAM memory 1502, after having loaded them from the program ROM 1503 or the hard disk (HD) 1506. Such a software application, when executed by the CPU 1501, causes the steps of the flowchart of the application to be performed.
[0148] Any step of the algorithm of the application can be implemented in software by a programmable computer such as a PC ("Personal Computer"), a DSP ("Digital Signal Processor") or a microcontroller executing a set of instructions or a program; or in hardware by a machine or an application-specific component such as an FPGA ("Field-Programmable Gate Array") or an ASIC ("Application-Specific Integrated Circuit").
[0149] Although the application has been described above with reference to particular embodiments, the application is not limited to particular embodiments and modifications within the scope of the application will be apparent to those skilled in the art.
[0150] Many further modifications and variations will be apparent to those skilled in the art in reference to the foregoing description, the foregoing description being given by way of example only and without intending as a limitation on the scope of the application, the scope of the application being given by the appended claims only. In particular, different features from different embodiments can be interchanged, where appropriate.
[0151] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
Claims
1. A method of decoding video data, the video data comprising frames, each frame being partitioned into blocks, the method of decoding comprising: determining a first motion vector and a second motion vector for a block to be decoded in a frame to be decoded, the first motion vector specifying a first region in a first frame different from the frame to be decoded and the second motion vector specifying a second region in a second frame different from the frame to be decoded; in the case where a mode is used in the method of decoding in which the first motion vector and the second motion vector can be refined, determining whether to refine the first motion vector and the second motion vector in the method of decoding by using at least a sample value of the first region and a sample value of the second region; wherein, in the case where it is determined to refine, refining the first motion vector and the second motion vector; and wherein, in the case where the refining is performed, decoding the block to be decoded using inter prediction using the refined first motion vector and the refined second motion vector, wherein a block of pixels is decoded using an inverse transform; wherein a plurality of patterns for refining are available for use in the mode in the method of decoding; wherein the plurality of patterns comprises a horizontal pattern and a vertical pattern, wherein, when the horizontal pattern is used, a plurality of horizontal sub-pixel positions to the right of the position to be refined and a plurality of horizontal sub-pixel positions to the left of the position to be refined are candidates for a horizontal position of the refined motion vector, and wherein, when the vertical pattern is used, a plurality of vertical sub-pixel positions above the position to be refined and a plurality of vertical sub-pixel positions below the position to be refined are candidates for a vertical position of the refined motion vector.
2. The method of claim 1, wherein, The refined first motion vector and the refined second motion vector are 1 / 16 sub-pixel precision.
3. A decoding apparatus for video data, the video data comprising frames, each frame being partitioned into blocks, the decoding apparatus comprising one or more processors configured to decode a block of pixels by: determining a first motion vector and a second motion vector for a block to be decoded in a frame to be decoded, the first motion vector specifying a first region in a first frame different from the frame to be decoded and the second motion vector specifying a second region in a second frame different from the frame to be decoded; in the case where a mode is used in the decoding apparatus in which the first motion vector and the second motion vector can be refined, determining whether to refine the first motion vector and the second motion vector in the decoding apparatus by using at least a sample value of the first region and a sample value of the second region; wherein, in the case where it is determined to refine, refining the first motion vector and the second motion vector; and wherein, in the case where the refining is performed, decoding the block to be decoded using inter prediction using the refined first motion vector and the refined second motion vector, wherein a block of pixels is decoded using an inverse transform; wherein a plurality of patterns for refining are available for use in the mode by the decoding apparatus; wherein the plurality of patterns comprises a horizontal pattern and a vertical pattern, wherein, when the horizontal pattern is used, a plurality of horizontal sub-pixel positions to the right of the position to be refined and a plurality of horizontal sub-pixel positions to the left of the position to be refined are candidates for a horizontal position of the refined motion vector, and wherein, when the vertical pattern is used, a plurality of vertical sub-pixel positions above the position to be refined and a plurality of vertical sub-pixel positions below the position to be refined are candidates for a vertical position of the refined motion vector. wherein, when a horizontal pattern is used, the multiple horizontal sub-pixel positions to the right of the position to be refined and the multiple horizontal sub-pixel positions to the left of the position to be refined are candidates for the horizontal position of the refined motion vector, and wherein, when a vertical pattern is used, the multiple vertical sub-pixel positions above the position to be refined and the multiple vertical sub-pixel positions below the position to be refined are candidates for the vertical position of the refined motion vector.
4. The apparatus of claim 3, wherein, The first and second refined motion vectors are 1 / 16 sub-pixel precision.
5. A non-transitory computer readable storage medium storing instructions of a computer program for implementing the method of any of claims 1 to 2.
6. An apparatus for decoding video data comprising frames, the apparatus comprising: one or more processors; a non-transitory computer readable storage medium storing instructions of a computer program that, when executed by the one or more processors, cause implementation of the method of any of claims 1 to 2.
7. An apparatus for decoding video data comprising frames, comprising means for performing the method of any of claims 1 to 2.
Citation Information
Patent Citations
Video encoder with block merging and methods for use therewith
CN104754342A
Method and apparatus for encoding and decoding video data by subpixel motion vector refinement
CN116506638A
Device and method for fast block-matching motion estimation in video encoders
US20060245497A1
Apparatus and method for video processing
US20140211854A1