Method and apparatus for encoding and decoding video data through sub-pixel motion vector refinement
By adopting the decoder-side motion vector derivation method in video coding and refining the motion vector using multi-step search and template signal characteristics, the problem of insufficient motion vector accuracy in FRUC mode is solved, and the coding efficiency and quality are improved.
Patent Information
- Application Number
- CN202310685629.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-04
- Filing Date
- 2018-06-27
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2038-06-27
AI Technical Summary
Existing video coding technologies have not fully utilized the refinement step on the decoder side in motion vector prediction, especially in the FRUC mode, to improve the accuracy of motion vectors, resulting in insufficient coding efficiency.
The decoder-side motion vector derivation method uses a multi-step search step to evaluate the motion vector at 1/16 sub-pixel or higher resolution, combining the template internal signal and matching type characteristics to perform motion vector refinement, including diamond and cross pattern search, and signal content evaluation based on adjacent pixel blocks.
Improved coding efficiency for video coding, especially in FRUC mode, by enhancing the accuracy of motion vectors through finer sub-pixel position evaluation and pattern matching, thereby improving coding quality and compression performance.
Smart Images

Figure CN116506641B_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application 201880037185.X, entitled “Method and device for encoding or decoding video data through sub-pixel motion vector refinement”, filed on June 27, 2018. Technical Field
[0002] The present disclosure relates to methods and apparatus for encoding or decoding video data, and more particularly to encoding according to a particular coding mode using a decoder-side motion vector derivation mode, referred to as a frame rate up-conversion mode or FRUC mode. Background Art
[0003] Predictive coding of video data is based on the division of frames into multiple pixel blocks. For each pixel block, a prediction block is searched for in the available data. The prediction block can be a block in a reference frame different from the current frame in INTER coding mode, or it can be generated from adjacent pixels in the current frame in INTRA coding mode. Different coding modes are defined based on different ways of determining the prediction block. The result of the coding is an indication of the prediction block and a residual block, which is the difference between the block to be coded and the prediction block.
[0004] With regard to the INTER coding mode, the indication of the prediction block is a motion vector that gives the position of the prediction block in the reference image relative to the position of the block to be coded. The motion vector itself is predictively coded based on the motion vector predictor. The HEVC (High Efficiency Video Coding) standard defines several known coding modes for predictive coding of motion vectors, namely the AMVP (Advanced Motion Vector Prediction) mode and the merge derivation process. These modes are based on the construction of a candidate list of motion vector predictors and the signaling of the index of the motion vector predictor to be used for coding in this list. Typically, the residual motion vector is also signaled.
[0005] Recently, a new coding mode for motion vector prediction, called FRUC, has been introduced that defines a decoder-side derivation process for motion vector predictors that is not signaled at all. The result of the derivation process is used as the motion vector predictor without the decoder transmitting any indices or residual motion vectors.
[0006] In FRUC merge mode, the derivation process includes a refinement step to improve the accuracy of the obtained motion vector at the sub-pixel level. This process involves evaluating different sub-pixel positions around the obtained motion vector according to different patterns. Summary of the Invention
[0007] The present invention has been conceived to improve the known refinement steps. It aims to increase the coding efficiency by taking into account the characteristics of the signal and / or matching type inside the template.
[0008] According to a first aspect of the present invention, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the method comprising, for a pixel block:
[0009] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0010] - Evaluate the motion vector list to select a motion vector;
[0011] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions near the selected motion vector; wherein
[0012] - Select at least some sub-pixel positions at 1 / 16 sub-pixel or higher resolution.
[0013] In one embodiment, refining the selected motion vector comprises a plurality of search steps; and wherein at least one of the search steps involves sub-pixel positions at 1 / 16 sub-pixel or higher resolution.
[0014] In one embodiment, the plurality of search steps comprises at least three consecutive steps; each of the three consecutive search steps involves sub-pixel positions at a given resolution; and wherein each given resolution associated with the last two search steps is greater than the given resolution of the previous search step.
[0015] In one embodiment, the plurality of search steps include at least one search step based on a diamond pattern at a first sub-pixel resolution, and two search steps based on a cross pattern at a sub-pixel resolution greater than the first sub-pixel resolution.
[0016] In one embodiment, at least some of the searching steps are performed at a sub-pixel resolution depending on the type of matching used to encode the pixel block.
[0017] In one embodiment, the first search step is a search step based on a diamond pattern at a 1 / 4 sub-pixel resolution, the second search step is a search step based on a cross pattern at a 1 / 8 sub-pixel resolution, and the third search step is a search step based on a cross pattern at a 1 / 16 sub-pixel resolution.
[0018] According to another aspect of the present invention, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the method comprising, for a pixel block:
[0019] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0020] - Evaluate the motion vector list to select a motion vector;
[0021] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions near the selected motion vector; wherein
[0022] - evaluating the sub-pixel position according to at least one pattern selected from a plurality of patterns; and wherein:
[0023] - The plurality of patterns include horizontal patterns and vertical patterns.
[0024] In one embodiment, the plurality of patterns further includes at least a diagonal pattern.
[0025] In one embodiment, a pattern of the plurality of patterns is selected based on edge directions detected in adjacent pixel blocks.
[0026] In one embodiment, at least one pattern of the plurality of patterns is defined at 1 / 16 sub-pixel resolution or higher.
[0027] According to another aspect of the present invention, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the method comprising, for a pixel block:
[0028] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0029] - Evaluate the motion vector list to select a motion vector;
[0030] - determining to refine the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector based on signal content of at least one neighboring pixel block.
[0031] In one embodiment, the signal content is the frequency in a block of adjacent pixels.
[0032] According to another aspect of the present invention, there is provided a method for encoding or decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the method comprising, for a pixel block:
[0033] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0034] -Evaluate the motion vector list to select a motion vector; where:
[0035] - the derivation of the motion vector list of motion vectors is based on a template defined by a pattern of neighboring pixels of a pixel block;
[0036] - A template for derivation of a motion vector list of motion vectors is determined based on a content signal of the template.
[0037] In one embodiment, the signal content is the frequencies in the template.
[0038] According to another aspect of the invention, there is provided a computer program product for a programmable apparatus, the computer program product comprising a sequence of instructions for implementing the method according to the invention when loaded into and executed by the programmable apparatus.
[0039] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores instructions of a computer program for implementing the method according to the present invention.
[0040] According to another aspect of the present invention, there is provided a decoding device for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding device comprising a processor configured to decode a pixel block by:
[0041] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0042] - Evaluate the motion vector list to select a motion vector;
[0043] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions near the selected motion vector; wherein
[0044] - Select at least some sub-pixel positions at 1 / 16 sub-pixel or higher resolution.
[0045] According to another aspect of the present invention, there is provided a decoding device for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding device comprising a processor configured to decode a pixel block by:
[0046] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0047] - Evaluate the motion vector list to select a motion vector;
[0048] - refining the selected motion vector by evaluating the motion vector at sub-pixel positions near the selected motion vector; wherein
[0049] - evaluating the sub-pixel position according to at least one pattern selected from a plurality of patterns; and wherein:
[0050] - The plurality of patterns include horizontal patterns and vertical patterns.
[0051] According to another aspect of the present invention, there is provided a decoding device for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding device comprising a processor configured to decode a pixel block by:
[0052] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0053] - Evaluate the motion vector list to select a motion vector;
[0054] - determining to refine the selected motion vector by evaluating the motion vector at sub-pixel positions in the vicinity of the selected motion vector based on signal content of at least one neighboring pixel block.
[0055] According to another aspect of the present invention, there is provided a decoding device for decoding video data comprising frames, each frame being divided into a plurality of pixel blocks, the decoding device comprising a processor configured to decode a pixel block by:
[0056] - a motion vector list that uses a mode that obtains motion information via a decoder-side motion vector derivation method to derive motion vectors;
[0057] -Evaluate the motion vector list to select a motion vector; where:
[0058] - the derivation of the motion vector list of motion vectors is based on a template defined by a pattern of neighboring pixels of a pixel block;
[0059] - A template for derivation of a motion vector list of motion vectors is determined based on a content signal of the template.
[0060] At least a portion of the method according to the present invention may be computer-implemented. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which are generally referred to herein as "circuits," "modules," or "systems." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium, the computer program product being expressed as having computer-usable program code embodied in the medium.
[0061] Since the present invention can be implemented using software, it can be embodied as computer-readable code for provision to a programmable device on any suitable carrier medium. Tangible, non-transitory carrier media can include storage media such as floppy disks, CD-ROMs, hard drives, magnetic tape devices, or solid-state storage devices. Transient carrier media can include signals such as electric, electronic, optical, acoustic, magnetic, or electromagnetic signals, such as microwave or radio frequency signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Embodiments of the invention will now be described, by way of example only, with reference to the following drawings, in which:
[0063] Figure 1 shows the HEVC encoder architecture;
[0064] Figure 2 The principle of the decoder is shown;
[0065] Figure 3 Template matching and bilateral matching in FRUC merge mode are shown;
[0066] Figure 4 shows the decoding of FRUC merge information;
[0067] Figure 5 Encoder evaluations for merge mode and merge FRUC mode are shown;
[0068] Figure 6 shows the derivation of the merged FRUC pattern at the coding unit and sub-coding unit levels of JEM;
[0069] Figure 7 The template around the current block used for the JEM template matching method is shown;
[0070] Figure 8 Motion vector refinement is shown;
[0071] Figure 9 shows the adaptive sub-pixel resolution for motion vector refinement in an embodiment of the present invention;
[0072] Figure 10 An example of the results obtained using prior art techniques and one embodiment of the present invention is given;
[0073] Figure 11 shows some motion vector refinement search shapes used in one embodiment of the present invention;
[0074] Figure 12 An adaptive search shape for motion vector refinement according to an embodiment of the present invention is shown;
[0075] Figure 13shows adaptive motion vector refinement according to one embodiment of the present invention;
[0076] Figure 14 Shows template selection for motion vector refinement according to one embodiment of the present invention;
[0077] Figure 15 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. DETAILED DESCRIPTION
[0078] Figure 1 The HEVC encoder architecture is shown. In the video encoder, the original sequence 101 is divided into pixel blocks 102 called coding units. The coding mode then affects each block. Two families of coding modes are commonly used in HEVC: modes based on spatial prediction, or INTRA modes 103, and modes based on motion estimation 104 and motion compensation 105, or temporal prediction. INTRA coding units are typically predicted from the coded pixels at their causal boundaries through a process called INTRA prediction.
[0079] Temporal prediction first consists in finding a reference region closest to the coding unit in a previous or future frame, called a reference frame 116, in a motion estimation step 104. This reference region constitutes a prediction block. Next, in a motion compensation step 105, the prediction block is used to predict the coding unit to calculate the residual.
[0080] In both the spatial and temporal prediction cases, the residual is computed by subtracting the coding unit from the original prediction value.
[0081] In INTRA prediction, the prediction direction is encoded. In temporal prediction, at least one motion vector is encoded. However, in order to further reduce the bit rate cost associated with motion vector encoding, the motion vector is not encoded directly. In fact, assuming that the motion is uniform, it is particularly interesting to encode the motion vector as the difference between the motion vector and the motion vectors around it. For example, in the H.264 / AVC coding standard, the motion vector is encoded relative to the median vector calculated between the three blocks above and to the left of the current block. Only the difference calculated between the median vector and the current block motion vector (also called the residual motion vector) is encoded in the bitstream. This is processed in the module "Mv prediction and encoding" 117. The value of each coded vector is stored in the motion vector field 118. The neighboring motion vectors used for prediction are extracted from the motion vector field 118.
[0082] The mode that optimizes the rate-distortion performance is then selected in module 106. To further reduce redundancy, a transform (typically a DCT) is applied to the residual block in module 107, and quantization is applied to the coefficients in module 108. The quantized coefficient block is then entropy coded in module 109, and the result is inserted into the bitstream 110.
[0083] The encoder then decodes the coded frame in modules 111 to 116 for future motion estimation. These steps allow the encoder and decoder to have the same reference frame. To reconstruct the coded frame, the residual is inversely quantized in module 111 and inversely transformed in module 112 to provide a "reconstructed" residual in the pixel domain. Depending on the coding mode (INTER or INTRA), this residual is added to the INTER predictor 114 or the INTRA predictor 113.
[0084] This first reconstruction is then filtered in module 115 by one or more post filters. These post filters are integrated into the encoding and decoding loops. This means that they need to be applied to the reconstructed frame on both the encoder and decoder sides so that the same reference frame is used on both sides. The purpose of this post filtering is to remove compression artifacts.
[0085] exist Figure 2 The principle of the decoder has been shown in . First, the video stream 201 is entropy decoded in module 202. Then, the residual data is inversely quantized in module 203 and inversely transformed in module 204 to obtain pixel values. The mode data is also entropy decoded according to the mode of performing INTRA type decoding or INTER type decoding. In the case of INTRA mode, the INTRA prediction value is determined according to the INTRA prediction mode 205 specified in the bitstream. If the mode is INTER, motion information is extracted from the bitstream 201. It consists of a reference frame index and a motion vector residual. The motion vector prediction value is added to the motion vector residual to obtain a motion vector 210. The motion vector is then used to locate the reference area 206 in the reference frame. Note that the motion vector field data 211 is updated by the decoded motion vector so as to be used for the prediction of the next decoded motion vector. The first reconstruction of the decoded frame is then post-filtered (207) using a post-filter exactly the same as the post-filter used on the encoder side. The output of the decoder is the decompressed video 209.
[0086] The HEVC standard uses 3 different INTER modes: Inter mode, Merge mode and Merge Skip mode. The main difference between these modes is the data signaling in the bitstream. For motion vector coding, the current HEVC standard includes a competitive motion vector prediction scheme compared to its predecessors. This means that on the encoder side several candidates compete according to the rate-distortion criterion in order to find the best motion vector predictor or the best motion information for inter or merge mode, respectively. The index of the best candidate or the best predictor corresponding to the motion information is inserted into the bitstream. The decoder can derive the same set of predictors or candidates and use the best predictor or candidate based on the decoded index.
[0087] The design of the derivation of prediction values and candidates is very important to achieve the best coding efficiency without significantly affecting the complexity. In HEVC, two motion vector derivations are used: one for inter mode (Advanced Motion Vector Prediction (AMVP)) and the other for merge mode (Merge Derivation Process).
[0088] As already mentioned, candidates for the merge mode ("classic" or "skip") represent all motion information: direction, list and reference frame indices, and motion vectors. Several candidates are generated through the merge derivation process described below, each with an index. In the current HEVC design, the maximum number of candidates for both merge modes is equal to 5.
[0089] For the current version of JEM, two types of searches can be performed: template matching and bilateral matching. Figure 3 The principle of bilateral matching 301 is to find the best match between two blocks (sometimes also referred to as templates by analogy with template matching described below) along the motion trajectory of the current coding unit.
[0090] The principle of template matching 302 is to derive the motion information of the current coding unit by calculating the matching cost between the reconstructed pixels around the current block and the neighboring pixels around the block pointed to by the estimated motion vector. The template corresponds to the pattern of neighboring pixels around the current block and corresponds to the corresponding pattern of neighboring pixels around the prediction block.
[0091] For both matching types (template matching or bilateral matching), the different matching costs calculated are compared to find the best matching cost. The motion vector or motion vector pair that obtains the best match is selected as the derived motion information. More details can be found in JVET-F1001.
[0092] Both matching methods provide the possibility to derive full motion information, motion vectors, reference frames and prediction types. Motion information derivation at the decoder side (denoted as "FRUC" in JEM) is applicable to all HEVC inter modes: AMVP, Merge and Merge-Skip.
[0093] For AMVP, all motion information is signaled: uni-prediction or bi-prediction, reference frame index, predictor index motion vector and residual motion vector, and the FRUC method is applied to determine the new predictor which is set as the first predictor in the predictor list. Therefore, its index is 0.
[0094] For merge and merge skip modes, a FRUC flag is signaled for the CU. When the FRUC flag is false ("false"), the merge index is signaled and the regular merge mode is used. When the FRUC flag is true ("true"), an additional FRUC mode flag is signaled to indicate which method (bilateral matching or template matching) will be used to derive motion information for the block. Note that bilateral matching only applies to B frames and not to P frames.
[0095] For merge and merge skip modes, the motion vector field is defined for the current block. This means that vectors are defined for sub-coding units that are smaller than the current coding unit. In addition, with classic merge, one motion vector for each list can form the motion information of a block.
[0096] Figure 4 is a flow chart illustrating the signaling of the FRUC flag for the merge mode of a block. In HEVC terminology, a block can be a coding unit or a prediction unit.
[0097] In the first step 401, the skip flag is decoded to determine whether the coding unit is coded in skip mode. If this flag is tested as false in step 402, the merge flag is then decoded in step 403 and tested in step 405. If the coding unit is coded in skip or merge mode, the merge FRUC flag is decoded in step 404. If the coding unit is not coded in skip or merge mode, the intra prediction information for the classic AMVP inter mode is decoded in step 406. If the FRUC flag for the current coding unit is tested as true in step 407, and if the current slice is a B slice, the match mode flag is decoded in step 408. Note that bilateral matching in FRUC only applies to B slices. If the slice is not a B slice and FRUC is selected, the mode must be template matching, and the match mode flag is not present. If the coding unit is not FRUC, the classic merge index is decoded in step 409.
[0098] The FRUC merge mode competes with the classic merge mode (and other possible merge modes) on the encoder side. Figure 5 The current coding mode evaluation method in JEM is described. First, the classic merge mode of HEVC is evaluated in step 501. In step 502, the candidate list is first evaluated using a simple SAD (sum of absolute differences) between the original block and each candidate in the candidate list. Then, the true rate-distortion (RD) cost of each candidate in the list of constrained candidates is evaluated, as shown by steps 504 to 508. In the evaluation, the rate-distortion with residual (step 505) and the rate-distortion without residual (step 506) are evaluated. Finally, the best merge candidate is determined in step 509, which can be with or without residual.
[0099] The FRUC merge mode is then evaluated in steps 510 through 516. For each matching method in step 510 (i.e., bilateral and template matching), the motion vector field of the current block is obtained in step 511, and full rate-distortion cost estimates with or without residuals are calculated in steps 512 and 513. Based on these rate-distortion costs, the best motion vector 516 with or without residuals is determined in step 515. Finally, the best mode between the classic merge mode and the FRUC merge mode is determined in step 517, before possible evaluation of other modes.
[0100] Figure 6 The encoder-side FRUC merge evaluation method is described. For each matching type (step 601), namely template matching type and bilateral matching type, the CU level is first evaluated by module 61, and then the sub-CU level is evaluated by module 62. The goal is to find the motion information of each sub-CU in the current CU 603.
[0101] Module 61 handles the coding unit level evaluation. A list of motion information is derived in step 611. For each motion information in the list, the distortion costs are calculated and compared with each other in step 612. The best motion vector for the template or the best motion vector pair for the bilateral 613 is the vector that minimizes the cost. Then, a motion vector refinement step 614 is applied to improve the accuracy of the obtained motion vector. With the FRUC method, bilinear interpolation is used for template matching estimation instead of the classic discrete cosine transform interpolation filter (DCTIF) interpolation filter. This reduces the memory access around the block to only one pixel, instead of 7 pixels around the block with conventional DCTIF. In fact, the bilinear interpolation filter only needs 2 pixels to obtain the sub-pixel value in one direction.
[0102] After motion vector refinement, a better motion vector for the current coding unit is obtained in step 615. This motion vector will be used for sub-coding unit level evaluation.
[0103] In step 602, the current coding unit is subdivided into several sub-coding units. The sub-coding units are square blocks, which depend on the division depth of the coding unit in the quadtree structure. The minimum size is 4×4.
[0104] For each sub-CU, the sub-CU level evaluation module 62 evaluates the best motion vector. A motion vector list is derived in step 621, including the best motion vector obtained at the CU level in step 615. For each motion vector, a distortion cost is evaluated in step 622. However, this cost also includes a cost representing the distance between the best motion vector obtained at the coding unit level and the current motion vector to avoid diverging motion vector fields. Based on the minimum cost, the best motion vector 623 is obtained. This vector 623 is then refined by the MV refinement process 624 in the same manner as at the CU level in step 614.
[0105] At the end of the process, for one matching type, the motion information of each sub-CU is obtained. On the encoder side, the best RD cost between the two matching types is compared to select the best matching type. On the decoder side, this information is decoded from the bitstream (in Figure 4 Step 408).
[0106] For the template FRUC matching mode, templates 702, 703 include 4 rows on the top of the block and 4 rows on the left side of block 701 for estimating rate distortion cost, as shown in FIG. Figure 7 Two different templates are used, the left template 702 and the upper template 703. The current block or matching block 701 is not used to determine the distortion.
[0107] Figure 8 shows the use of additional search around the identified best prediction value (613 or 623) Figure 6 The motion vector refinement steps 614 and 624 in FIG.
[0108] The method takes as input the best motion vector predictor 801 identified (612 or 622) in a list.
[0109] In step 802, a diamond search is applied at a resolution corresponding to the 1 / 4 pixel position. The diamond search is based on a diamond pattern as shown in FIG81, which is at the 1 / 4 pixel resolution and centered around the best motion vector. This step results in a new best motion vector 803 at the 1 / 4 pixel resolution.
[0110] In step 804, the best motion vector position 803 obtained by the diamond search becomes the center of a cross search (referring to a cross pattern) at 1 / 4 pixel resolution. The cross search pattern is shown in Figure 82 at 1 / 4 pixel resolution, centered on the best motion vector 803. This step results in a new best motion vector 805 at 1 / 4 pixel resolution.
[0111] In step 806, the new best motion vector position 805 obtained by the search step 804 becomes the center of a cross search at 1 / 8 pixel resolution. This step results in a new best motion vector 807 at 1 / 8 pixel resolution. Figure 83 shows these three search step patterns at 1 / 8 resolution, with all positions tested.
[0112] The present invention has been conceived to improve the known refinement steps. It aims to increase the coding efficiency by taking into account the characteristics of the signal and / or matching type inside the template.
[0113] In one embodiment of the present invention, the sub-pixel refinement accuracy is improved in each step of the refinement method.
[0114] The refinement step is performed by evaluating motion vectors at sub-pixel positions. Sub-pixel positions are determined at a given sub-pixel resolution. The sub-pixel resolution determines the number of sub-pixel positions between two pixels. The higher the resolution, the higher the number of sub-pixel positions between two pixels. For example, 1 / 8 pixel resolution corresponds to 8 sub-pixel positions between two pixels.
[0115] Figure 9 In the first diagram 91, the 1 / 16 pixel resolution is shown. Figure 8 92 . As depicted in the example of FIG92 , the first step maintains a diamond search at 1 / 4 pixel, followed by a cross search at 1 / 8 pixel, and finally a cross search at 1 / 16 pixel. This embodiment offers the possibility of obtaining a more accurate position around the good position. Furthermore, the motion vectors are, on average, closer to the initial position than the previous sub-constraints.
[0116] In an additional embodiment, the refinement is applied only to template matching and not to bilateral matching.
[0117] Figure 10is a flow chart representing these embodiments. The best motion vector 1001 corresponding to the best motion vector (613 or 623) obtained in the previous step is set to the center position of a diamond search 1002 at a 1 / 4 pixel position. Diamond search 1002 results in a new best motion vector 1003. The matching type is tested in step 1004. If the matching type is template matching, the motion vector 1003 obtained by the diamond search is used as the center of a cross search 1009 at 1 / 8 pixel accuracy, thereby obtaining a new best motion vector 1010. A cross search 1011 at 1 / 16 pixel accuracy is performed on this new best motion vector 1010 to obtain a final best motion vector 1012.
[0118] If the matching type is not template matching, a regular cross search of 1 / 4 pixels is performed in step 1005 to obtain a new best motion vector 1006, followed by a 1 / 8 pixel search step 1007 to obtain the final best motion vector 1008.
[0119] This embodiment improves coding efficiency, in particular for single prediction. For dual prediction, averaging between two similar blocks is sometimes analogous to an increase in sub-pixel resolution. For example, if both blocks used for dual prediction come from the same reference frame and the difference between their two motion vectors is equal to a lower sub-pixel resolution, then in this case dual prediction corresponds to an increase in sub-pixel resolution. For single prediction, there is no such additional averaging between the two blocks. Therefore, it is more important to increase the sub-pixel resolution for single prediction, especially when this higher resolution does not need to be signaled in the bitstream.
[0120] In one embodiment, the diamond or cross shaped search pattern is replaced by a horizontal, diagonal or vertical pattern. Figure 11 Various diagrams that can be used in this embodiment are shown. Diagram 1101 shows positions searched based on a 1 / 8 pixel horizontal pattern. Diagram 1102 shows positions searched based on a 1 / 8 pixel vertical pattern. Diagram 1103 shows positions searched based on a 1 / 8 pixel diagonal pattern. Diagram 1104 shows positions searched based on another 1 / 8 pixel diagonal pattern. Diagram 1105 shows positions searched based on a 1 / 16 pixel horizontal pattern.
[0121] A horizontal pattern is a pattern of sub-pixel positions aligned horizontally. A vertical pattern is a pattern of sub-pixel positions aligned vertically. A diagonal pattern is a pattern of sub-pixel positions aligned diagonally.
[0122] The advantage of these patterns is mainly interesting when the block contains an edge, since it provides a better refinement of this edge for the prediction. In fact, in classical motion vector estimation, the refinement of the motion vector gives higher results when additional test positions are chosen in the perpendicular axis of the detected edge.
[0123] Figure 12 An example of a method for selecting a pattern according to this embodiment is given. This flowchart can be used to change, for example Figure 8 One or more pattern searches are performed in modules 802, 804, and 806 in step 1201. When the match type is template matching, step 1202 extracts the left template of the current block. A determination is then made as to whether the block contains an edge. If so, the direction of the edge is determined in step 1203. For example, a gradient can be calculated on the current block to determine the presence and direction of an edge. If the "directionLeft" is horizontal, step 1204 selects a vertical pattern, such as pattern 1102, for motion vector refinement.
[0124] If there is no edge in the left template, or if the identified direction is not horizontal, then the upper template of the current block is extracted in step 1206. It is determined whether the template contains an edge, and if so, the direction is determined in step 1207. If the "directionUp" is tested to be vertical in step 1208, then the pattern selected for motion vector refinement in step 1209 is a horizontal pattern, such as pattern 1101 or 1105. Otherwise, the initial pattern (81, 82, 83) is selected in step 1210.
[0125] Note that if the left block contains a horizontal edge, that edge should pass through the current block. Similarly, if the top block contains a vertical edge, that edge should pass through the current block. Conversely, if the left block contains a vertical edge, that edge should not pass through the current block, and so on.
[0126] If the match type is determined to be bilateral matching in step 1201, a template block is selected in step 1211. Then, in step 1212, it is determined whether the template contains a direction and which direction it is. If the direction is horizontal in step 1213, a vertical pattern is selected in step 1205. If the direction is vertical in step 1214, a horizontal pattern is selected in step 1209. For bilateral matching, other linear patterns can be used as diagonal patterns 1103 and 1104. In that case, the linear pattern that is most perpendicular to the direction of the edge determined in 1212 is selected. If no edge is detected in step 1212, the original pattern is retained.
[0127] An advantage of such a change of the pattern used for motion vector refinement is an increase in coding efficiency, since the pattern is adapted to the signal contained inside the template.
[0128] The pattern can also be changed. For example, diagram 1105 shows a horizontal pattern at 1 / 16 pixel accuracy, where more positions have been concentrated around the initial motion vector position with high motion vector accuracy.
[0129] In one embodiment, the signal content of the template is used to determine the number of locations to test for motion vector refinement. For example, the presence of high frequencies in the template around the current block is used to determine the application of the refinement step. The reason is that in the absence of sufficiently high frequencies, sub-pixel refinement of the motion vector is irrelevant.
[0130] Figure 13 An example of this embodiment is given.
[0131] If the matching mode is tested to be template matching in step 1301, then in step 1302, a left template is extracted, which is typically the adjacent 4 lines of the current block. Then, in step 1303, it is determined whether the block contains high frequencies. For example, this determination can be made by comparing the sum of the gradients with a threshold. If it is tested in step 1304 that the left template contains high frequencies, then a motion vector refinement step 1305 is applied, which corresponds to Figure 6 Steps 614 and 624.
[0132] If the left template does not contain enough high frequencies, then the upper template is extracted in step 1306. In step 1307, it is determined whether the extracted template contains high frequencies. If this is the case in step 1308, then the motion vector refinement step is applied in step 1305. Otherwise, the motion vector refinement step is skipped in step 1309.
[0133] If the match type is determined to be bilateral matching in step 1301, a template is extracted in step 1310. A determination is then made in step 1314 as to whether the extracted template contains high frequencies. If the extracted template is found not to contain high frequencies in step 1313, motion vector refinement is not applied in step 1309. Otherwise, motion vector refinement is applied in step 1305.
[0134] In yet another embodiment, the signal content of the template is used to determine the template to be used for distortion estimation, which is used to determine the motion vector for template matching type FRUC evaluation.
[0135] Figure 14 This embodiment is shown.
[0136] The process begins at step 1401. In steps 1402 and 1403, a left template and an upper template are extracted. For each extracted template, a determination is made in steps 1404 and 1405 as to whether the extracted template contains high frequencies. If the tested templates are found to contain high frequencies in steps 1406 and 1407, they are used for distortion estimation in FRUC evaluation in steps 1408 and 1409. When both templates contain high frequencies, they are both used in FRUC evaluation. If neither template contains high frequencies, the upper template is selected for distortion estimation in step 1411.
[0137] All these embodiments can be combined.
[0138] Figure 15 FIG1 is a schematic block diagram of a computing device 1500 for implementing one or more embodiments of the present invention. The computing device 1500 may be a device such as a microcomputer, a workstation, or a lightweight portable device. The computing device 1500 includes a communication bus connected to:
[0139] - Central processing unit 1501, such as a microprocessor, denoted as CPU;
[0140] - Random access memory 1502, represented as RAM, for storing executable code of the method according to an embodiment of the present invention, and registers suitable for recording variables and parameters necessary for implementing the method for encoding or decoding at least a portion of an image according to an embodiment of the present invention, the storage capacity of which can be expanded by, for example, an optional RAM connected to an expansion port;
[0141] - a read-only memory 1503, denoted as ROM, for storing computer programs for implementing embodiments of the present invention;
[0142] Network interface 1504 is typically connected to a communications network through which digital data to be processed is sent or received. Network interface 1504 can be a single network interface or comprised of a group of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of a software application running in CPU 1501.
[0143] - User interface 1505 may be used to receive input from a user or display information to a user;
[0144] - A hard disk 1506 denoted as HD may be provided as a mass storage device;
[0145] - The I / O module 1507 may be used to receive / send data from / to external devices such as a video source or a display.
[0146] The executable code may be stored in the read-only memory 1503, on the hard disk 1506, or on a removable digital medium (e.g., a disk). According to a variant, the executable code of the program may be received via the communication network via the network interface 1504, so as to be stored in one of the storage devices of the communication device 1500, such as the hard disk 1506, and then executed.
[0147] The central processing unit 1501 is adapted to control and direct the execution of portions of the instructions or software code of one or more programs according to embodiments of the present invention, which instructions are stored in one of the aforementioned storage devices. After power is applied, the CPU 1501 is capable of executing instructions related to software applications from the main RAM memory 1502, after having loaded those instructions from, for example, the program ROM 1503 or the hard disk (HD) 1506. Such software applications, when executed by the CPU 1501, cause the steps of the flowcharts of the present invention to be performed.
[0148] Any step of the algorithm of the present invention can be implemented in software by a programmable computer (such as a PC ("personal computer"), a DSP ("digital signal processor") or a microcontroller executing a set of instructions or a program; or can be implemented in hardware by a machine or a dedicated component (such as an FPGA ("field programmable gate array") or an ASIC ("application-specific integrated circuit")).
[0149] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to the specific embodiments, and modifications within the scope of the invention will be apparent to those skilled in the art.
[0150] Many further modifications and variations will be suggested to those skilled in the art upon reference to the foregoing illustrative embodiments, which are given by way of example only and are not intended to limit the scope of the invention, which is determined solely by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.
[0151] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
Claims
1. A method for encoding video data, for generating encoded data to be decoded in a decoding method, wherein the video data comprises frames, each frame being divided into blocks, the encoding method comprising: determining whether to use a first mode, the first mode being a mode in which at least refinement of a motion vector associated with a block to be decoded can be performed in a decoding method; In the case of using the first mode, determining whether to perform refinement on the motion vector based on at least sample values of video data in an area different from the block to be encoded; If it is determined that refinement is to be performed, refining the motion vector associated with the block to be encoded; and In case refinement is performed, the block is encoded using the refined motion vector, wherein a plurality of patterns for thinning can be used in a first mode in the decoding method; wherein the plurality of patterns include horizontal patterns and vertical patterns, wherein, when a horizontal pattern is used, a plurality of horizontal sub-pixel positions to the right of the position to be refined and a plurality of horizontal sub-pixel positions to the left of the position to be refined are candidates for the horizontal position of the refined motion vector, and When the vertical pattern is used, a plurality of vertical sub-pixel positions above the position to be refined and a plurality of vertical sub-pixel positions below the position to be refined are candidates for the vertical position of the refined motion vector.
2. The method according to claim 1, wherein The refined motion vectors are 1 / 16 sub-pixel accurate.
3. The method according to claim 1 or 2, wherein: The sample values in the block to be encoded are not used to determine whether to perform the refinement.
4. A method for decoding video data, wherein the video data includes frames, each frame is divided into blocks, and the decoding method comprises: determining whether to use a first mode, the first mode being a mode in which at least refinement of a motion vector associated with a block to be decoded can be performed in a decoding method; In case the first mode is used, determining whether to perform refinement on the motion vector based on at least sample values of video data in a region different from the block to be decoded; If it is determined that refinement is to be performed, refining the motion vector associated with the block to be decoded; and In case refinement is performed, the block is decoded using the refined motion vector, wherein a plurality of patterns for thinning can be used in a first mode in the decoding method; wherein the plurality of patterns include horizontal patterns and vertical patterns, wherein, when a horizontal pattern is used, a plurality of horizontal sub-pixel positions to the right of the position to be refined and a plurality of horizontal sub-pixel positions to the left of the position to be refined are candidates for the horizontal position of the refined motion vector, and When the vertical pattern is used, a plurality of vertical sub-pixel positions above the position to be refined and a plurality of vertical sub-pixel positions below the position to be refined are candidates for the vertical position of the refined motion vector.
5. The method according to claim 4, wherein The refined motion vectors are 1 / 16 sub-pixel accurate.
6. The method according to claim 4 or 5, wherein: The sample values in the block to be decoded are not used to determine whether to perform the refinement.
7. A video data encoding apparatus for generating encoded data to be decoded in a decoding apparatus, the video data comprising frames, each frame being divided into blocks, the video data encoding apparatus comprising one or more processors, the one or more processors being configured to encode blocks of pixels as follows: determining whether to use a first mode, the first mode being a mode in which at least refinement of a motion vector associated with a block to be decoded can be performed in a decoding device; In the case of using the first mode, determining whether to perform refinement on the motion vector based on at least sample values of video data in an area different from the block to be encoded; If it is determined that refinement is to be performed, refining the motion vector associated with the block to be encoded; and In case refinement is performed, the block is encoded using the refined motion vector, wherein a plurality of patterns for thinning are capable of being used by the decoding device in a first mode; wherein the plurality of patterns include horizontal patterns and vertical patterns, wherein, when a horizontal pattern is used, a plurality of horizontal sub-pixel positions to the right of the position to be refined and a plurality of horizontal sub-pixel positions to the left of the position to be refined are candidates for the horizontal position of the refined motion vector, and When the vertical pattern is used, a plurality of vertical sub-pixel positions above the position to be refined and a plurality of vertical sub-pixel positions below the position to be refined are candidates for the vertical position of the refined motion vector.
8. The device according to claim 7, wherein The refined motion vectors are 1 / 16 sub-pixel accurate.
9. The device according to claim 7 or 8, wherein The sample values in the block to be encoded are not used to determine whether to perform the refinement.
10. A device for decoding video data, the video data comprising frames, each frame being divided into blocks, the device comprising one or more processors, the one or more processors being configured to decode blocks of pixels as follows: determining whether to use a first mode, the first mode being a mode in which at least refinement of a motion vector associated with a block to be decoded can be performed in a decoding device; In case the first mode is used, determining whether to perform refinement on the motion vector based on at least sample values of video data in a region different from the block to be decoded; If it is determined that refinement is to be performed, refining the motion vector associated with the block to be decoded; and In case refinement is performed, the block is decoded using the refined motion vector, wherein a plurality of patterns for thinning are capable of being used by the decoding device in a first mode; wherein the plurality of patterns include horizontal patterns and vertical patterns, wherein, when a horizontal pattern is used, a plurality of horizontal sub-pixel positions to the right of the position to be refined and a plurality of horizontal sub-pixel positions to the left of the position to be refined are candidates for the horizontal position of the refined motion vector, and When the vertical pattern is used, a plurality of vertical sub-pixel positions above the position to be refined and a plurality of vertical sub-pixel positions below the position to be refined are candidates for the vertical position of the refined motion vector.
11. The device according to claim 10, wherein The refined motion vectors are 1 / 16 sub-pixel accurate.
12. The device according to claim 10 or 11, wherein The sample values in the block to be decoded are not used to determine whether to perform the refinement.
13. A non-transitory computer-readable storage medium storing instructions of a computer program for implementing the method according to any one of claims 1 to 3.
14. A non-transitory computer-readable storage medium storing instructions of a computer program for implementing the method according to any one of claims 4 to 6.
15. An apparatus for encoding video data comprising frames, the apparatus comprising: one or more processors; A non-transitory computer-readable storage medium storing instructions of a computer program, which, when executed by one or more processors, causes implementation of the method according to any one of claims 1 to 3.
16. An apparatus for decoding video data comprising frames, the apparatus comprising: one or more processors; A non-transitory computer-readable storage medium storing instructions of a computer program, which, when executed by one or more processors, causes implementation of the method according to any one of claims 4 to 6.
17. An apparatus for encoding video data comprising frames, comprising means for performing the method according to any one of claims 1 to 3.
18. An apparatus for decoding video data comprising frames, comprising means for performing the method according to any one of claims 4 to 6.