Motion vector decoding method and coding method

By using spatial and temporal candidate blocks to determine optimal motion vector predictions and resolutions, the method addresses challenges in video encoding and decoding, achieving improved accuracy and efficiency.

JP2025090760AActive Publication Date: 2025-06-17SAMSUNG ELECTRONICS CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025039829
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-10-31
Filing Date
2025-03-13
Publication Date
2025-06-17
Estimated Expiration
2035-11-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding methods face challenges in accurately predicting and efficiently encoding motion vectors, particularly in determining the optimal motion vector resolution and reducing computational complexity.

Method used

A method and apparatus for determining an optimal predicted motion vector and its resolution, using spatial and temporal candidate blocks to obtain motion vector candidates at multiple resolutions, and adaptively encoding or decoding video to reduce device complexity.

Benefits of technology

This approach enhances the accuracy of motion vector prediction and encoding, reduces computational complexity, and improves the efficiency of video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090760000001_ABST
    Figure 2025090760000001_ABST
Patent Text Reader

Abstract

To provide a method and a device of them, that predict and code a motion vector of a video image, and a method and a device of them, that predict and decode the motion vector of the video image.SOLUTION: A motion vector coding device contains: a prediction part that uses a spatial candidate block and a time-variant candidate block of the current block, acquires a prediction motion vector candidate of a plurality of predetermined motion vector resolutions, uses the prediction motion vector candidate, and determines a prediction motion vector of the current block, the motion vector of the current block, and the motion vector resolution of the current block; and a coding part that codes information indicating a prediction motion vector of the current block, and information indicating a residual motion vector between the motion vector of the current block and the prediction motion vector of the current block, and the motion vector resolution of the current block. The plurality of predetermined motion vector resolutions contains a resolution of a pixel unit that is larger than that of a one pixel unit.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video encoding method and a video decoding method, and more specifically, to a method and apparatus for predicting and encoding a motion vector of a video image, and a method and apparatus for predicting and decoding a motion vector of a video image.

Background Art

[0002] In a codec such as H.264 AVC (advanced video coding) and HEVC (high efficiency video coding), in order to predict the motion vector of the current block, the motion vector of a previously encoded block adjacent to the current block or a block at the same position in a previously encoded picture can be used as the predicted motion vector of the current block.

[0003] In a video encoding method and a video decoding method, in order to encode a video, one picture is divided into macroblocks, and each macroblock can be predicted and encoded using inter prediction or intra prediction.

[0004] The inter prediction is a method of removing temporal redundancy between pictures to compress a video, and motion estimation encoding is a typical example. The motion estimation encoding uses at least one reference picture to predict each block of the current picture. A reference block most similar to the current block is searched within a predetermined search range using a predetermined evaluation function.

[0005] Predict the current block based on the reference block, and encode the residual block generated by subtracting the predicted block generated as the prediction result from the current block. At this time, in order to perform the prediction more accurately, interpolation is performed on the search range of the reference picture to generate sub-pixels with a pixel unit smaller than the integer pel unit, and inter prediction can be performed based on the generated sub-pixels. Summary of the Invention

[0006] According to one embodiment, a motion vector decoding device, a motion vector encoding device, and a method thereof can determine an optimal predicted motion vector and the resolution of the motion vector, and efficiently encode or decode video to reduce the complexity of the device.

[0007] On the other hand, the technical problems and effects of the present invention are not limited to the features mentioned above. Other technical problems that have not been mentioned or are different will be clearly understood by those skilled in the art from the following description. Brief Description of the Drawings

[0008]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

DETAILED DESCRIPTION OF THE INVENTION

[0009] In a motion vector encoding device according to an embodiment, using a spatial candidate block and a temporal candidate block of a current block, a predicted motion vector candidate of a plurality of predetermined motion vector resolutions is obtained, and using the predicted motion vector candidate, a predicted motion vector of the current block, a motion vector of the current block, and a prediction unit that determines a motion vector resolution of the current block, and information indicating the predicted motion vector of the current block, the motion vector of the current block, a residual motion vector between the predicted motion vector of the current block, and an encoding unit that encodes information indicating the motion vector resolution of the current block, wherein the plurality of predetermined motion vector resolutions includes a resolution in pixel units larger than a resolution in pixel units of 1 pixel.

[0010] The prediction unit uses a set of first predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, searches for a reference block in pixel units of the first motion vector resolution, and uses a set of second predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, searches for a reference block in pixel units of the second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other, and the first predicted motion vector candidate set and the second predicted motion vector candidate set are obtained from different candidate blocks among the candidate blocks included in the spatial candidate block and the temporal candidate block.

[0011] The prediction unit uses a set of first predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, searches for a reference block in pixel units of the first motion vector resolution, uses a set of second predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, searches for a reference block in pixel units of the second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other, and the first set of predicted motion vector candidates and the second set of predicted motion vector candidates include different numbers of predicted motion vector candidates.

[0012] When the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the encoding unit encodes the residual motion vector after down-scaling it according to the resolution of the motion vector of the current block.

[0013] If the current block is the current encoding unit that constitutes the video, and the motion vector resolution is determined to be the same for each encoding unit, and there is a prediction unit predicted to be in the AMVP (advanced motion vector prediction) mode within the current encoding unit, the encoding unit encodes, as information indicating the motion vector resolution of the current block, information indicating the motion vector resolution of the prediction unit predicted to be in the AMVP mode once.

[0014] If the current block is the current encoding unit that constitutes the video, and the motion vector resolution is determined to be the same for each prediction unit, and there is a prediction unit predicted to be in the AMVP mode within the current encoding unit, the encoding unit encodes, as information indicating the motion vector resolution of the current block, information indicating the motion vector resolution for each prediction unit predicted to be in the AMVP mode existing within the current block.

[0015] In a motion vector encoding apparatus according to an embodiment, spatial candidate blocks and temporal candidate blocks of a current block are used to obtain prediction motion vector candidates at a plurality of predetermined motion vector resolutions, and the prediction motion vector candidates are used to determine a prediction motion vector of the current block, a motion vector of the current block, and a motion vector resolution of the current block. A prediction unit, and information indicating the prediction motion vector of the current block, the motion vector of the current block, a residual motion vector between the motion vector of the current block and the prediction motion vector of the current block, and information indicating the motion vector resolution of the current block. The encoding unit includes an encoding unit that encodes the information, and the prediction unit uses a set of first prediction motion vector candidates including one or more prediction motion vector candidates selected from the prediction motion vector candidates, and searches for a reference block in pixel units of the first motion vector resolution. Using a set of second prediction motion vector candidates including one or more prediction motion vector candidates selected from the prediction motion vector candidates, searching for a reference block in pixel units of the second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other, and the first prediction motion vector candidate set and the second prediction motion vector candidate set are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks, or include different numbers of prediction motion vector candidates.

[0016] In a motion vector encoding apparatus according to an embodiment, a merge candidate list including at least one merge candidate related to a current block is generated, and a motion vector of one of the merge candidates included in the merge candidate list is used to determine and encode the motion vector of the current block. The merge candidate list includes motion vectors of the candidates included in the merge candidate list and motion vectors downscaled by a plurality of predetermined motion vector resolutions.

[0017] The downscaling is characterized by selecting any one of the pixels located around the pixel indicated by the motion vector of the minimum motion vector resolution instead of the pixel indicated by the motion vector of the minimum motion vector resolution, and adjusting to indicate the selected pixel based on the resolution of the motion vector of the current block.

[0018] In a motion vector decoding device according to an embodiment, using a spatial candidate block and a temporal candidate block of a current block, obtaining prediction motion vector candidates of a plurality of predetermined motion vector resolutions, obtaining information indicating the prediction motion vector of the current block among the prediction motion vector candidates, a obtaining unit that obtains a residual motion vector between the motion vector of the current block and the prediction motion vector of the current block, and information indicating the motion vector resolution of the current block, and a decoding unit that restores the motion vector of the current block based on the residual motion vector, the information indicating the prediction motion vector of the current block, and the motion vector resolution information of the current block, wherein the plurality of predetermined motion vector resolutions include resolutions in pixel units larger than the resolution in pixel units of one pixel.

[0019] The prediction motion vector candidates of the plurality of predetermined motion vector resolutions include a set of first prediction motion vector candidates including one or more prediction motion vector candidates of a first motion vector resolution, and a set of second prediction motion vector candidates including one or more prediction motion vector candidates of a second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other, and the first prediction motion vector candidate set and the second prediction motion vector candidate set are obtained from different candidate blocks among the candidate blocks included in the spatial candidate block and the temporal candidate block, or include different numbers of prediction motion vector candidates.

[0020] When the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the decoding unit restores the residual motion vector by up-scaling it according to the minimum motion vector resolution.

[0021] The current block is the current encoding unit that constitutes the video, and the motion vector resolution is determined to be the same for each encoding unit. If there is a prediction unit predicted to be in the AMVP mode within the current encoding unit, the acquisition unit acquires, from the bitstream, as information indicating the motion vector resolution of the current block, information indicating the motion vector resolution of the prediction unit predicted to be in the AMVP mode once.

[0022] In a motion vector decoding apparatus according to an embodiment, a merge candidate list including at least one merge candidate related to a current block is generated, a motion vector of one candidate among the merge candidates included in the merge candidate list is used to determine and decode the motion vector of the current block, and the merge candidate list includes motion vectors obtained by down-scaling the motion vectors of the candidates included in the merge candidate list according to a plurality of predetermined motion vector resolutions.

[0023] In a motion vector decoding apparatus according to an embodiment, spatial candidate blocks and temporal candidate blocks of a current block are used to obtain predicted motion vector candidates at a plurality of predetermined motion vector resolutions, information indicating a predicted motion vector of the current block is obtained from among the predicted motion vector candidates, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block are obtained, and a decoding unit that restores the motion vector of the current block based on the residual motion vector, the information indicating the predicted motion vector of the current block, and the motion vector resolution information of the current block is included, and the predicted motion vector candidates at the plurality of predetermined motion vector resolutions include a set of first predicted motion vector candidates including one or more predicted motion vector candidates at a first motion vector resolution and a set of second predicted motion vector candidates including one or more predicted motion vector candidates at a second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other, and the first predicted motion vector candidate set and the second predicted motion vector candidate set are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks, or include different numbers of predicted motion vector candidates.

[0024] A computer-readable recording medium recording a program for causing a computer to execute the motion vector decoding method according to an embodiment is provided.

[0025] Hereinafter, with reference to FIGS. 1A to 7B, a motion vector resolution encoding apparatus, a decoding apparatus, and a method thereof for a video encoding apparatus, a video decoding apparatus, and a method thereof according to an embodiment are proposed. Hereinafter, the video encoding apparatus and the method thereof may each include a motion vector encoding apparatus and a motion vector encoding method described later. Further, the video decoding apparatus and the method thereof may each include a motion vector decoding apparatus and a motion vector decoding method described later.

[0026] Also, referring to FIGS. 8 to 20, a video encoding technique and a video decoding technique based on an encoding unit of a tree structure according to an embodiment applicable to the previously proposed video encoding method and video decoding method are disclosed. Also, referring to FIGS. 21 to 27, an embodiment applicable to the previously proposed video encoding method and video decoding method is disclosed.

[0027] Hereinafter, "video" can indicate a still image or a moving picture of video, that is, video itself.

[0028] Hereinafter, "sample" is data assigned to a sampling position of video and means data to be processed. For example, in a video in a spatial region, a pixel is also a sample.

[0029] Hereinafter, "current block" means a block of an encoding unit or a prediction unit of the current video to be encoded or decoded.

[0030] Throughout the specification, when a part states that a certain component "includes" or is "configured" in a certain component, it means that, unless there is a special contrary statement, it does not exclude other components and may further include other components. Also, the term "part" used in the specification means a hardware component such as software, FPGA, or ASIC, and a "part" can perform a certain role. However, "part" is not limited in meaning to software or hardware. A "part" may also be configured to be on a recordable medium that can be addressed, or may also be configured to cause a further processor to reproduce. Therefore, by way of example, a "part" includes components such as software components, object-oriented software components, class components, and task components, and processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided from among the components and "parts" are further combined into fewer components and "parts" or further separated into additional components and "parts".

[0031] First, referring to FIGS. 1A to 7B, there is disclosed a motion vector encoding apparatus and method for encoding video, and a motion vector decoding apparatus and method for decoding video according to an embodiment.

[0032] FIG. 1A shows a block diagram of a motion vector encoding apparatus 10 according to an embodiment.

[0033] In video encoding, inter prediction means a prediction method that uses the similarity between the current video and other videos. Among the reference videos restored prior to the current video, a reference region similar to the current region of the current video is detected, the distance in coordinates between the current region and the reference region is represented by a motion vector, and the difference in pixel values between the current region and the reference region is represented by residual data. Therefore, by performing inter prediction on the current region, instead of directly outputting the video information of the current region, an index indicating the reference video, a motion vector, and residual data can be output, improving the efficiency of encoding / decoding.

[0034] The motion vector encoding device 10 according to one embodiment can encode the motion vectors used for performing inter prediction for each block of each video of a video. The type of the block can also be square or rectangular, or any geometric shape, and is not limited to a data unit of a certain size. A block according to one embodiment can also be a maximum coding unit, a coding unit, a prediction unit, a transform unit, etc. among the coding units with a tree structure. The video encoding / decoding method based on the coding units with a tree structure will be described later with reference to FIGS. 8 to 20.

[0035] The motion vector encoding device 10 may include a prediction unit 11 and a coding unit 13.

[0036] The motion vector encoding device 10 according to one embodiment can perform encoding of the motion vectors for inter prediction for each video block of each video.

[0037] The motion vector encoding device 10 can refer to the motion vectors of blocks different from the current block to determine the motion vector of the current block for motion vector prediction, PU merging, or AMVP (advanced motion vector prediction).

[0038] The motion vector encoding device 10 according to one embodiment can determine the motion vector of the current block by referring to the motion vectors of other blocks that are temporally or spatially adjacent to the current block. The motion vector encoding device 10 can determine a prediction candidate including the motion vectors of candidate blocks that can also be reference targets for the motion vector of the current block. The motion vector encoding device 10 can determine the motion vector of the current block by referring to one motion vector selected from among the prediction candidates.

[0039] The motion vector encoding device 10 according to one embodiment divides into prediction units, which are units serving as the basis for prediction divided from coding units, and through motion estimation, searches for the prediction block most similar to the current coding unit in the reference picture adjacent to the current picture, and can determine a motion parameter indicating the motion information between the current block and the prediction block.

[0040] The prediction unit according to one embodiment starts to be divided from the coding unit and is divided only once without being divided in a quadtree form. For example, one coding unit is divided into a plurality of prediction units, and the prediction units generated by the division are not further divided additionally.

[0041] For motion estimation, the motion vector encoding device 10 can expand the resolution of the reference picture up to n times (n is an integer) in the horizontal direction and up to n times in the vertical direction, and determine the motion vector of the current block with the accuracy of 1 / n pixel positions. In that case, n is referred to as the minimum motion vector resolution of the picture.

[0042] For example, when the pixel unit of the minimum motion vector resolution is 1 / 4 pixel, the motion vector encoding device 10 can quadruple the resolution of the picture in the vertical and horizontal directions and determine the motion vector with an accuracy of up to 1 / 4 pixel position. However, depending on the characteristics of the video, it may not be sufficient to determine the motion vector in units of 1 / 4 pixel. Conversely, determining the motion vector in units of 1 / 4 pixel may also be less efficient than determining the motion vector at the 1 / 2 pixel position. Therefore, the motion vector encoding device 10 according to one embodiment can adaptively determine the resolution of the motion vector of the current block and encode the determined predicted motion vector, actual motion vector, and motion vector resolution.

[0043] The prediction unit 11 according to one embodiment can use one of the motion vector prediction candidates to determine the optimal motion vector for the inter prediction of the current block.

[0044] The prediction unit 11 uses the spatial candidate block and temporal candidate block of the current block to obtain prediction motion vector candidates with a plurality of predetermined motion vector resolutions, and uses the prediction motion vector candidates to determine the prediction motion vector of the current block, the motion vector of the current block, and the motion vector resolution of the current block. The spatial candidate block may include at least one peripheral block that is spatially adjacent to the current block. Also, the temporal candidate block may include at least one block located at the same position as the current block and a peripheral block that is spatially adjacent to the block at the same position in a reference picture having a POC (picture order count) different from that of the current block. The prediction unit 11 according to one embodiment can copy, combine, or transform at least one prediction motion vector candidate as it is to determine the motion vector of the current block.

[0045] The prediction unit 11 can determine the predicted motion vector of the current block, the motion vector of the current block, and the motion vector resolution of the current block by using the predicted motion vector candidates. The plurality of predetermined motion vector resolutions may include resolutions in pixel units larger than the resolution in pixel units of one pixel. That is, the plurality of predetermined motion vector resolutions may include resolutions such as 2-pixel units, 3-pixel units, 4-pixel units, etc. However, the plurality of predetermined motion vector resolutions do not necessarily include resolutions of one pixel unit or more, and may be composed of only resolutions of one pixel unit or less.

[0046] The prediction unit 11 can use a set of first predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates to search for a reference block in pixel units of the first motion vector resolution, and use a set of second predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates to search for a reference block in pixel units of the second motion vector resolution. The first motion vector resolution and the second motion vector resolution are also different from each other. The first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks. The first set of predicted motion vector candidates and the second set of predicted motion vector candidates may include different numbers of predicted motion vector candidates.

[0047] The method by which the prediction unit 11 determines the predicted motion vector of the current block, the motion vector of the current block, and the motion vector resolution of the current block by using the predicted motion vector candidates of the plurality of predetermined motion vector resolutions will be described later with reference to FIG. 4B.

[0048] The symbolization unit 13 can symbolize information indicating the predicted motion vector of the current block, the motion vector of the current block, the residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the resolution of the motion vector of the current block. Since the symbolization unit 13 can use the residual motion vector between the actual motion vector and the predicted motion vector and use fewer bits to symbolize the motion vector of the current block, the compression rate of video symbolization can be improved. As will be described later, the symbolization unit 13 can symbolize an index indicating the resolution of the motion vector of the current block. Further, the symbolization unit 13 can down-scale and symbolize the residual motion vector based on the difference between the minimum motion vector resolution and the resolution of the motion vector of the current block.

[0049] FIG. 1B shows a flowchart of a motion vector symbolization method according to an embodiment.

[0050] In step 12, the motion vector symbolization apparatus 10 according to an embodiment can use the motion vectors of the spatial candidate blocks and the temporal candidate blocks of the current block to obtain candidate predicted motion vectors with a plurality of predetermined motion vector resolutions. The plurality of predetermined motion vector resolutions may include resolutions in pixel units larger than the resolution in pixel units of one pixel.

[0051] In step 14, the motion vector symbolization apparatus 10 according to an embodiment can use the candidate predicted motion vectors obtained in step 12 to determine the predicted motion vector of the current block, the motion vector of the current block, and the resolution of the motion vector of the current block.

[0052] In step 16, the motion vector symbolization apparatus 10 according to an embodiment can symbolize the information indicating the predicted motion vector, the motion vector of the current block, the residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and the information indicating the resolution of the motion vector of the current block determined in step 14.

[0053] FIG. 2A shows a block diagram of a motion vector decoding apparatus according to an embodiment.

[0054] The motion vector decoding apparatus 20 can parse the received bitstream and determine a motion vector for performing inter prediction of the current block.

[0055] The acquisition unit 21 can acquire predicted motion vector candidates with a plurality of predetermined motion vector resolutions using the spatial candidate blocks and the temporal candidate blocks of the current block. The spatial candidate blocks may include at least one peripheral block that is spatially adjacent to the current block. Also, the temporal candidate blocks may include at least one block located at the same position as the current block and peripheral blocks that are spatially adjacent to the block at the same position within a reference picture having a POC different from that of the current block. The plurality of predetermined motion vector resolutions may include resolutions in pixel units larger than the resolution in pixel units of one pixel. That is, the plurality of predetermined motion vector resolutions may include resolutions such as 2-pixel units, 3-pixel units, 4-pixel units, etc. However, the plurality of predetermined motion vector resolutions do not necessarily include resolutions of one pixel unit or more and may be configured only with resolutions of one pixel unit or less.

[0056] The predicted motion vector candidates with a plurality of predetermined motion vector resolutions may include a set of first predicted motion vector candidates including one or more predicted motion vector candidates with a first motion vector resolution, and a set of second predicted motion vector candidates including one or more predicted motion vector candidates with a second motion vector resolution. The first motion vector resolution and the second motion vector resolution are different from each other. The first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks. Also, the first set of predicted motion vector candidates and the second set of predicted motion vector candidates may include different numbers of predicted motion vector candidates.

[0057] The acquisition unit 21 can acquire information indicating the predicted motion vector of the current block from among the received predicted motion vector candidates, and can acquire the residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information related to the resolution of the motion vector of the current block.

[0058] The decoding unit 23 can restore the motion vector of the current block based on the residual motion vector acquired by the acquisition unit 21, the information indicating the predicted motion vector of the current block, and the motion vector resolution information of the current block. The decoding unit 23 can upscale and restore the data related to the received residual motion vector based on the difference between the minimum motion vector resolution and the motion vector resolution of the current block.

[0059] Referring to FIGS. 5A to 5D, a method by which the motion vector decoding apparatus 20 according to various embodiments parses the received bit stream and acquires the motion vector resolution related to the current block will be described later.

[0060] FIG. 2B shows a flowchart of a motion vector decoding method according to an embodiment.

[0061] In step 22, the motion vector decoding apparatus 20 according to an embodiment can acquire predicted motion vector candidates of a plurality of predetermined motion vector resolutions using the spatial candidate blocks and the temporal candidate blocks of the current block.

[0062] In step 24, the motion vector decoding apparatus 20 according to an embodiment can acquire information indicating the predicted motion vector of the current block from among the predicted motion vector candidates, and can acquire the residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block. The plurality of predetermined motion vector resolutions may include resolutions in pixel units larger than the resolution in pixel units of 1 pixel.

[0063] In step 26, the motion vector decoding device 20 according to one embodiment can reconstruct the motion vector of the current block based on the residual motion vector obtained in step 24, information indicating the predicted motion vector of the current block, and motion vector resolution information of the current block.

[0064] FIG. 3A illustrates interpolation for motion compensation based on multiple resolutions.

[0065] The motion vector encoding device 10 can determine a motion vector of a predetermined multiple motion vector resolution for inter-predicting the current block. The predetermined multiple motion vector resolution is 2 k The resolution may include a pixel unit (k is an integer). If k is greater than 0, the motion vector may indicate only some pixels in the reference image, and if k is less than 0, an n-tap (n is an integer) FIR filter (finite impulse response filter) may be used to perform interpolation to generate sub-pixel pixels, and the generated sub-pixel pixels may be indicated. For example, the motion vector encoding device 10 may determine the minimum motion vector resolution to be 1 / 4 pixel units, and determine the resolutions of a plurality of predetermined motion vectors to be 1 / 4, 1 / 2, 1, and 2 pixel units.

[0066] For example, sub-pixels (a to l) in 1 / 2 pixel units can be generated by interpolating using an n-tap FIR filter. For vertical 1 / 2 sub-pixels, sub-pixel a can be generated by interpolating using integer pixel units A1, A2, A3, A4, A5, and A6, and sub-pixel b can be generated by interpolating using integer pixel units B1, B2, B3, B4, B5, and B6. Sub-pixels c, d, e, and f can be generated in the same manner.

[0067] The pixel values of the horizontal sub-pixels are calculated as follows. For example, a = (A1 - 5×A2 + 20×A3 + 20×A4 - 5×A5 + A6) / 32, b = (B1 - 5×B2 + 20×B3 + 20×B4 - 5×B5 + B6) / 32 are calculated in this way. The pixel values of sub-pixels c, d, e, and f are also calculated by the same method.

[0068] Similar to the horizontal sub-pixels, the vertical sub-pixels can also be interpolated and generated using a 6-tap FIR filter. Sub-pixel g can be generated using A1, B1, C1, D1, E1, and F1, and sub-pixel h can be generated using A2, B2, C2, D2, E2, and F2.

[0069] The pixel values of the vertical sub-pixels are also calculated by the same method as the pixel values of the horizontal sub-pixels. For example, it can be calculated as g = (A1 - 5×B1 + 20×C1 + 20×D1 - 5×E1 + F1) / 32.

[0070] The 1 / 2-pixel unit sub-pixel m in the diagonal direction is interpolated using other 1 / 2-pixel unit sub-pixels. In other words, the pixel value of sub-pixel m is calculated as m = (a - 5×b + 20×c + 20×d - 5×e + f) / 32.

[0071] If 1 / 2-pixel unit sub-pixels are generated, 1 / 4-pixel unit sub-pixels can be generated using integer-pixel unit pixels and 1 / 2-pixel unit sub-pixels. Interpolation is performed using two adjacent pixels to generate 1 / 4-pixel unit sub-pixels. Alternatively, the 1 / 4-pixel unit sub-pixels are generated by directly applying an interpolation filter to the integer-pixel unit pixel values without using the 1 / 2-pixel unit sub-pixel values.

[0072] The aforementioned interpolation filter is described by taking a 6-tap filter as an example, but the motion vector encoding device 10 can use filters with other numbers of taps to interpolate pictures. For example, the interpolation filter may include 4-tap, 7-tap, 8-tap, and 12-tap filters.

[0073] As shown in FIG. 3A, if interpolation is performed on the reference picture to generate sub-pixels in units of 1 / 2 pixel and sub-pixels in units of 1 / 4 pixel, the interpolated reference picture is compared with the current block, and the block with the minimum SAD (sum of absolute difference) or rate-distortion cost is searched in units of 1 / 4 pixel, and a motion vector having a resolution of 1 / 4 pixel unit is determined.

[0074] FIG. 3B shows motion vector resolutions of 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel units when the minimum motion vector resolution is 1 / 4 pixel unit. (a), (b), (c), and (d) in FIG. 3B show the coordinates of the pixels (displayed as black squares) that can indicate the motion vectors with resolutions of 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel units, respectively, based on the coordinate (0, 0).

[0075] The motion vector encoding device 10 according to an embodiment can search for a block similar to the current block in the reference picture based on the sub-pixel unit with respect to the motion vector determined at the integer pixel position in order to perform motion compensation in units of sub-pixels.

[0076] For example, the motion vector encoding device 10 determines a motion vector at the integer pixel position, doubles the resolution of the reference picture, and searches for the most similar prediction block in the range of (-1, -1) to (1, 1) based on the motion vector determined at the integer pixel position. Next, by further doubling the resolution and searching for the most similar prediction block in the range of (-1, -1) to (1, 1) at a resolution of 4 times based on the motion vector at the 1 / 2 pixel position, the motion vector at the final resolution of 1 / 4 pixel can be determined.

[0077] For example, when the motion vector at an integer pixel position is (-4, -3) with respect to the coordinate (0, 0), at a resolution of 1 / 2 pixel, the motion vector becomes (-8, -6). If it moves by about (0, -1), the motion vector at a resolution of 1 / 2 pixel is finally determined to be (-8, -7). Also, the motion vector at a resolution of 1 / 4 pixel is changed to (-16, -14), and if it moves by about (-1, 0), the final motion vector at a resolution of 1 / 4 pixel is determined to be (-17, -14).

[0078] The motion vector encoding device 10 according to an embodiment can search for a block similar to the current block within the reference picture based on a pixel position larger than the 1-pixel position, with the motion vector determined at the integer pixel position, in order to perform motion compensation in a pixel unit larger than the 1-pixel unit. Hereinafter, a pixel position larger than the 1-pixel position (for example, 2 pixels, 3 pixels, 4 pixels) is referred to as a super pixel.

[0079] For example, when the motion vector at an integer pixel position is (-4, -8) with respect to the coordinate (0, 0), at a resolution of 2 pixels, the motion vector is determined to be (-2, -4). To encode the motion vector in units of 1 / 4 pixel, more bits are consumed compared to the motion vector in units of integer pixels, but an accurate inter prediction in units of 1 / 4 pixel can be performed, and the number of bits consumed for residual block encoding can be reduced.

[0080] However, if interpolation is performed in a pixel unit smaller than 1 / 4 pixel unit, for example, 1 / 8 pixel unit, to generate sub-pixels and estimate the motion vector in units of 1 / 8 pixel based on them, an excessive number of bits are consumed for encoding the motion vector, and rather, the compression ratio of encoding decreases.

[0081] Also, when there is a lot of noise in the video or when the texture is scarce, the resolution can be set in super pixel units to perform motion estimation, and the compression ratio of encoding can be improved.

[0082] FIG. 4A shows candidate blocks of the current block for obtaining predicted motion vector candidates.

[0083] The prediction unit 11 can obtain candidates for one or more predicted motion vectors of the current block in order to perform motion prediction related to the current block on the reference picture of the current block to be encoded. The prediction unit 11 can obtain at least one of a spatial candidate block and a temporal candidate block of the current block in order to obtain a candidate for the predicted motion vector.

[0084] When the current block is predicted with reference to reference frames having different POCs, the prediction unit 11 can use a block located around the current block, a co-located block belonging to a reference frame that is temporally different (with a different POC) from the current block, and a block around the co-located block to obtain a predicted motion vector candidate.

[0085] For example, the spatial candidate block may include at least one of a left block A1 411, an upper block B1 412, an upper left block B2 413, an upper right block B0 414, and a lower left block A0 425, which are adjacent blocks of the current block 410. The temporal candidate block may include at least one of a co-located block 430 belonging to a reference frame having a different POC from the current block, and an adjacent block H 431 of the co-located block 430. The motion vector encoding device 10 can obtain motion vectors of the temporal candidate block and the spatial candidate block as predicted motion vector candidates.

[0086] The prediction unit 11 can obtain prediction motion vector candidates with a plurality of predetermined motion vector resolutions. Each prediction motion vector candidate can have a different resolution from one another. That is, the prediction unit 11 can use the first prediction motion vector candidate among the prediction motion vector candidates to search for a reference block in pixel units of the first motion vector resolution, and use the second prediction motion vector candidate to search for a reference block in pixel units of the second motion vector resolution. The first prediction motion vector candidate and the second prediction motion vector candidate are obtained using different blocks among the blocks belonging to the spatial candidate block and the temporal candidate block.

[0087] The prediction unit 11 can determine the number and type of the set of candidate blocks (that is, the set of prediction motion vector candidates) to be different according to the resolution of the motion vector.

[0088] For example, when the minimum motion vector resolution is 1 / 4 pixel and 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel units are used as the plurality of predetermined motion vector resolutions, the prediction unit 11 can generate a predetermined number of prediction motion vector candidates for each resolution. The prediction motion vector candidates for each resolution are also the motion vectors of different candidate blocks. By using different candidate blocks for each resolution and obtaining the prediction motion vector candidates, the prediction unit 11 can increase the probability of searching for the optimal reference block, and improve the coding efficiency by reducing the rate-distortion cost.

[0089] To determine the motion vector of the current block, the prediction unit 11 can use each predicted motion vector candidate to determine the search start position in the reference picture and search for the optimal reference block based on the resolution of each predicted motion vector candidate. That is, if the predicted motion vector candidates of the current block are obtained, the motion vector encoding device 10 searches for the reference block in pixel units of a predetermined resolution corresponding to each predicted motion vector candidate, compares the rate-distortion costs based on the difference values between the motion vector of the current block and each predicted motion vector, and can determine the predicted motion vector having the minimum cost.

[0090] The encoding unit 13 can encode the residual motion vector, which is the difference vector between the determined one predicted motion vector and the actual motion vector of the current block, and the information indicating the motion vector resolution used for the inter prediction of the current block.

[0091] The encoding unit 13 can determine and encode the residual motion vector as shown in Equation (1). MVx is the x component of the actual motion vector of the current block, and MVy is the y component of the actual motion vector of the current block. pMVx is the x component of the predicted motion vector of the current block, and pMVy is the y component of the predicted motion vector of the current block. MVDx is the x component of the residual motion vector of the current block, and MVDy is the y component of the residual motion vector of the current block.

[0092] MVDx = MVx - pMVx MVDy = MVy - pMVy (1) The decoding unit 23 can restore the motion vector of the current block by using the information indicating the predicted motion vector obtained from the bitstream and the residual motion vector. The decoding unit 23 can determine the final motion vector by adding up the predicted motion vector and the residual motion vector as shown in Equation (2).

[0093] MVx = pMVx + MVDx MVy = pMCy + MVDy (2) If the minimum motion vector resolution is in sub-pixel units, the motion vector encoding device 10 can multiply the predicted motion vector and the actual motion vector by an integer value and represent the motion vector by the integer value. If the predicted motion vector with a resolution of 1 / 4 pixel unit starting from the coordinates (0, 0) indicates the coordinates (1 / 2, 3 / 2), and the minimum motion vector resolution is 1 / 4 pixel unit, the motion vector encoding device 10 can encode, as the predicted motion vector, a vector (2, 6) which is the value obtained by multiplying the predicted motion vector by the integer 4. If the minimum motion vector resolution is 1 / 8 pixel unit, the predicted motion vector can be multiplied by the integer 8, and a vector (4, 12) can be encoded as the predicted motion vector.

[0094] FIG. 4B shows the generation process of the predicted motion vector candidates according to an embodiment.

[0095] As described above, the prediction unit 11 can obtain predicted motion vector candidates with a plurality of predetermined motion vector resolutions.

[0096] For example, when the minimum motion vector resolution is 1 / 4 pixel and the plurality of predetermined motion vector resolutions include 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixels, the prediction unit 11 can generate a predetermined number of predicted motion vector candidates for each resolution. The prediction unit 11 can configure the sets 460, 470, 480, and 490 of the predicted motion vector candidates to be different from each other according to the plurality of predetermined motion vector resolutions. Each of the sets 460, 470, 490, and 490 may include predicted motion vectors taken from different candidate blocks and may include different numbers of predicted motion vectors. By using different candidate blocks for each resolution, the prediction unit 11 can increase the probability of searching for the optimal prediction block and improve the encoding efficiency by reducing the rate-distortion cost.

[0097] The predicted motion vector candidate 460 with a resolution of 1 / 4 pixel unit is determined from two different temporal or spatial candidate blocks. For example, the prediction unit 11 obtains the motion vector of the left block 411 of the current block and the motion vector of the upper block 412 as the predicted motion vector candidate 460 with a resolution of 1 / 4 pixel, and can search for the optimal motion vector for each predicted motion vector candidate. That is, the prediction unit 11 determines the search start position using the motion vector of the left block 411 of the current block, searches for the reference block in 1 / 4 pixel units, and can determine the optimal predicted motion vector and the reference block. In addition, the prediction unit 11 determines the search start position using the motion vector of the upper block 412, searches for the reference block in 1 / 4 pixel units, and can determine another optimal predicted motion vector and the reference block.

[0098] The predicted motion vector candidate 470 with a resolution of 1 / 2 pixel unit is determined from one temporal or spatial candidate block. The predicted motion vector candidate with a resolution of 1 / 2 pixel unit is also a predicted motion vector different from the predicted motion vector candidate with a resolution of 1 / 4 pixel unit. For example, the prediction unit 11 can obtain the motion vector of the upper right block 414 of the current block as the predicted motion vector candidate 470 with a resolution of 1 / 2 pixel. That is, the prediction unit 11 determines the search start position using the motion vector of the upper right block 414, searches for the reference block in 1 / 2 pixel units, and can determine another optimal predicted motion vector and the reference block.

[0099] The predicted motion vector candidate 480 with a resolution of one pixel unit is determined from one temporal or spatial candidate block. The predicted motion vector with a resolution of one pixel unit is also a predicted motion vector different from the predicted motion vectors 460 and 470 used at resolutions of 1 / 4 pixel unit and 1 / 2 pixel unit. For example, the prediction unit 11 can determine the motion vector of the temporal candidate block 430 as the predicted motion vector candidate 480 of one pixel. That is, the prediction unit 11 can use the motion vector of the temporal candidate block 430 to determine the search start position, search for the reference block in units of one pixel, and determine other optimal predicted motion vectors and reference blocks.

[0100] The predicted motion vector candidate 490 with a resolution of two pixel units is determined from one temporal or spatial candidate block. The predicted motion vector with a resolution of two pixel units is also a predicted motion vector different from the predicted motion vectors used at other resolutions. For example, the prediction unit 11 can determine the motion vector of the lower left block 425 as the predicted motion vector 490 of two pixels. That is, the prediction unit 11 can use the motion vector of the lower left block 425 to determine the search start position, search for the reference block in units of two pixels, and determine other optimal predicted motion vectors and reference blocks.

[0101] The prediction unit 11 can compare the rate-distortion cost based on the motion vector of the current block and each predicted motion vector, and finally determine the motion vector of the current block, one predicted motion vector, and one motion vector resolution 495.

[0102] When the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the motion vector encoding device 10 can downscale and encode the residual motion vector according to the resolution of the motion vector of the current block. Also, when the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the motion vector decoding device 20 can upscale and restore the residual motion vector according to the minimum motion vector resolution.

[0103] When the minimum motion vector resolution is 1 / 4 pixel unit and the resolution of the motion vector of the current block is determined to be 1 / 2 pixel unit, the encoding unit 13 can calculate the residual motion vector by adjusting the determined actual motion vector and the predicted motion vector to 1 / 2 pixel unit in order to reduce the magnitude of the residual motion vector.

[0104] The encoding unit 13 can reduce the magnitudes of the actual motion vector and the predicted motion vector by half, and also reduce the magnitude of the residual motion vector by half. That is, when the minimum motion vector resolution is 1 / 4 pixel unit, the predicted motion vector (MVx, MVy) expressed by multiplying by 4 can be further divided by 2 to represent the predicted motion vector. For example, when the minimum motion vector resolution is 1 / 4 pixel unit and the predicted motion vector is (-24, -16), the motion vector at the resolution of 1 / 2 pixel unit will be (-12, -8), and the motion vector at the resolution of 2 pixel units will be (-3, -2). The following formula (3) shows the process of reducing the magnitudes of the actual motion vector (MVx, MVy) and the predicted motion vector (pMVx, pMCy) by half using bit shift operations.

[0105] MVDx=(MVx)>>1-(pMVx)>>1 MVDy=(MVy)>>1-(pMCy)>>1 (3) The decoding unit 23 can determine the final motion vector (MVx, MVy) related to the current block as shown in Equation (4) by adding the received residual motion vector (MVDx, MVDy) to the finally determined predicted motion vector (pMVx, pMCy). The decoding unit 23 can upscale the predicted motion vector obtained as shown in Equation (4) and the residual motion vector using a bit shift operation to restore the motion vector related to the current block.

[0106] MVx=(pMVx)<<1+(MVDx)<<1 MVy=(pMCy)<<1+(MVDy)<<1 (4) According to one embodiment, if the minimum motion vector resolution is 1 / 2 n pixel, and for the current block, when a motion vector with a resolution of 2 k pixels per unit is determined, the residual motion vector can be determined using the following Equation (5).

[0107] MVDx=(MVx)>>(k+n)-(pMVx)>>(k+n) MVDy=(MVy)>>(k+n)-(pMCy)>>(k+n) (5) The decoding unit 23 can determine the final motion vector of the current block as shown in Equation (6) by adding the received residual motion vector to the finally determined predicted motion vector.

[0108] MVx=(pMVx)<<(k+n)+(MVDx)<<(k+n) MVy=(pMCy)<<(k+n)+(MVDy)<<(k+n) (6) If the magnitude of the residual motion vector decreases, the number of bits representing the residual motion vector decreases, and the coding efficiency improves.

[0109] As described with reference to FIGS. 3A and 3B, 2 kIn the motion vector for each pixel, when k is less than 0, the reference picture performs interpolation to generate pixels at non-integer positions. Conversely, when the motion vector is such that k is greater than 0, in the reference picture, only pixels existing at positions that are multiples of 2 k are searched for. Therefore, when the motion vector resolution of the current block is greater than or equal to 1 pixel unit, the decoding unit 23 can omit the interpolation of the reference picture according to the motion vector resolution of the current block to be decoded.

[0110] FIG. 5A shows an encoding unit and a prediction unit according to an embodiment.

[0111] If the current block is the current encoding unit that constitutes the video, and the motion vector resolution for inter prediction is determined to be the same for each encoding unit, and if there is one or more prediction units predicted to be in the AMVP mode within the current encoding unit, the encoding unit 13 can encode only once the information indicating the motion vector resolution of the current block and the information indicating the motion vector resolution of the prediction units predicted to be in the AMVP mode, and transmit it to the motion vector decoding device 20. The acquisition unit 21 can acquire only once the information indicating the motion vector resolution of the current block and the information indicating the motion vector resolution of the prediction units predicted to be in the AMVP mode from the bit stream.

[0112] FIG. 5B shows a part of the prediction_unit syntax according to an embodiment for transmitting the adaptively determined motion vector resolution. FIG. 5B is a syntax that defines the operation in which the motion vector decoding device 20 according to an embodiment acquires the information indicating the motion vector resolution of the current block.

[0113] For example, as illustrated in FIG. 5A, when the current coding unit 560 has a size of 2Nx2N, the prediction unit 563 has the same size (2Nx2N) as the coding unit 560, and is predicted using the AMVP mode, the coding unit 13 encodes the information indicating the motion vector resolution of the coding unit 560 once for the current coding unit 560, and the acquisition unit 21 can acquire the information indicating the motion vector resolution of the coding unit 560 once from the bitstream.

[0114] When the size of the current coding unit 570 is 2Nx2N and it is divided into two prediction units 573 and 577 of size 2NxN, since the prediction unit 573 is predicted to be in the merge mode, the coding unit 13 does not transmit the information indicating the motion vector resolution related to the prediction unit 573. Since the prediction unit 577 is predicted to be in the AMVP mode, the coding unit 13 transmits the information indicating the motion vector resolution related to the prediction unit 573 once, and the acquisition unit 21 can acquire the information indicating the motion vector resolution related to the prediction unit 577 once from the bitstream. That is, the coding unit 13 transmits the information indicating the motion vector resolution related to the prediction unit 577 once as the information indicating the motion vector resolution of the current coding unit 570, and the decoding unit 23 can receive the information indicating the motion vector resolution related to the prediction unit 577 once as the information indicating the motion vector resolution of the current coding unit 570. The information indicating the motion vector resolution can also be in the form of an index indicating any one of a plurality of predetermined motion vectors, such as "cu_resolution_idx[x0][y0]".

[0115] Referring to the syntax of FIG. 5B, for the current coding unit, the initial values of "parsedMVResolution 510", which is information indicating whether the motion vector resolution has been extracted for the current coding unit, and "mv_resolution_idx 512", which is information indicating the motion vector resolution of the current coding unit, can be set to 0 respectively. First, in the case of prediction unit 573, since it is predicted to be in the merge mode and does not satisfy condition 513, acquisition unit 21 does not receive the information "cu_resolution_idx[x0][y0]" indicating the motion vector resolution of the current prediction unit.

[0116] In the case of prediction unit 577, since it is predicted to be in the AMVP mode, it satisfies condition 513, and since "parsedMVResolution" has a value of 0, it satisfies condition 514, and acquisition unit 21 can receive "cu_resolution_idx[x0][y0] 516". Since "cu_resolution_idx[x0][y0]" is received, "parsedMVResolution" is set to 1 (518). The received "cu_resolution_idx[x0][y0]" is stored in "mv_resolution_idx" (520).

[0117] When the size of the current coding unit 580 is 2Nx2N and it is divided into two prediction units 583 and 587 of size 2NxN, and both prediction unit 583 and prediction unit 587 are predicted to be in the AMVP mode, the coding unit 13 encodes and transmits the information indicating the motion vector resolution of the current coding unit 580 once, and the acquisition unit 21 can receive the information indicating the motion vector resolution of the coding unit 580 from the bitstream once.

[0118] Referring to the syntax of FIG. 5B, in the case of prediction unit 583, since condition 513 is satisfied and the value of "parsedMVResolution" is 0, condition 514 is satisfied, and acquisition unit 21 can acquire "cu_resolution_idx[x0][y0]516". Since decoding unit 23 has acquired "cu_resolution_idx[x0][y0]", it sets "parsedMVResolution" to 1 (518) and stores the acquired "cu_resolution_idx[x0][y0]" in "mv_resolution_idx" (520). However, in the case of prediction unit 587, since "parsedMVResolution" already has a value of 1, condition statement 514 cannot be satisfied, so acquisition unit 21 does not acquire "cu_resolution_idx[x0][y0]516". That is, since acquisition unit 21 has already received the information indicating the motion vector resolution related to the current coding unit 580 from prediction unit 583, there is no need to acquire the information indicating the motion vector resolution (the same as the motion vector resolution of prediction unit 583) from prediction unit 587.

[0119] If the size of the current coding unit 590 is 2Nx2N and it is divided into two prediction units 593 and 597 of size 2NxN, and both prediction unit 593 and prediction unit 597 are predicted to be in merge mode, then condition statement 513 is not satisfied, so acquisition unit 21 does not acquire the information indicating the motion vector resolution for the current coding unit 590.

[0120] According to another embodiment, the encoding unit 13 has a current block which is a current encoding unit constituting a video, and for each prediction unit, the motion vector resolution for inter prediction is determined to be the same. If there is one or more prediction units predicted to be in the AMVP mode within the current encoding unit, the encoding unit 13 can transmit, as information indicating the motion vector resolution of the current block, information indicating the motion vector resolution for each prediction unit predicted to be in the AMVP mode existing in the current block to the motion vector decoding device 20. The acquisition unit 21 can acquire, from the bitstream, information indicating the motion vector resolution for each prediction unit existing in the current block as information indicating the motion vector resolution of the current block.

[0121] FIG. 5C shows a part of the prediction_unit syntax according to another embodiment for transmitting the adaptively determined motion vector resolution. FIG. 5C is a syntax that defines the operation of the motion vector decoding device 20 according to another embodiment to acquire information indicating the motion vector resolution of the current block.

[0122] The difference between the syntax of FIG. 5C and the syntax of FIG. 5B is that there is no "parsedMVResolution 510" indicating whether the motion vector resolution has been extracted for the current encoding unit. Therefore, the acquisition unit 21 acquires (524) information "cu_resolution_idx[x0,y0]" indicating the motion vector resolution for each prediction unit existing in the current encoding unit.

[0123] For example, referring to FIG. 5A again, the size of the current coding unit 580 is 2Nx2N and is divided into two prediction units 583 and 587 of size 2NxN. When both the prediction unit 583 and the prediction unit 587 are predicted to be in the AMVP mode, and the motion vector resolution of the prediction unit 583 is 1 / 4 pixel and the motion vector resolution of the prediction unit 587 is 2 pixels, the acquisition unit 21 acquires information indicating the motion vector resolution of 1 / 2 of the motion vector resolution of two motion vector resolutions (i.e., the prediction unit 583) and 1 / 4 of the motion vector resolution of the prediction unit 587 for one coding unit 580 (524).

[0124] When the coding unit 13 according to another embodiment adaptively determines the motion vector resolution for the current prediction unit regardless of the prediction mode of the prediction unit, the information indicating the motion vector resolution is coded and transmitted for each prediction unit regardless of the prediction mode, and the acquisition unit 11 can acquire the information indicating the motion vector resolution from the bit stream for each prediction unit regardless of the prediction mode.

[0125] FIG. 5D shows a part of the prediction_unit syntax according to another embodiment for transmitting the adaptively determined motion vector resolution. FIG. 5D is a syntax that defines the operation in which the motion vector decoding device 20 according to another embodiment acquires information indicating the motion vector resolution of the current block.

[0126] Although it is assumed that the syntax described in FIGS. 5B and 5C is applicable only when the adaptive motion vector resolution determination method is limited to the case where the prediction mode is the AMVP mode, FIG. 5D is an example of a syntax assuming the case where the adaptive motion vector resolution determination method is applicable regardless of the prediction mode. Referring to the syntax of FIG. 5D, the acquisition unit 21 receives "cu_resolution_idx[x0][y0]" for each prediction unit regardless of the prediction mode of the prediction unit (534).

[0127] The index "cu_resolution_idx[x0][y0]" indicating the motion vector resolution described with reference to FIGS. 5B to 5D is encoded and transmitted in unary or fixed length. Also, when only two motion vector resolutions are used, "cu_resolution_idx[x0][y0]" is also data in flag format.

[0128] The motion vector encoding device 10 can adaptively configure a plurality of predetermined motion vector resolutions used for encoding in units of slices or blocks. Also, the motion vector decoding device 10 can adaptively configure a plurality of predetermined motion vector resolutions used for decoding in units of slices or blocks. The plurality of predetermined motion vector resolutions adaptively configured in units of slices or blocks can be used as a motion vector resolution candidate group. That is, the motion vector encoding device 10 and the motion vector decoding device 20 can be configured to have different types and numbers of motion vector candidate groups for the current block based on the information of the already encoded or decoded peripheral blocks.

[0129] For example, when the minimum motion vector resolution of the motion vector encoding device 10 or the motion vector decoding device 20 is 1 / 4 pixel unit, as a motion vector resolution candidate group fixed identically for all videos, resolutions of 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel units can be used. Instead of using a motion vector resolution candidate group fixed identically for all videos, the motion vector encoding device 10 or the motion vector decoding device 20, when the motion vector resolution of the already encoded peripheral block is small, can use 1 / 8 pixel, 1 / 4 pixel, and 1 / 2 pixel units as the motion vector resolution candidate group of the current block, and when the motion vector resolution of the peripheral block is large, can use 1 / 2 pixel, 1 pixel, and 2 pixel units as the motion vector resolution candidate group of the current block. The motion vector encoding device 10 or the motion vector decoding device 20 can also be configured to vary the type and number of the motion vector resolution candidate groups in units of slices or blocks based on the magnitude of the motion vector and other information.

[0130] The type and number of resolutions constituting the motion vector resolution candidate group can be set and used in exactly the same way by the motion vector encoding device 10 and the motion vector decoding device 20 all the time, or can be analogized in the same way based on the information of the peripheral block and other information. Alternatively, the information related to the motion vector resolution candidate group used by the motion vector encoding device 10 can be encoded in the bitstream and clearly transmitted to the motion vector decoding device 20.

[0131] FIG. 6A shows an embodiment that uses a plurality of resolutions to form a merge candidate list.

[0132] In order to reduce the amount of data related to motion information transmitted for each prediction unit, the motion vector encoding device 10 can utilize the merge mode that is set as the motion information of the current block based on the motion information of spatial / temporal neighboring blocks. The motion vector encoding device 10 constructs the same merge candidate list in both the encoding device and the decoding device for predicting motion information, and by transmitting the candidate selection information in the list to the decoding device, the amount of motion-related data can be effectively reduced.

[0133] When there is a prediction unit for which the current block is predicted using the merge mode, the motion vector decoding device 20 constructs a merge candidate list related to the current block in the same way as the motion vector encoding device 10, obtains the candidate selection information in the list from the bitstream, and can decode the motion vector of the current block.

[0134] The merge candidate list may include spatial candidates based on the motion information of spatial neighboring blocks and temporal candidates based on the motion information of temporal neighboring blocks. The motion vector encoding device 10 and the motion vector decoding device 20 can include spatial candidates and temporal candidates of a plurality of predetermined motion vector resolutions in the merge candidate list in a predetermined order.

[0135] The method for determining the motion vector resolution described with reference to FIGS. 3A to 4B is not applicable only when the prediction unit of the current block is encoded using the AMVP mode, and is also applicable when using a prediction mode (for example, the merge mode) that immediately uses one of the predicted motion vector candidates as the final motion vector without transmitting the residual motion vector.

[0136] That is, the motion vector encoding device 10 and the motion vector decoding device 20 can adjust the predicted motion vector candidates to be suitable for a plurality of predetermined motion vector resolutions, and determine the adjusted predicted motion vector candidates as the motion vector of the current block. That is, the motion vectors of the candidate blocks included in the merge candidate list may include the motion vectors with the minimum motion vector resolution and the motion vectors downscaled by a plurality of predetermined motion vector resolutions. The method of downscaling will be described later with reference to FIGS. 7A and 7B.

[0137] For example, assume that the minimum motion vector resolution is 1 / 4 pixel, the plurality of predetermined motion vector resolution candidates are 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixels, and the merge candidate list is composed of (A1, B1, B0, A0, B2, co-located block). As shown in FIG. 6A, the acquisition unit 21 of the motion vector decoding device 20 constructs predicted motion vector candidates at 1 / 4 resolution, adjusts the predicted motion vectors at 1 / 4 resolution to the resolutions of 1 / 2 pixel, 1 pixel, and 2 pixels, and can sequentially acquire merge candidate lists of multiple resolutions according to the order of the resolutions.

[0138] FIG. 6B shows another embodiment of constructing a merge candidate list using multiple resolutions.

[0139] As shown in FIG. 6B, the motion vector encoding device 10 and the motion vector decoding device 20 can also construct predicted motion vector candidates at the resolution of 1 / 4 pixel unit, adjust each predicted motion vector candidate at the resolution of 1 / 4 pixel unit to the resolutions of 1 / 2 pixel, 1 pixel, and 2 pixel units, and sequentially acquire the merge candidate list according to the order of the predicted motion vector candidates.

[0140] When there is a prediction unit for which the current block predicts using the merge mode, the motion vector decoding device 20 can determine the motion vector related to the current block based on the merge candidate lists of multiple resolutions acquired by the acquisition unit 21 and the information related to the merge candidate index acquired from the bit stream.

[0141] FIG. 7A shows pixels indicated by two motion vectors with different resolutions.

[0142] The motion vector encoding device 10 can adjust the motion vector so as to indicate a surrounding pixel instead of a pixel indicated by an existing high-resolution motion vector in order to adjust (adjust) the high-resolution motion vector to a corresponding low-resolution motion vector. Selecting any one of the surrounding pixels is rounding.

[0143] For example, with reference to the coordinates (0, 0), to adjust a motion vector with a resolution of 1 / 4 pixel indicating (19, 27) to a motion vector with a resolution of 1 pixel unit, the motion vector (19, 27) with a resolution of 1 / 4 pixel is divided by the integer 4, and rounding occurs during the division process. For the sake of convenience of explanation, hereinafter, it is assumed that the motion vector of each resolution starts from the coordinates (0, 0) and indicates the coordinates (x, y) (x and y are integers).

[0144] Referring to FIG. 7A, the resolution of the minimum motion vector is 1 / 4 pixel unit. To adjust the motion vector 715 with a resolution of 1 / 4 pixel unit to a motion vector with a resolution of 1 pixel unit, the four integer pixels 720, 730, 740, 750 surrounding the pixel 710 indicated by the 1 / 4 pixel motion vector 715 also become candidate pixels indicated by the corresponding 1 pixel motion vectors 725, 735, 745, 755. That is, if the value of the coordinates 710 is (19, 27), the coordinates 1020 will be (7, 24), the coordinates 730 will be (16, 28), the coordinates 740 will be (20, 28), and the coordinates 750 will also be (20, 24).

[0145] When the motion vector encoding device 10 according to an embodiment adjusts a motion vector 710 with a resolution of 1 / 4 pixel unit to a corresponding motion vector with a resolution of 1 pixel unit, it can be determined to indicate the integer pixel 740 at the upper right end. That is, if the motion vector with a resolution of 1 / 4 pixel unit starts from the coordinates (0, 0) and points to the coordinates (19, 27), the corresponding motion vector with a resolution of 1 pixel unit starts from the coordinates (0, 0) and points to the coordinates (20, 28), and the final motion vector with a resolution of 1 pixel unit also becomes (5, 7).

[0146] When the motion vector encoding device 10 according to an embodiment adjusts a high-resolution motion vector to a low-resolution motion vector, the adjusted low-resolution motion vector always points to the upper right end of the pixel indicated by the high-resolution motion vector. The motion vector encoding device 10 according to another embodiment causes the adjusted low-resolution motion vector to always point to the upper left end, the lower left end, or the lower right end pixel of the pixel indicated by the high-resolution motion vector.

[0147] The motion vector encoding device 10 can differently select the pixel indicated by the corresponding low-resolution motion vector from among the four pixels at the upper left end, the upper right end, the lower left end, and the lower right end located around the pixel indicated by the high-resolution motion vector according to the resolution of the motion vector of the current block for the high-resolution motion vector.

[0148] For example, referring to FIG. 7B, the 1 / 2 pixel motion vector can be adjusted to point to the pixel 1080 at the upper left end of the pixel 1060 indicated by the 1 / 4 pixel motion vector, the 1 pixel motion vector can be adjusted to point to the pixel 1070 at the upper right end of the pixel indicated by the 1 / 4 pixel motion vector, and the 2 pixel motion vector can be adjusted to point to the pixel 1090 at the lower right end of the pixel indicated by the 1 / 4 pixel motion vector.

[0149] When the motion vector encoding device 10 is to indicate any one of the peripheral pixels instead of the pixel indicated by the existing high-resolution motion vector, the position of the indicated pixel can be determined based on at least one of the resolution, the motion vector candidate of 1 / 4 pixel, the information of the peripheral block, the encoding information, and an arbitrary pattern.

[0150] On the other hand, for the sake of convenience of explanation, in FIGS. 3A to 7B, only the operations performed by the motion vector encoding device 10 are described, and the operations in the motion vector decoding device 20 are omitted or only the operations performed by the motion vector decoding device 20 are described and the operations in the motion vector encoding device 10 are omitted. However, it will be easily understood by those of ordinary skill in the technical field to which the present embodiment belongs that corresponding operations are also performed in each of the motion vector encoding device 10 and the motion vector decoding device 20, and in each of the motion vector decoding device 20 and the motion vector encoding device 10.

[0151] Hereinafter, with reference to FIGS. 8 to 20, a video encoding method and its apparatus, and a video decoding method and its apparatus based on a tree-structured encoding unit and a conversion unit according to an embodiment are disclosed. The motion vector encoding device 10 described with reference to FIGS. 1A to 7B may be included in the video encoding device 800. That is, the motion vector encoding device 10 can encode the information indicating the predicted motion vector for the inter prediction of the video encoded by the video encoding device 800, the residual motion vector, and the information indicating the motion vector resolution by the method described with reference to FIGS. 1A to 7B.

[0152] FIG. 8 illustrates a block diagram of a video encoding device 800 based on an encoding unit with a tree structure according to an embodiment of the present invention.

[0153] A video encoding device 800 with video prediction based on an encoding unit with a tree structure according to an embodiment includes an encoding unit determination unit 820 and an output unit 830. Hereinafter, for convenience of explanation, a video encoding device 800 with video prediction based on an encoding unit with a tree structure according to an embodiment is referred to in short as the "video encoding device 800".

[0154] The encoding unit determination unit 820 can partition the current picture based on the maximum encoding unit, which is the maximum-size encoding unit for the current picture of the video. If the current picture is larger than the maximum encoding unit, the video data of the current picture is divided into at least one maximum encoding unit. The maximum encoding unit according to an embodiment is a data unit such as 32x32, 64x64, 128x128, 256x256 in size, and is also a square data unit with a size that is a power of 2 in both the vertical and horizontal directions.

[0155] The encoding unit according to an embodiment is characterized by a maximum size and a depth. The depth indicates the number of times the encoding unit is spatially divided from the maximum encoding unit. The deeper the depth, the more the encoding unit by depth is divided from the maximum encoding unit to the minimum encoding unit. The depth of the maximum encoding unit is defined as the top depth, and the minimum encoding unit is defined as the bottom encoding unit. Since the size of the encoding unit by depth becomes smaller as the depth of the maximum encoding unit increases, the encoding unit of the upper depth may include a plurality of encoding units of the lower depth.

[0156] As described above, the video data of the current picture is divided into maximum encoding units according to the maximum size of the encoding unit, and each maximum encoding unit may include encoding units that are divided by depth. Since the maximum encoding unit according to an embodiment is divided by depth, the video data in the spatial domain included in the maximum encoding unit is hierarchically classified by depth.

[0157] The maximum depth that limits the total number of times the height and width of the maximum encoding unit can be hierarchically divided, and the maximum size of the encoding unit are set in advance.

[0158] The segmentation unit determination unit 820 encodes at least one segmented area into which the area of the maximum coding unit is segmented for each depth, and determines the depth at which the final coding result is output for each at least one segmented area. That is, the segmentation unit determination unit 820 encodes video data with coding units by depth for each maximum coding unit of the current picture, selects the depth at which the minimum coding error occurs, and determines it as the final depth. The determined final depth and the video data for each maximum coding unit are output to the output unit 830.

[0159] The video data within the maximum coding unit is encoded based on coding units by depth with at least one depth below the maximum depth, and the coding results based on the respective coding units by depth are compared. Based on the comparison result of the coding errors of the coding units by depth, the depth with the minimum coding error is selected. For each respective maximum coding unit, at least one final depth is determined.

[0160] The size of the maximum coding unit is hierarchically segmented as the depth increases, and the number of coding units increases. Also, even if they are coding units of the same depth included in one maximum coding unit, the coding error related to each data is measured, and the division into lower depths is determined. Therefore, even for the data included in one maximum coding unit, the coding error by depth differs depending on the position, so the final depth is determined differently depending on the position. Therefore, for one maximum coding unit, the final depth is set to 1 or more, and the data of the maximum coding unit is partitioned by coding units with 1 or more final depths.

[0161] Therefore, according to one embodiment, the encoding unit determination unit 820 determines the encoding unit based on the tree structure included in the current maximum encoding unit. The "encoding unit based on the tree structure" according to one embodiment includes the encoding units at the final depth and the determined depth among all the depth-based encoding units included in the current maximum encoding unit. The encoding unit at the final depth is hierarchically determined by the depth within the same region in the maximum encoding unit, and is independently determined for other regions. Similarly, the final depth related to the current region is independently determined from the final depth related to other regions.

[0162] The maximum depth according to one embodiment is an index related to the number of divisions from the maximum encoding unit to the minimum encoding unit. The first maximum depth according to one embodiment can indicate the total number of divisions from the maximum encoding unit to the minimum encoding unit. The second maximum depth according to one embodiment can indicate the total number of depth levels from the maximum encoding unit to the minimum encoding unit. For example, when the depth of the maximum encoding unit is 0, the depth of the encoding unit obtained by dividing the maximum encoding unit once is set to 1, and the depth of the encoding unit obtained by dividing it twice is set to 2. In that case, if the encoding unit obtained by dividing the maximum encoding unit four times is the minimum encoding unit, there are depth levels of 0, 1, 2, 3, and 4, so the first maximum depth is set to 4, and the second maximum depth is set to 5.

[0163] Predictive encoding and transformation of the maximum encoding unit are performed. Similarly, predictive encoding and transformation are also performed for each maximum encoding unit, for each depth below the maximum depth, based on the depth-based encoding units.

[0164] Each time the maximum encoding unit is divided by depth, the number of depth-based encoding units increases. Therefore, encoding including predictive encoding and transformation must be performed for all the depth-based encoding units generated as the depth increases. Hereinafter, for the sake of convenience of explanation, predictive encoding and transformation will be described based on the encoding unit at the current depth among at least one maximum encoding unit.

[0165] The video encoding device 800 according to one embodiment can variously select the size or form of the data unit for encoding video data. For encoding video data, steps such as predictive encoding, transformation, and entropy encoding are involved. Throughout all steps, the same data unit is used, and the data unit may also be changed step by step.

[0166] For example, the video encoding device 800 can select a data unit different from the encoding unit not only for the encoding unit for encoding video data but also for performing predictive encoding of the video data of the encoding unit.

[0167] For predictive encoding of the maximum encoding unit, predictive encoding is performed based on the encoding unit of the final depth according to one embodiment, that is, the encoding unit that is not further divided. Hereinafter, the encoding unit that is not further divided and serves as the basis for predictive encoding is referred to as a "prediction unit". The partition obtained by dividing the prediction unit may include the prediction unit and a data unit in which at least one of the height and width of the prediction unit is divided. The partition is a data unit in the form obtained by dividing the prediction unit of the encoding unit, and the prediction unit is also a partition of the same size as the encoding unit.

[0168] For example, when the encoding unit of size 2Nx2N (where N is a positive integer) is not further divided, it becomes a prediction unit of size 2Nx2N, and the size of the partition can also be 2Nx2N, 2NxN, Nx2N, NxN, etc. The partition mode according to one embodiment may selectively include not only symmetric partitions in which the height or width of the prediction unit is divided at a symmetric ratio but also partitions divided at an asymmetric ratio such as 1:n or n:1, partitions divided into geometric forms, partitions of arbitrary forms, etc.

[0169] The prediction mode of the prediction unit is at least one of the intra mode, the inter mode, and the skip mode. For example, the intra mode and the inter mode are performed for partitions of sizes 2Nx2N, 2NxN, Nx2N, and NxN. Also, the skip mode is performed only for partitions of size 2Nx2N. For each prediction unit within the coding unit, coding is performed independently, and the prediction mode with the minimum coding error is selected.

[0170] Also, the video coding device 800 according to an embodiment can perform conversion of the video data of the coding unit based not only on the coding unit for coding of the video data but also on a data unit different from the coding unit. For conversion of the coding unit, the conversion is performed based on a conversion unit smaller than or the same size as the coding unit. For example, the conversion unit may include a data unit for the intra mode and a conversion unit for the inter mode.

[0171] In a manner similar to the coding unit with a tree structure according to an embodiment, the conversion unit within the coding unit is also recursively divided into smaller-sized conversion units, and the residual data of the coding unit is partitioned by the conversion unit with a tree structure according to the conversion depth.

[0172] Regarding the conversion unit according to an embodiment as well, the height and width of the coding unit are divided, and a conversion depth indicating the number of divisions until the conversion unit is set. For example, if the size of the conversion unit of the current coding unit of size 2Nx2N is 2Nx2N, it is set to conversion depth 0, if the size of the conversion unit is NxN, it is set to conversion depth 1, and if the size of the conversion unit is N / 2xN / 2, it is set to conversion depth 2. That is, for the conversion unit as well, the conversion unit with a tree structure is set according to the conversion depth.

[0173] The depth-wise segmentation information requires not only the depth but also prediction-related information and conversion-related information. Therefore, the coding unit determination unit 820 can determine not only the depth that generates the minimum coding error, but also the partition mode in which the prediction unit is partitioned into partitions, the prediction mode for each prediction unit, the size of the conversion unit for conversion, and the like.

[0174] Regarding the determination method of the coding unit, prediction unit / partition, and conversion unit according to the tree structure of the maximum coding unit according to an embodiment, it will be described in detail with reference to FIGS. 17 to 19.

[0175] The coding unit determination unit 820 can measure the coding error of the depth-wise coding unit by using a rate-distortion optimization technique based on a Lagrangian multiplier.

[0176] The output unit 830 outputs the video data of the maximum coding unit and the depth-wise segmentation information encoded based on at least one depth determined by the coding unit determination unit 820 in the form of a bitstream.

[0177] The encoded video data is also the encoding result of the residual data of the video.

[0178] The depth-wise segmentation information may include depth information, partition mode information of the prediction unit, prediction mode information, division information of the conversion unit, and the like.

[0179] The final depth information is defined using depth-by-depth split information indicating whether to encode at the current depth or to encode in the encoding units of a lower depth without encoding at the current depth. If the current depth of the current encoding unit is the depth, then the current encoding unit is encoded in the encoding unit of the current depth, so the split information of the current depth is defined so as not to be further split into lower depths. On the contrary, if the current depth of the current encoding unit is not the depth, then encoding using the encoding unit of the lower depth must be attempted, so the split information of the current depth is defined so as to be split into the encoding units of the lower depth.

[0180] If the current depth is not the depth, encoding is performed on the encoding units split into the encoding units of the lower depth. Since there is one or more encoding units of the lower depth within the encoding unit of the current depth, encoding is repeatedly performed for each encoding unit of the lower depth, and recursive encoding is performed for each encoding unit of the same depth.

[0181] In one maximum encoding unit, the encoding units of the tree structure are determined, and at least one split information must be determined for each encoding unit of the depth. Therefore, at least one split information is determined for one maximum encoding unit. Also, the data of the maximum encoding unit is hierarchically partitioned by depth, and the depth is different depending on the position, so the depth and split information are set for the data.

[0182] Therefore, the output unit 830 according to one embodiment is assigned encoding information related to the depth and encoding mode for at least one of the encoding units, prediction units, and minimum units included in the maximum encoding unit.

[0183] The minimum unit according to one embodiment is a square data unit of the size obtained by dividing the minimum encoding unit, which is the lowest depth, into four parts. The minimum unit according to one embodiment is also the largest square data unit included in all the encoding units, prediction units, partition units, and conversion units included in the maximum encoding unit.

[0184] For example, the encoded information output via the output unit 830 is classified into encoded information by depth-level encoding unit and encoded information by prediction unit. The encoded information by depth-level encoding unit may include prediction mode information and partition size information. The encoded information transmitted by prediction unit may include information related to the estimation direction of the inter mode, information related to the reference video index of the inter mode, information related to the motion vector, information related to the chroma component of the intra mode, information related to the interpolation method of the intra mode, and the like.

[0185] Information related to the maximum size of the encoding unit defined by picture, slice, or GOP, and information related to the maximum depth are inserted into the header of the bitstream, the sequence parameter set, the picture parameter set, or the like.

[0186] Also, information related to the maximum size of the conversion unit allowed for the current video and information related to the minimum size of the conversion unit are also output via the header of the bitstream, the sequence parameter set, the picture parameter set, or the like. The output unit 830 can encode and output reference information related to prediction, prediction information, slice type information, and the like.

[0187] According to the simplest form of the embodiment of the video encoding apparatus 800, the encoding unit by depth level is an encoding unit having a size obtained by halving the height and width of the encoding unit of the upper depth of one layer. That is, if the size of the encoding unit of the current depth is 2Nx2N, the size of the encoding unit of the lower depth is NxN. Further, the current encoding unit of 2Nx2N size includes a maximum of four encoding units of the lower depth of NxN size.

[0188] Therefore, the video encoding device 800 can determine encoding units of optimal forms and sizes for each maximum encoding unit based on the size and maximum depth of the maximum encoding unit determined in consideration of the characteristics of the current picture, and can configure encoding units with a tree structure. Also, for each maximum encoding unit, since encoding can be performed in various prediction modes, conversion methods, etc., an optimal encoding mode is determined in consideration of the video characteristics of encoding units of various video sizes.

[0189] Therefore, if a video with a very high resolution or a very large amount of data is encoded in units of existing macroblocks, the number of macroblocks per picture becomes excessively large. As a result, the amount of compression information generated for each macroblock also increases, so the transmission burden of the compression information becomes large, and the data compression efficiency tends to decrease. Therefore, the video encoding device according to one embodiment can consider the size of the video, increase the maximum size of the encoding unit, and adjust the encoding unit in consideration of the video characteristics, so the video compression efficiency increases.

[0190] FIG. 9 illustrates a block diagram of a video decoding device 900 based on encoding units with a tree structure according to one embodiment.

[0191] The motion vector decoding device 20 described with reference to FIGS. 2A to 7B may be included in the video decoding device 900. That is, the motion vector decoding device 20 receives and parses a bitstream related to the encoded video for information indicating a predicted motion vector for performing inter prediction of the video decoded by the video decoding device 900, a residual motion vector, and information indicating a motion vector resolution, and can restore the motion vector based on the parsed information.

[0192] A video decoding apparatus 900 with video prediction based on an encoding unit with a tree structure according to an embodiment includes a receiving unit 910, a video data and encoding information extraction unit 920, and a video data decoding unit 930. Hereinafter, for convenience of explanation, a video decoding apparatus 900 with video prediction based on an encoding unit with a tree structure according to an embodiment is referred to by the abbreviation "video decoding apparatus 900".

[0193] The definitions of various terms such as the encoding unit, depth, prediction unit, transformation unit, and various partitioning information for the decoding operation of the video decoding apparatus 900 according to an embodiment are the same as those described with reference to FIG. 8 and the video encoding apparatus 800.

[0194] The receiving unit 910 receives and parses a bitstream related to the encoded video. The video data and encoding information extraction unit 920 extracts video data encoded for each encoding unit by the encoding unit with a tree structure for each maximum encoding unit from the parsed bitstream and outputs it to the video data decoding unit 930. The video data and encoding information extraction unit 920 can extract information related to the maximum size of the encoding unit of the current picture from the header related to the current picture, the sequence parameter set, or the picture parameter set.

[0195] Also, the video data and encoding information extraction unit 920 extracts the final depth and partitioning information related to the encoding unit with a tree structure for each maximum encoding unit from the parsed bitstream. The extracted final depth and partitioning information are output to the video data decoding unit 930. That is, the video data in the bit string is divided into maximum encoding units, and the video data decoding unit 930 decodes the video data for each maximum encoding unit.

[0196] The depth and partitioning information for each maximum encoding unit are set for one or more depth information, and the partitioning information for each depth may include the partition mode information, prediction mode information, and partitioning information of the transformation unit of the encoding unit. Also, the partitioning information for each depth may be extracted as the depth information.

[0197] The depth and splitting information for each maximum coding unit extracted by the video data and coding information extraction unit 920 are depth and splitting information determined by repeatedly performing coding for each coding unit for each depth for each maximum coding unit at the coding end, as in the video coding device 800 according to one embodiment, so as to generate a minimum coding error. Therefore, the video decoding device 900 can decode data by a coding method that generates a minimum coding error and restore the video.

[0198] According to one embodiment, since the coding information related to the depth and coding mode is assigned to a predetermined data unit among the coding unit, prediction unit, and minimum unit, the video data and coding information extraction unit 920 can extract the depth and splitting information for each predetermined data unit. If the depth and splitting information of the maximum coding unit are recorded for each predetermined data unit, the predetermined data units having the same depth and splitting information can be inferred as data units included in the same maximum coding unit.

[0199] The video data decoding unit 930 decodes the video data of each maximum coding unit based on the depth and splitting information for each maximum coding unit, and restores the current picture. That is, the video data decoding unit 930 can decode the video data encoded based on the decoded partition mode, prediction mode, and transform unit for each coding unit among the coding units having a tree structure included in the maximum coding unit. The decoding process may include a prediction process including intra prediction and motion compensation, and an inverse transform process.

[0200] The video data decoding unit 930 can perform intra prediction or motion compensation for each coding unit according to each partition and prediction mode based on the partition mode information and prediction mode information of the prediction unit of the coding unit by depth.

[0201] In addition, the video data decoding unit 930 can read conversion unit information with a tree structure for each encoding unit for inverse conversion by maximum encoding unit, and perform inverse conversion based on the conversion unit for each encoding unit. Through the inverse conversion, the pixel values in the spatial region of the encoding unit are restored.

[0202] The video data decoding unit 930 can determine the depth of the current maximum encoding unit by using the depth-wise segmentation information. If the segmentation information indicates that it cannot be further segmented at the current depth, then the current depth is the depth. Therefore, the video data decoding unit 930 can decode the encoding unit of the current depth for the video data of the current maximum encoding unit by using the partition mode, prediction mode, and conversion unit size information of the prediction unit.

[0203] That is, observe the encoding information set for a predetermined data unit among the encoding unit, prediction unit, and minimum unit, gather the data units having the encoding information including the same segmentation information, and the video data decoding unit 930 regards it as one data unit decoded by the same encoding mode. For each encoding unit determined in this way, obtain the information related to the encoding mode, and perform the decoding of the current encoding unit.

[0204] FIG. 10 illustrates the concept of an encoding unit according to an embodiment.

[0205] Examples of the encoding unit may include encoding units with sizes from an encoding unit with a size expressed as width x height and a size of 64x64 to sizes of 32x32, 16x16, and 8x8. The encoding unit with a size of 64x64 is divided into partitions with sizes of 64x64, 64x32, 32x64, and 32x32. The encoding unit with a size of 32x32 is divided into partitions with sizes of 32x32, 32x16, 16x32, and 16x16. The encoding unit with a size of 16x16 is divided into partitions with sizes of 16x16, 16x8, 8x16, and 8x8. The encoding unit with a size of 8x8 is divided into partitions with sizes of 8x8, 8x4, 4x8, and 4x4.

[0206] For video data 1010, the resolution is set to 1920x1080, the maximum size of the encoding unit is set to 64, and the maximum depth is set to 2. For video data 1020, the resolution is set to 1920x1080, the maximum size of the encoding unit is set to 64, and the maximum depth is set to 3. For video data 1030, the resolution is set to 352x288, the maximum size of the encoding unit is set to 16, and the maximum depth is set to 1. The maximum depth shown in FIG. 10 indicates the total number of divisions from the maximum encoding unit to the minimum encoding unit.

[0207] When the resolution is high or the data volume is large, in order to not only improve the encoding efficiency but also accurately reflect the video characteristics, it is desirable that the maximum size of the encoding size is relatively large. Therefore, for video data 1010 and 1020 with a higher resolution compared to video data 1030, the maximum size of the encoding size is selected to be 64.

[0208] Since the maximum depth of video data 1010 is 2, the encoding unit 1015 of video data 1010 may include the encoding units with major axis sizes of 32 and 16, which are obtained by dividing the maximum encoding unit with a major axis size of 64 twice, resulting in a depth of two levels. On the other hand, since the maximum depth of video data 1030 is 1, the encoding unit 1035 of video data 1030 may include the encoding unit with a major axis size of 8, which is obtained by dividing the encoding unit with a major axis size of 16 once, resulting in a depth of one level.

[0209] Since the maximum depth of video data 1020 is 3, the encoding unit 1025 of video data 1020 may include the encoding units with major axis sizes of 32, 16, and 8, which are obtained by dividing the maximum encoding unit with a major axis size of 64 three times, resulting in a depth of three levels. The deeper the depth, the better the ability to represent detailed information.

[0210] FIG. 11 illustrates a block diagram of a video encoding unit 1100 based on an encoding unit according to an embodiment.

[0211] According to one embodiment, the video encoding unit 1100 performs operations involved in encoding video data in the picture encoding unit 1520 of the video encoding apparatus 800. That is, the intra prediction unit 1120 performs intra prediction for each prediction unit on the encoding unit in the intra mode among the current video 1105, and the inter prediction unit 1115 performs inter prediction for each prediction unit on the encoding unit in the inter mode by using the current video 1105 and the reference video obtained in the reconstructed picture buffer 1110. After the current video 1105 is divided into maximum coding units, encoding is sequentially performed. At this time, encoding is performed on the coding units in which the maximum coding unit is divided into a tree structure.

[0212] Residual data is generated by removing the prediction data related to the encoding unit of each mode output from the intra prediction unit 1120 or the inter prediction unit 1115 from the data related to the encoding unit to be encoded of the current video 1105. The residual data is output as quantization coefficients quantized for each conversion unit through the conversion unit 1125 and the quantization unit 1130. The quantized conversion coefficients are restored to the residual data in the spatial domain through the inverse quantization unit 1145 and the inverse conversion unit 1150. The restored residual data in the spatial domain is added to the prediction data related to the encoding unit of each mode output from the intra prediction unit 1120 or the inter prediction unit 1115, thereby being restored to the data in the spatial domain related to the encoding unit of the current video 1105. The restored data in the spatial domain is generated as a reconstructed video through the deblocking unit 1155 and the SAO execution unit 1160. The generated reconstructed video is stored in the reconstructed picture buffer 1110. The reconstructed video stored in the reconstructed picture buffer 1110 is used as a reference video for inter prediction of other videos. The conversion coefficients quantized in the conversion unit 1125 and the quantization unit 1130 are output as a bitstream 1140 through the entropy encoding unit 1135.

[0213] For the video encoding unit 1100 according to an embodiment to be applied to the video encoding apparatus 800, the components of the video encoding unit 1100, namely, the inter prediction unit 1115, the intra prediction unit 1120, the conversion unit 1125, the quantization unit 1130, the entropy encoding unit 1135, the inverse quantization unit 1145, the inverse conversion unit 1150, the deblocking unit 1155, and the SAO execution unit 1160, can perform operations based on each encoding unit among the encoding units with a tree structure for each maximum coding unit.

[0214] In particular, the intra prediction unit 1120 and the inter prediction unit 1115 consider the maximum size and maximum depth of the current maximum coding unit, determine the partition mode and prediction mode of each encoding unit among the encoding units with a tree structure, and the conversion unit 1125 can determine how to divide the conversion units by quadtree within each encoding unit among the encoding units with a tree structure.

[0215] FIG. 12 illustrates a block diagram of a video decoding unit 1200 based on encoding units according to an embodiment.

[0216] The entropy decoding unit 1215 parses the encoded video data to be decoded and the encoding information necessary for decoding from the bitstream 1205. The encoded video data is quantized conversion coefficients, and the inverse quantization unit 1220 and the inverse conversion unit 1225 restore the residual data from the quantized conversion coefficients.

[0217] The intra prediction unit 1240 performs intra prediction for each prediction unit for the encoding unit in the intra mode. The inter prediction unit 1235 performs inter prediction for the encoding unit in the inter mode in the current video using the reference video obtained in the restoration picture buffer 1230 for each prediction unit.

[0218] By adding the prediction data and the residual data related to the coding unit of each mode that has passed through the intra prediction unit 1240 or the inter prediction unit 1235, the data of the spatial region related to the coding unit of the current video 1105 is restored. The restored data of the spatial region is output as a restored video 1260 after passing through the deblocking unit 1245 and the SAO execution unit 1250. Also, the restored video stored in the restored picture buffer 1230 is output as a reference video.

[0219] In order to decode video data in the picture decoder 930 of the video decoder 900, the step-by-step operations after the entropy decoder 1215 of the video decoder 1200 according to an embodiment are performed.

[0220] Since the video decoder 1200 is applied to the video decoder 900 according to an embodiment, the components of the video decoder 1200, namely the entropy decoder 1215, the inverse quantization unit 1220, the inverse transform unit 1225, the intra prediction unit 1240, the inter prediction unit 1235, the deblocking unit 1245, and the SAO execution unit 1250, can perform operations based on each coding unit of the coding units with a tree structure for each maximum coding unit.

[0221] In particular, the intra prediction unit 1240 and the inter prediction unit 1235 can determine the partition mode and the prediction mode for each coding unit of the coding units with a tree structure, and the inverse transform unit 1225 can determine how to divide the transform units with a quadtree structure for each coding unit.

[0222] FIG. 13 illustrates coding units and partitions by depth according to an embodiment.

[0223] The video encoding device 800 according to one embodiment and the video decoding device 900 according to one embodiment use hierarchical encoding units in order to consider video characteristics. The maximum height, maximum width, and maximum depth of the encoding unit are adaptively determined according to the characteristics of the video and can also be variously set according to the user's requirements. The size of the encoding unit by depth is determined by the preset maximum size of the encoding unit.

[0224] The hierarchical structure 1300 of the encoding unit according to one embodiment illustrates the case where the maximum height and maximum width of the encoding unit are 64 and the maximum depth is 3. At this time, the maximum depth indicates the total number of divisions from the maximum encoding unit to the minimum encoding unit. Along the vertical axis of the hierarchical structure 1300 of the encoding unit according to one embodiment, as the depth increases, the height and width of the encoding unit by depth are respectively divided. Also, along the horizontal axis of the hierarchical structure 1300 of the encoding unit, the prediction units and partitions that are the basis for predictive encoding of each encoding unit by depth are illustrated.

[0225] That is, the encoding unit 1310 is the maximum encoding unit in the hierarchical structure 1300 of the encoding unit, has a depth of 0, and the size of the encoding unit, that is, the height and width are 64x64. As the depth increases along the vertical axis, there are the encoding unit 1320 at depth 1 with a size of 32x32, the encoding unit 1330 at depth 2 with a size of 16x16, and the encoding unit 1340 at depth 3 with a size of 8x8. The encoding unit 1340 at depth 3 with a size of 8x8 is the minimum encoding unit.

[0226] For each depth, the prediction units and partitions of the encoding unit are arranged along the horizontal axis. That is, if the encoding unit 1310 with a size of 64x64 at depth 0 is the prediction unit, the prediction unit is divided into the partition 1310 with a size of 64x64, the partition 1312 with a size of 64x32, the partition 1314 with a size of 32x64, and the partition 1316 with a size of 32x32 included in the encoding unit 1310 with a size of 64x64.

[0227] Similarly, the prediction units of the encoding unit 1320 with a size of 32x32 at depth 1 are divided into a partition 1320 with a size of 32x32, a partition 1322 with a size of 32x16, a partition 1324 with a size of 16x32, and a partition 1326 with a size of 16x16 included in the encoding unit 1320 with a size of 32x32.

[0228] Similarly, the prediction units of the encoding unit 1330 with a size of 16x16 at depth 2 are divided into a partition 1330 with a size of 16x16, a partition 1332 with a size of 16x8, a partition 1334 with a size of 8x16, and a partition 1336 with a size of 8x8 included in the encoding unit 1330 with a size of 16x16.

[0229] Similarly, the prediction units of the encoding unit 1340 with a size of 8x8 at depth 3 are divided into a partition 1340 with a size of 8x8, a partition 1342 with a size of 8x4, a partition 1344 with a size of 4x8, and a partition 1346 with a size of 4x4 included in the encoding unit 1340 with a size of 8x8.

[0230] The encoding unit determination unit 820 of the video encoding apparatus 800 according to an embodiment has to perform encoding for each encoding unit of each depth included in the maximum encoding unit 1310 in order to determine the depth of the maximum encoding unit 1310.

[0231] The number of encoding units by depth for including data of the same range and the same size increases as the depth becomes deeper. For example, for data including one encoding unit at depth 1, four encoding units at depth 2 are required. Therefore, in order to compare the encoding results of the same data by depth, one encoding unit at depth 1 and four encoding units at depth 2 have to be used for encoding respectively.

[0232] For each depth - specific encoding, along the horizontal axis of the hierarchical structure 1300 of the encoding units, encoding is performed for each prediction unit of the depth - specific encoding units, and a representative encoding error, which is the minimum encoding error at that depth, is selected. Also, along the vertical axis of the hierarchical structure 1300 of the encoding units, the depth increases, encoding is performed for each depth, the depth - specific representative encoding errors are compared, and the minimum encoding error is searched for. In the maximum encoding unit 1310, the depth and partition where the minimum encoding error occurs are selected as the depth and partition mode of the maximum encoding unit 1310.

[0233] FIG. 14 illustrates the relationship between encoding units and conversion units according to an embodiment.

[0234] A video encoding device 800 according to an embodiment, or a video decoding device 900 according to an embodiment, encodes or decodes video for each maximum encoding unit with an encoding unit that is smaller than or the same size as the maximum encoding unit. In the encoding process, the size of the conversion unit for conversion is selected based on a data unit that is not as large as each encoding unit.

[0235] For example, in a video encoding device 800 according to an embodiment, or a video decoding device 900 according to an embodiment, when the current encoding unit 1410 is 64x64 in size, conversion is performed using a 32x32 - sized conversion unit 1420.

[0236] Also, after encoding the data of the 64x64 - sized encoding unit 1410 by performing conversion with conversion units of 32x32, 16x16, 8x8, and 4x4 sizes that are 64x64 or smaller, the conversion unit with the minimum error from the original is selected.

[0237] FIG. 15 illustrates encoding information according to an embodiment.

[0238] The output unit 830 of the video encoding device 800 according to an embodiment can encode and transmit, as division information, information 1500 related to the partition mode, information 1510 related to the prediction mode, and information 1520 related to the transform unit size for each encoding unit of each depth.

[0239] The information 1500 related to the partition mode indicates information related to the form of the partition into which the prediction unit of the current encoding unit is divided as the data unit for the predictive encoding of the current encoding unit. For example, the current encoding unit CU_0 of size 2Nx2N is divided and used in one of the partition types of partition 1502 of size 2Nx2N, partition 1504 of size 2NxN, partition 1506 of size Nx2N, and partition 1508 of size NxN. In that case, the information 1500 related to the partition mode of the current encoding unit is set to indicate one of the partition 1502 of size 2Nx2N, partition 1504 of size 2NxN, partition 1506 of size Nx2N, and partition 1508 of size NxN.

[0240] The information 1510 related to the prediction mode indicates the prediction mode of each partition. For example, through the information 1510 related to the prediction mode, it is set that the partition indicated by the information 1500 related to the partition mode is subjected to predictive encoding in one of the intra mode 1512, the inter mode 1514, and the skip mode 1516.

[0241] Also, the information 1520 related to the transform unit size indicates based on which transform unit the current encoding unit is to be transformed. For example, the transform unit is one of the first intra-transform unit size 1522, the second intra-transform unit size 1524, the first inter-transform unit size 1526, and the second inter-transform unit size 1528.

[0242] For each encoding unit by depth, the video data and encoding information extraction unit 1610 of the video decoder 900 according to an embodiment can extract information 1500 related to the partition mode, information 1510 related to the prediction mode, and information 1520 related to the conversion unit size, and use them for decoding.

[0243] FIG. 16 illustrates an encoding unit by depth according to an embodiment.

[0244] To show the change in depth, division information is used. The division information indicates whether the encoding unit of the current depth is divided into encoding units of a lower depth.

[0245] The prediction units 1610 for the prediction encoding of the encoding units 1600 of depth 0 and size 2N_0x2N_0 may include a partition mode 1612 of size 2N_0x2N_0, a partition mode 1614 of size 2N_0xN_0, a partition mode 1616 of size N_0x2N_0, and a partition mode 1618 of size N_0xN_0. Only the partitions 1612, 1614, 1616, 1618 in which the prediction units are divided into symmetric ratios are illustrated, but as described above, the partition mode is not limited to them, and may include asymmetric partitions, arbitrary-shaped partitions, geometric-shaped partitions, and the like.

[0246] For each partition mode, prediction encoding is repeatedly performed for each one partition of size 2N_0x2N_0, two partitions of size 2N_0xN_0, two partitions of size N_0x2N_0, and four partitions of size N_0xN_0. For partitions of size 2N_0x2N_0, size N_0x2N_0, size 2N_0xN_0, and size N_0xN_0, prediction encoding is performed in the intra mode and the inter mode. The skip mode performs prediction encoding only for the partition of size 2N_0x2N_0.

[0247] If the encoding error is minimized by one of the partitioning modes 1612, 1614, 1616 of size 2N_0x2N_0, 2N_0xN_0, and N_0x2N_0, there is no need to further divide into lower depths.

[0248] If the encoding error is minimized by the partitioning mode 1618 of size N_0xN_0, divide while changing depth 0 to 1 (1620), and repeatedly perform encoding on depth 2 and the encoding unit 1630 of the partitioning mode of size N_0xN_0 to search for the minimum encoding error.

[0249] For the prediction unit 1640 for predictive encoding of the encoding unit 1630 of depth 1 and size 2N_1x2N_1 (= N_0xN_0), it may include the partitioning mode 1642 of size 2N_1x2N_1, the partitioning mode 1644 of size 2N_1xN_1, the partitioning mode 1646 of size N_1x2N_1, and the partitioning mode 1648 of size N_1xN_1.

[0250] Also, if the encoding error is minimized by the partitioning mode 1648 of size N_1xN_1, divide while changing depth 1 to depth 2 (1650), and repeatedly perform encoding on depth 2 and the encoding unit 1660 of size N_2xN_2 to search for the minimum encoding error.

[0251] When the maximum depth is d, the encoding units by depth are set until depth d - 1, and the division information is set until depth d - 2. That is, when divided from depth d - 2 and encoded until depth d - 1, the prediction unit 1690 for predictive encoding of the encoding unit 1680 of depth d - 1 and size 2N_(d - 1)x2N_(d - 1) may include the partitioning mode 1692 of size 2N_(d - 1)x2N_(d - 1), the partitioning mode 1694 of size 2N_(d - 1)xN_(d - 1), the partitioning mode 1696 of size N_(d - 1)x2N_(d - 1), and the partitioning mode 1698 of size N_(d - 1)xN_(d - 1).

[0252] Among the partition modes, for each partition of size 2N_(d - 1) x 2N_(d - 1), two partitions of size 2N_(d - 1) x N_(d - 1), two partitions of size N_(d - 1) x 2N_(d - 1), and four partitions of size N_(d - 1) x N_(d - 1), encoding through predictive encoding is repeatedly performed, and the partition mode that generates the minimum encoding error is searched for.

[0253] Even if the encoding error by the partition mode 1698 of size N_(d - 1) x N_(d - 1) is the minimum, since the maximum depth is d, the encoding unit CU_(d - 1) at depth d - 1 does not go through the process of further splitting to lower depths, and the depth related to the current maximum encoding unit 1600 is determined to be depth d - 1, and the partition mode is determined to be N_(d - 1) x N_(d - 1). Also, since the maximum depth is d, no splitting information is set for the encoding unit 1652 at depth d - 1.

[0254] The data unit 1699 is regarded as the "minimum unit" related to the current maximum encoding unit. The minimum unit according to one embodiment is also a square data unit of the size obtained by dividing the minimum encoding unit, which is the lowest depth, into four parts. Through such an iterative encoding process, the video encoding device 800 according to one embodiment compares the encoding errors by depth of the encoding unit 1600, selects the depth at which the minimum encoding error occurs, determines the depth, and the partition mode and prediction mode of the depth are set as the encoding mode of the depth.

[0255] In this way, all the minimum encoding errors by depth of depths 0, 1, …, d - 1, d are compared, and the depth with the minimum error is selected and determined by the depth. The depth, and the partition mode and prediction mode of the prediction unit are encoded and transmitted as splitting information. Also, since the encoding unit must be split from depth 0 to the depth, only the splitting information of the depth is set to "0", and the splitting information by depth excluding the depth must be set to "1".

[0256] The video data and encoding information extraction unit 920 of the video decoding apparatus 900 according to an embodiment extracts the depth related to the encoding unit 1600 and the information related to the prediction unit, and can use them for decoding the encoding unit 1612. The video decoding apparatus 900 according to an embodiment can use the depth division information to recognize the depth whose division information is "0" as the depth, and use the division information related to the depth for decoding.

[0257] FIG. 17, FIG. 18 and FIG. 19 illustrate the relationships among the encoding unit, the prediction unit, and the conversion unit according to an embodiment.

[0258] The encoding unit 1710 is an encoding unit by depth determined by the video encoding apparatus 800 according to an embodiment with respect to the maximum encoding unit. The prediction unit 1760 is a partition of the prediction unit of each encoding unit by depth in the encoding unit 1710, and the conversion unit 1770 is the conversion unit of each encoding unit by depth.

[0259] If the depth of the maximum encoding unit is 0, for the encoding units 1712 and 1054 by depth, the depth is 1, for the encoding units 1714, 1716, 1718, 1728, 1750, 1752 by depth, the depth is 2, for the encoding units 1720, 1722, 1724, 1726, 1730, 1732, 1748 by depth, the depth is 3, and for the encoding units 1740, 1742, 1744, 1746 by depth, the depth is 4.

[0260] In the prediction unit 1760, some partitions 1714, 1716, 1722, 1732, 1748, 1750, 1752, 1754 are in a form where the encoding unit is divided. That is, the partitions 1714, 1722, 1750, 1754 are in the 2NxN partition mode, the partitions 1716, 1748, 1752 are in the Nx2N partition mode, and the partition 1732 is in the NxN partition mode. The prediction unit and the partition of the encoding unit by depth are smaller than or the same as their respective encoding units.

[0261] In the conversion unit 1770, for the video data of the partial conversion unit 1752, conversion or inverse conversion is performed in a data unit with a size smaller than that of the encoding unit. Also, the conversion units 1714, 1716, 1722, 1732, 1748, 1750, 1752, and 1754 are data units of different sizes or forms when compared with the prediction unit 1760 and the partition in the prediction unit 1760. That is, the video encoding device 800 according to one embodiment and another video decoding device 900 according to one embodiment are based on separate data units even for the intra prediction / motion estimation / motion compensation operations and the conversion / inverse conversion operations related to the same encoding unit.

[0262] As a result, encoding is recursively performed for each maximum encoding unit and for each hierarchical-structured encoding unit by region, and an optimal encoding unit is determined, thereby constituting an encoding unit with a recursive tree structure. The encoding information may include split information related to the encoding unit, partition mode information, prediction mode information, and conversion unit size information. Table 1 below shows an example that can be set in the video encoding device 800 according to one embodiment and the video decoding device 900 according to one embodiment.

[0263]

Table 1

[0264] The splitting information indicates whether the current coding unit is to be split into coding units of a lower depth. If the splitting information for the current depth d is 0, then since the current coding unit is at a depth where it is not further split into lower coding units, partition mode information, prediction mode, and transform unit size information are defined for that depth. If the splitting information requires further splitting in one step, then independent coding must be performed for each of the four split coding units of the lower depth.

[0265] The prediction mode can be indicated by one of an intra mode, an inter mode, and a skip mode. The intra mode and the inter mode are defined for all partition modes, and the skip mode is defined only for the partition mode 2Nx2N.

[0266] The partition mode information can indicate symmetric partition modes 2Nx2N, 2NxN, Nx2N, and NxN where the height or width of the prediction unit is divided in symmetric ratios, and asymmetric partition modes 2NxnU, 2NxnD, nLx2N, nRx2N where the division is in asymmetric ratios. The asymmetric partition modes 2NxnU and 2NxnD are in forms where the height is divided in ratios of 1:3 and 3:1 respectively, and the asymmetric partition modes nLx2N and nRx2N are in forms where the width is divided in ratios of 1:3 and 3:1 respectively.

[0267] The transform unit size is set to two sizes in the intra mode and two sizes in the inter mode. That is, if the transform unit splitting information is 0, then the size of the transform unit is set to the size 2Nx2N of the current coding unit. If the transform unit splitting information is 1, then a transform unit of the size into which the current coding unit is split is set. Also, if the partition mode related to the current coding unit of size 2Nx2N is a symmetric partition mode, then the size of the transform unit is set to NxN, and if it is an asymmetric partition mode, then it is set to N / 2xN / 2.

[0268] According to one embodiment, the encoded information of the encoding unit with a tree structure is assigned to at least one of the encoding units of depth, prediction units, and minimum units. The encoding unit of depth may include one or more prediction units and minimum units that hold the same encoded information.

[0269] Therefore, by checking the encoded information held by adjacent data units respectively, it can be confirmed whether they are included in the encoding units of the same depth. Also, by using the encoded information held by the data unit, the encoding unit of that depth can be confirmed, so the distribution of depths within the maximum encoding unit can be inferred.

[0270] Therefore, in that case, when the current encoding unit predicts by referring to the surrounding data units, the encoded information of the data units within the encoding units of different depths adjacent to the current encoding unit is directly referred to and used.

[0271] In another embodiment, when the current encoding unit performs predictive encoding by referring to the surrounding encoding units, the encoded information of the adjacent encoding units of different depths is used, and within the encoding units of different depths, the data adjacent to the current encoding unit is searched, whereby the surrounding encoding units are also referred to.

[0272] FIG. 20 illustrates the relationship between the encoding unit, prediction unit, and transform unit according to the encoding mode information in Table 1.

[0273] The maximum encoding unit 2000 includes encoding units of depth 2002, 2004, 2006, 2012, 2014, 2016, 2018. Since one of the encoding units 2018 is an encoding unit of depth, the split information is set to 0. The partition mode information of the encoding unit 2018 with a size of 2Nx2N is set to one of the partition modes 2Nx2N 2022, 2NxN 2024, Nx2N 2026, NxN 2028, 2NxnU 2032, 2NxnD 2034, nLx2N 2036, and nRx2N 2038.

[0274] The transformation unit division information (TU size flag) is a type of transformation index, and the size of the transformation unit corresponding to the transformation index is changed according to the prediction unit type or partition mode of the coding unit.

[0275] For example, when the partition mode information is set to one of the symmetric partition modes 2Nx2N 2022, 2NxN 2024, Nx2N 2026, and NxN 2028, if the transformation unit division information is 0, a transformation unit 2042 with a size of 2Nx2N is set, and if the transformation unit division information is 1, a transformation unit 2044 with a size of NxN is set.

[0276] When the partition mode information is set to one of the asymmetric partition modes 2NxnU 2032, 2NxnD 2034, nLx2N 2036, and nRx2N 2038, if the transformation unit division information (TU size flag) is 0, a transformation unit 2052 with a size of 2Nx2N is set, and if the transformation unit division information is 1, a transformation unit 2054 with a size of N / 2xN / 2 is set.

[0277] The transformation unit division information (TU size flag) described with reference to FIG. 20 is a flag having a value of 0 or 1. However, the transformation unit division information according to an embodiment is not limited to a 1-bit flag, and can be increased to 0, 1, 2, 3,... etc. according to the setting, and the transformation unit can also be hierarchically divided. The transformation unit division information is used as an embodiment of the transformation index.

[0278] In that case, if the conversion unit division information according to one embodiment is used together with the maximum size of the conversion unit and the minimum size of the conversion unit, the size of the actually used conversion unit is represented. The video encoding apparatus 800 according to one embodiment can encode the maximum conversion unit size information, the minimum conversion unit size information, and the maximum conversion unit division information. The encoded maximum conversion unit size information, minimum conversion unit size information, and maximum conversion unit division information are inserted into the SPS. The video decoding apparatus 900 according to one embodiment can use the maximum conversion unit size information, the minimum conversion unit size information, and the maximum conversion unit division information for video decoding.

[0279] For example, (a) if the current encoding unit has a size of 64x64 and the maximum conversion unit size is 32x32, then (a-1) when the conversion unit division information is 0, the size of the conversion unit is set to 32x32, (a-2) when the conversion unit division information is 1, the size of the conversion unit is set to 16x16, and (a-3) when the conversion unit division information is 2, the size of the conversion unit is set to 8x8.

[0280] As another example, (b) if the current encoding unit has a size of 32x32 and the minimum conversion unit size is 32x32, then (b-1) when the conversion unit division information is 0, the size of the conversion unit is set to 32x32, and since the size of the conversion unit is not smaller than 32x32, no further conversion unit division information is set.

[0281] As yet another example, (c) if the current encoding unit has a size of 64x64 and the maximum conversion unit division information is 1, then the conversion unit division information is 0 or 1, and no other conversion unit division information is set.

[0282] Therefore, when defining the maximum transform unit division information as "MaxTransformSizeIndex", the minimum transform unit size as "MinTransformSize", and the transform unit size when the transform unit division information is 0 as "RootTuSize", the minimum transform unit size "CurrMinTuSize" possible in the current coding unit is defined by the following formula (I).

[0283] CurrMinTuSize =max(MinTransformSize, RootTuSize / (2^MaxTransformSizeIndex)) (I) When compared with the minimum transform unit size "CurrMinTuSize" possible in the current coding unit, "RootTuSize", which is the transform unit size when the transform unit division information is 0, can indicate the maximum transform unit size that can be adopted on the system. That is, according to formula (I), "RootTuSize / (2^MaxTransformSizeIndex)" is the transform unit size obtained by dividing the transform unit size "RootTuSize" (when the transform unit division information is 0) by the number corresponding to the maximum transform unit division information. Since "MinTransformSize" is the minimum transform unit size, the smaller value among them is also the minimum transform unit size "CurrMinTuSize" possible in the current coding unit.

[0284] The maximum transform unit size "RootTuSize" according to an embodiment varies depending on the prediction mode.

[0285] For example, if the current prediction mode is the inter mode, "RootTuSize" is determined by the following formula (II). In formula (II), "MaxTransformSize" indicates the maximum transform unit size, and "PUSize" indicates the current prediction unit size.

[0286] RootTuSize=min(MaxTransformSize, PUSize) (II) That is, if the current prediction mode is the inter mode, "RootTuSize", which is the transform unit size when the transform unit division information is 0, is set to the smaller value of the maximum transform unit size and the current prediction unit size.

[0287] If the prediction mode of the current partition unit is the intra mode, "RootTuSize" is determined by the following formula (III). "PartitionSize" indicates the size of the current partition unit.

[0288] RootTuSize = min(MaxTransformSize, PartitionSize) (III) That is, if the current prediction mode is the intra mode, "RootTuSize", which is the transform unit size when the transform unit division information is 0, is set to the smaller value of the maximum transform unit size and the current partition unit size.

[0289] However, it should be noted that the current maximum transform unit size "RootTuSize" according to one embodiment that varies depending on the prediction mode of the partition unit is only one embodiment, and the factors determining the current maximum transform unit size are not limited thereto.

[0290] With the video encoding technique based on the tree-structured coding unit described with reference to FIGS. 8 to 20, for each tree-structured coding unit, the video data in the spatial region is encoded, and with the video decoding technique based on the tree-structured coding unit, while decoding is performed for each maximum coding unit, the video data in the spatial region is restored, and the video, which is a picture and a picture sequence, is restored. The restored video is played by a playback device, stored in a recording medium, or transmitted via a network.

[0291] On the other hand, the above-described embodiment of the present invention can be created as a program executed by a computer, and is embodied by a general-purpose digital computer that uses a computer-readable recording medium and operates the program. The computer-readable recording medium includes recording media such as magnetic recording media (e.g., ROM (read-only memory), floppy (registered trademark) disk, hard disk, etc.) and optical reading media (e.g., CD-ROM (compact disc read only memory), DVD (digital versatile disc), etc.).

[0292] For the sake of convenience of explanation, the video encoding method and / or video encoding method described above with reference to FIGS. 1A to 20 is referred to as the "video encoding method of the present invention". Also, the video decoding method and / or video decoding method described above with reference to FIGS. 1A to 20 is referred to as the "video decoding method of the present invention".

[0293] Also, the video encoding device, video encoding device 800, or video encoding device configured by video encoding unit 1100 described above with reference to FIGS. 1A to 20 is referred to as the "video encoding device of the present invention". Also, the video decoding device 900 or video decoding device configured by video decoding unit 1200 described above with reference to FIGS. 1A to 20 is referred to as the "video decoding device of the present invention".

[0294] An embodiment in which the computer-readable recording medium on which the program is stored is disk 21000 according to an embodiment will be described in detail below.

[0295] FIG. 21 illustrates the physical structure of a disk 26000 storing a program according to an embodiment. The disk 26000 described as a recording medium may also be a hard drive, a CD-ROM disk, a Blu-ray (registered trademark) disk, or a DVD disk. The disk 26000 is composed of a number of concentric tracks Tr, and each track Tr is divided into a predetermined number of sectors Se along the circumferential direction. A program for implementing the quantization parameter determination method, the video encoding method, and the video decoding method described above is assigned and stored in a specific area of the disk 26000 storing the program according to the above-described embodiment.

[0296] A computer system achieved by using a recording medium storing a program for implementing the above-described video encoding method and video decoding method will be described with reference to FIG. 22.

[0297] FIG. 22 illustrates a disk drive 26800 for recording and reading a program using the disk 26000. The computer system 26700 can store, using the disk drive 26800, a program for implementing at least one of the video encoding method and the video decoding method of the present invention on the disk 26000. To execute the program stored on the disk 26000 on the computer system 26700, the disk drive 26800 reads the program from the disk 26000 and transmits the program to the computer system 26700.

[0298] A program for implementing at least one of the video encoding method and the video decoding method of the present invention is stored not only in the disk 26000 illustrated in FIGS. 21 and 22 but also in a memory card, a ROM cassette, and an SSD (solid state drive).

[0299] A system to which the video encoding method and the video decoding method according to the above-described embodiment are applied will be described.

[0300] FIG. 23 illustrates the overall structure of a content supply system 11000 for providing a content distribution service. The service area of the communication system is divided into cells of a predetermined size, and radio base stations 11700, 11800, 11900, 12000 serving as base stations are installed in each cell.

[0301] The content supply system 11000 includes a number of independent devices. For example, independent devices such as a computer 12100, a PDA (personal digital assistant) 12200, a camera 12600, and a mobile phone 12500 are connected to the Internet 11100 via an Internet service provider 11200, a communication network 11400, and radio base stations 11700, 11800, 11900, 12000.

[0302] However, the content supply system 11000 is not limited to the structure illustrated in FIG. 23, and devices are selectively connected. The independent devices may also be directly connected to the communication network 11400 without passing through the radio base stations 11700, 11800, 11900, 12000.

[0303] The video camera 12300 is an imaging device that can capture video images like a digital video camera. The mobile phone 12500 can adopt at least one communication method among various protocols such as the PDC (personal digital communications) method, the CDMA (code division multiple access) method, the W-CDMA (wideband code division multiple access) method, the GSM (global system for mobile communications (registered trademark)) method, and the PHS (personal handyphone system) method.

[0304] The video camera 12300 is connected to the streaming server 11300 via the wireless base station 11900 and the communication network 11400. The streaming server 11300 can stream the content transmitted by the user using the video camera 12300 in real-time broadcast. The content received from the video camera 12300 is encoded by the video camera 12300 or the streaming server 11300. The video data captured by the video camera 12300 is also transmitted to the streaming server 11300 via the computer 12100.

[0305] The video data captured by the camera 12600 is also transmitted to the streaming server 11300 via the computer 12100. The camera 12600 is an imaging device that can capture both still images and video images like a digital camera. The video data received from the camera 12600 is encoded by the camera 12600 or the computer 12100. Software for video encoding and video decoding is stored on a computer-readable recording medium such as a CD-ROM disk, a floppy disk, a hard disk drive, an SSD, or a memory card that can be accessed by the computer 12100.

[0306] Also, when a video is captured by a camera mounted on the mobile phone 12500, the video data is received from the mobile phone 12500.

[0307] The video data is encoded by an LSI (large scale integrated circuit) system mounted on the video camera 12300, the mobile phone 12500, or the camera 12600.

[0308] In a content supply system 11000 according to an embodiment, for example, content recorded by a user using a video camera 12300, a camera 12600, a mobile phone 12500, or another imaging device, such as live concert recording content, is encoded and transmitted to a streaming server 11300. The streaming server 11300 can stream and transmit the content data to other clients that have requested the content data.

[0309] The client is a device that can decode the encoded content data, and is also, for example, a computer 12100, a PDA 12200, a video camera 12300, or a mobile phone 12500. Therefore, the content supply system 11000 causes the client to receive and play back the encoded content data. Further, the content supply system 11000 causes the client to receive the encoded content data, decode it in real time, and play it back, enabling personal broadcasting.

[0310] The video encoding device and video decoding device of the present invention are applied to the encoding operation and decoding operation of the independent devices included in the content supply system 11000.

[0311] Referring to FIGS. 24 and 25, an embodiment of the mobile phone 12500 in the content supply system 11000 will be described in detail.

[0312] FIG. 24 illustrates an external structure of a mobile phone 12500 to which the video encoding method and video decoding method of the present invention according to an embodiment are applied. The mobile phone 12500 is also a smartphone whose functions are not limited and whose functions can be changed or extended through application programs.

[0313] The mobile phone 12500 includes a built-in antenna 12510 for exchanging RF signals with a wireless base station 12000, and a display screen 12520 such as an LCD (liquid crystal display) screen or an OLED (organic light emitting diodes) screen for displaying video captured by a camera 12530 or video received and decoded by the antenna 12510. The smartphone 12510 includes an operation panel 12540 including control buttons and a touch panel. When the display screen 12520 is a touch screen, the operation panel 12540 further includes a touch sensing panel of the display screen 12520. The smartphone 12510 includes a speaker 12580 for outputting voice and sound, or other forms of audio output units, and a microphone 12550 for inputting voice and sound, or other forms of audio input units. The smartphone 12510 further includes a camera 12530 such as a CCD camera for shooting video and still images. In addition, the smartphone 12510 may include a recording medium 12570 for storing encoded or decoded data such as video and still images captured by the camera 12530, received by e-mail, or acquired in other forms, and a slot 12560 for attaching the recording medium 12570 to the mobile phone 12500. The recording medium 12570 may also be other forms of flash memory such as an SD card or an EEPROM (electrically erasable and programmable read only memory) built into a plastic case.

[0314] FIG. 25 illustrates the internal structure of the mobile phone 12500. In order to organize the control of each part of the mobile phone 12500 composed of the display screen 12520 and the operation panel 12540, a power supply circuit 12700, an operation input control unit 12640, a video encoding unit 12720, a camera interface 12630, an LCD control unit 12620, a video decoding unit 12690, a multiplexer / demultiplexer (MUX / DEMUX) 12680, a recording / deciphering unit 12670, a modulation / demodulation unit 12660, and an acoustic processing unit 12650 are connected to the central control unit 12710 via a synchronization bus 12730.

[0315] When the user operates the power button and sets the mobile phone from the "power off" state to the "power on" state, the power supply circuit 12700 supplies power from the battery pack to each part of the mobile phone 12500, so that the mobile phone 12500 is set to the operation mode.

[0316] The central control unit 12710 includes a CPU (central processing unit), a ROM, and a RAM (random access memory).

[0317] When the mobile phone 12500 transmits communication data externally, digital signals are generated in the mobile phone 12500 under the control of the central control unit 12710. For example, in the audio processing unit 12650, digital audio signals are generated, in the video encoding unit 12720, digital video signals are generated, and text data of messages are generated via the operation panel 12540 and the operation input control unit 12640. If the digital signals are transmitted to the modulation / demodulation unit 12660 under the control of the central control unit 12710, the modulation / demodulation unit 12660 modulates the frequency band of the digital signals, and the communication circuit 12610 performs D / A conversion (digital-analog conversion) processing and frequency conversion processing on the band-modulated digital audio signals. The transmission signal output from the communication circuit 12610 is sent to the voice communication base station or the radio base station 12000 via the antenna 12510.

[0318] For example, when the mobile phone 12500 is in the call mode, the acoustic signal acquired by the microphone 12550 is converted into a digital acoustic signal in the acoustic processing unit 12650 under the control of the central control unit 12710. The generated digital acoustic signal is converted into a transmission signal through the modulation / demodulation unit 12660 and the communication circuit 12610, and is sent out via the antenna 12510.

[0319] When a text message such as an e-mail is transmitted in the data communication mode, the text data of the message is input using the operation panel 12540, and the text data is transmitted to the central control unit 12610 via the operation input control unit 12640. Under the control of the central control unit 12610, the text data is converted into a transmission signal through the modulation / demodulation unit 12660 and the communication circuit 12610, and is sent to the radio base station 12000 via the antenna 12510.

[0320] In the data communication mode, in order to transmit video data, the video data captured by the camera 12530 is provided to the video encoding unit 12720 via the camera interface 12630. The video data captured by the camera 12530 is immediately displayed on the display screen 12520 via the camera interface 12630 and the LCD control unit 12620.

[0321] The structure of the video encoding unit 12720 corresponds to the structure of the video encoding apparatus of the present invention described above. The video encoding unit 12720 encodes the video data provided from the camera 12530 by the video encoding method of the present invention described above, converts it into compressed and encoded video data, and can output the encoded video data to the multiplexing / demultiplexing unit 12680. During the recording of the camera 12530, the acoustic signal acquired by the microphone 12550 of the mobile phone 12500 is also converted into digital acoustic data via the acoustic processing unit 12650, and the digital acoustic data is transmitted to the multiplexing / demultiplexing unit 12680.

[0322] The multiplexing / demultiplexing unit 12680 multiplexes the encoded video data provided from the video encoding unit 12720 together with the acoustic data provided from the acoustic processing unit 12650. The multiplexed data is converted into a transmission signal via the modulation / demodulation unit 12660 and the communication circuit 12610, and is sent out via the antenna 12510.

[0323] In the process of the mobile phone 12500 receiving communication data from the outside, the signal received via the antenna 12510 is converted into a digital signal via frequency recovery processing and A / D conversion processing. The modulation / demodulation unit 12660 demodulates the frequency band of the digital signal. The digitally demodulated signal is transmitted to the video decoding unit 12690, the acoustic processing unit 12650 or the LCD control unit 12620 depending on the type.

[0324] When the mobile phone 12500 is in the call mode, it amplifies the signal received via the antenna 12510 and generates a digital audio signal through frequency conversion and A / D conversion (analog-digital conversion) processing. The received digital audio signal is converted into an analog audio signal under the control of the central control unit 12710 via the modulation / demodulation unit 12660 and the audio processing unit 12650, and the analog audio signal is output via the speaker 12580.

[0325] In the data communication mode, when data of a video file accessed from a website on the Internet is received, the signal received from the radio base station 12000 via the antenna 12510 outputs multiplexed data as a processing result of the modulation / demodulation unit 12660, and the multiplexed data is transmitted to the multiplexing / demultiplexing unit 12680.

[0326] To decode the multiplexed data received via the antenna 12510, the multiplexing / demultiplexing unit 12680 demultiplexes the multiplexed data and separates an encoded video data stream and an encoded audio data stream. The encoded video data stream is provided to the video decoding unit 12690 by the synchronization bus 12730, and the encoded audio data stream is provided to the audio processing unit 12650.

[0327] The structure of the video decoding unit 12690 corresponds to the structure of the video decoding device of the present invention described above. The video decoding unit 12690 utilizes the video decoding method of the present invention described above to decode the encoded video data, generate restored video data, and can provide the restored video data to the display screen 12520 via the LCD control unit 12620.

[0328] As a result, the video data of the video file accessed from the Internet website is displayed on the display screen 12520. At the same time, the audio processing unit 12650 can also convert the audio data into an analog audio signal and provide the analog audio signal to the speaker 12580. Thereby, the audio data included in the video file accessed from the Internet website is also reproduced by the speaker 12580.

[0329] The mobile phone 12500, or other forms of communication terminal, is a transceiver terminal that includes both the video encoding device and the video decoding device of the present invention, a transmission terminal that includes only the aforementioned video encoding device of the present invention, or a reception terminal that includes only the video decoding device of the present invention.

[0330] The communication system of the present invention is not limited to the structure described with reference to FIG. 25. For example, FIG. 26 illustrates a digital broadcast system to which a communication system according to an embodiment is applied.

[0331] The digital broadcast system according to an embodiment of FIG. 26 can receive digital broadcasts transmitted via a satellite network or a terrestrial network by using the video encoding device and the video decoding device of the present invention.

[0332] Specifically, the broadcasting station 12890 transmits a video data stream to the communication satellite or the broadcasting satellite 12900 via radio waves. The broadcasting satellite 12900 transmits a broadcast signal, and the broadcast signal is received by the satellite broadcast receiver by the antenna 12860 at home. In each household, the encoded video stream is decoded and reproduced by the TV receiver 12810, the set-top box 12870, or other devices.

[0333] In the playback device 12830, when the video decoding device of the present invention is implemented, the playback device 12830 can read and decode an encoded video stream recorded on a recording medium 12820 such as a disk and a memory card. Thereby, the restored video signal is played back on, for example, the monitor 12840.

[0334] The video decoding device of the present invention is also mounted on a set-top box 12870 connected to an antenna 12860 for satellite / terrestrial wave broadcasting or a cable antenna 12850 for cable TV reception. The output data of the set-top box 12870 is also played back on the TV monitor 12880.

[0335] As another example, instead of the set-top box 12870, the video decoding device of the present invention may also be mounted on the TV receiver 12810 itself.

[0336] An automobile 12920 equipped with an appropriate antenna 12910 can also receive signals transmitted from a satellite 12800 or a radio base station 11700. Decoded video is played back on the display screen of an in-vehicle navigation system 12930 mounted on the automobile 12920.

[0337] The video signal is encoded by the video encoding device of the present invention, recorded on a recording medium, and stored. Specifically, a video signal is stored on a DVD disk 12960 by a DVD recorder, or a video signal is stored on a hard disk by a hard disk recorder 12950. As another example, the video signal may also be stored on an SD card 12970. If the hard disk recorder 12950 includes the video decoding device of the present invention according to an embodiment, the video signal recorded on the DVD disk 12960, the SD card 12970, or other forms of recording media is played back on the monitor 12880.

[0338] The automotive navigation system 12930 may not include the camera 12530, camera interface 12630, and video encoding unit 12720 of FIG. 25. For example, the computer 12100 and TV receiver 12810 may also not include the camera 12530, camera interface 12630, and video encoding unit 12720 of FIG. 25.

[0339] FIG. 27 illustrates a network structure of a cloud computing system using a video encoding device and a video decoding device according to an embodiment.

[0340] The cloud computing system of the present invention includes a cloud computing server 14000, a user DB 14100, computing resources 14200, and a user terminal.

[0341] The cloud computing system provides an on-demand outsourcing service of computing resources via an information communication network such as the Internet according to a request from a user terminal. In a cloud computing environment, a service provider integrates computing resources of data centers located at different physical positions using virtualization technology to provide services required by users. Service users do not install and use computing resources such as applications, storage, operating systems (OS), and security on terminals owned by each user, but can select and use services on a virtual space generated via virtualization technology at a desired time and to a desired extent.

[0342] The user terminal of a specific service user is connected to the cloud computing server 14100 via an information communication network including the Internet and a mobile communication network. The user terminal is provided with cloud computing services, particularly a video playback service, from the cloud computing server 14100. The user terminal can also be any electronic device capable of connecting to the Internet, such as a desktop PC 14300, a smart TV 14400, a smartphone 14500, a notebook computer 14600, a PMP (portable multimedia player) 14700, a tablet PC 14800, etc.

[0343] The cloud computing server 14100 can integrate a large number of computing resources 14200 distributed in the cloud network and provide them to the user terminal. The large number of computing resources 14200 includes various data services and may also include data uploaded from the user terminal. In this way, the cloud computing server 14100 integrates video databases distributed in various places using virtualization technology and provides the services required by the user terminal.

[0344] User information of users subscribing to cloud computing services is stored in the user DB 14100. Here, the user information may include login information and personal credit information such as address and name. Also, the user information may include an index of videos. Here, the index may include a list of videos that have been played, a list of videos being played, the stop time of the video being played, etc.

[0345] The information related to the video stored in the user DB 14100 is shared among user devices. Therefore, for example, when a playback request is made from the notebook computer 14600 and the notebook computer 14600 is provided with a predetermined video service, the playback history of the predetermined video service is stored in the user DB 14100. When a playback request for the same video service is received from the smartphone 14500, the cloud computing server 14100 refers to the user DB 14100, searches for the predetermined video service, and plays it. When the smartphone 14500 receives a video data stream via the cloud computing server 14100, the operation of decrypting the video data stream and playing the video is similar to the operation of the mobile phone 12500 described above with reference to FIG. 24.

[0346] The cloud computing server 14100 can also refer to the playback history of the predetermined video service stored in the user DB 14100. For example, the cloud computing server 14100 receives a playback request for the video stored in the user DB 14100 from the user terminal. If the video was being played previously, the cloud computing server 14100 will have different streaming methods depending on whether the user terminal selects to play from the beginning or resume from the previous stop point. For example, when the user terminal requests to play from the beginning, the cloud computing server 14100 streams the video to the user terminal from the first frame. On the other hand, when the terminal requests to resume playing from the previous stop point, the cloud computing server 14100 streams the video to the user terminal from the frame at the stop point.

[0347] At that time, the user terminal device may include the video decoding device of the present invention described with reference to FIGS. 1A to 20. As another example, the user terminal device may include the video encoding device of the present invention described with reference to FIGS. 1A to 20. Further, the user terminal device may include both the video encoding device and the video decoding device of the present invention described with reference to FIGS. 1A to 20.

[0348] An embodiment in which the video encoding method, the video decoding method, the video encoding device, and the video decoding device described with reference to FIGS. 1A to 20 are utilized is described with reference to FIGS. 21 to 27. However, an embodiment in which the video encoding method and the video decoding method described with reference to FIGS. 1A to 20 are stored in a recording medium or the video encoding device and the video decoding device are implemented in a device is not limited to the embodiment of FIGS. 21 to 27.

[0349] The method, process, device, product, and / or system according to the present invention are simple, cost-effective, not complex, and very diverse and accurate. Further, by applying well-known components to the process, device, product, and system according to the present invention, they can be immediately utilized, and efficient and economical manufacturing, application, and utilization can be implemented. Another important aspect of the present invention conforms to the current trend that requires cost reduction, system simplification, and performance improvement. Useful aspects that can be seen in such embodiments of the present invention will, as a result, at least raise the level of the current technology.

[0350] The present invention has been described in connection with specific preferred embodiments, but other inventions to which alternatives, modifications, and variations are applied to the present invention will be apparent to those skilled in the art in light of the foregoing description. That is, the claims are to be construed to include all such inventions to which such alternatives, modifications, and variations have been made. Accordingly, all of the content described in the specification and drawings must be construed in an illustrative and non-limiting sense. The methods, processes, devices, products, and / or systems according to the present invention are very diverse and accurate because they are simple, cost-effective, and not complex. In addition, by applying well-known components to the processes, devices, products, and systems according to the present invention, they can be immediately utilized, and efficient and economical manufacturing, application, and utilization can be realized. Another important aspect of the present invention is that it meets the current trend of demanding cost reduction, system simplification, and performance improvement. The useful aspects that can be seen in such embodiments of the present invention will, as a result, at least be able to raise the level of the current technology.

[0351] The present invention has been described in connection with specific preferred embodiments, but other inventions to which alternatives, modifications, and variations are applied to the present invention will be apparent to those skilled in the art in light of the foregoing description. That is, the claims are to be construed to include all such inventions to which such alternatives, modifications, and variations have been made. Accordingly, all of the content described in the specification and drawings must be construed in an illustrative and non-limiting sense.

[0352] Hereinafter, means based on the embodiments will be illustratively listed. (Appendix 1) In a motion vector encoding device, using the spatial candidate block and the temporal candidate block of the current block, obtaining a predicted motion vector candidate with a plurality of predetermined motion vector resolutions, and using the predicted motion vector candidate to determine the predicted motion vector of the current block, the motion vector of the current block, and the motion vector resolution of the current block; An encoding unit that encodes information indicating the predicted motion vector of the current block, the motion vector of the current block, the residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block, The apparatus according to claim 1, wherein the plurality of predetermined motion vector resolutions include resolutions in pixel units larger than the resolution in pixel units of 1 pixel. (Appendix 2) The prediction unit uses a set of first predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, and searches for a reference block in pixel units of the first motion vector resolution, uses a set of second predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, and searches for a reference block in pixel units of the second motion vector resolution, The first motion vector resolution and the second motion vector resolution are different from each other, The apparatus according to Appendix 1, wherein the first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate block and the temporal candidate block. (Appendix 3) The prediction unit uses a set of first predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, and searches for a reference block in pixel units of the first motion vector resolution, uses a set of second predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, and searches for a reference block in pixel units of the second motion vector resolution, The first motion vector resolution and the second motion vector resolution are different from each other, The apparatus according to Appendix 1, wherein the first set of predicted motion vector candidates and the second set of predicted motion vector candidates include different numbers of predicted motion vector candidates. (Appendix 4) The encoding unit when the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, downsizing and encoding the residual motion vector according to the resolution of the motion vector of the current block, the apparatus according to appended note 1. (Appended note 5) wherein the current block is a current encoding unit constituting a video, and the motion vector resolution is determined to be the same for each encoding unit, and if there is a prediction unit predicted to be in the AMVP (advanced motion vector prediction) mode within the current encoding unit the encoding unit encodes, as information indicating the motion vector resolution of the current block, information indicating the motion vector resolution of the prediction unit predicted to be in the AMVP mode once, the apparatus according to appended note 1. (Appended note 6) wherein the current block is a current encoding unit constituting a video, and the motion vector resolution is determined to be the same for each prediction unit, and if there is a prediction unit predicted to be in the AMVP mode within the current encoding unit the encoding unit encodes, as information indicating the motion vector resolution of the current block, information indicating the motion vector resolution for each prediction unit predicted to be in the AMVP mode existing within the current block, the apparatus according to appended note 1. (Appended note 7) In a motion vector encoding apparatus using a spatial candidate block and a temporal candidate block of a current block to obtain predicted motion vector candidates with a plurality of predetermined motion vector resolutions, and using the predicted motion vector candidates to determine a predicted motion vector of the current block, a motion vector of the current block, and a motion vector resolution of the current block, a prediction unit and an encoding unit that encodes information indicating the predicted motion vector of the current block, the motion vector of the current block, a residual motion vector between the predicted motion vector of the current block and the motion vector of the current block, and information indicating the motion vector resolution of the current block, wherein the prediction unit Using a set of first predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, a reference block is searched for in pixel units of the first motion vector resolution, Using a set of second predicted motion vector candidates including one or more predicted motion vector candidates selected from among the predicted motion vector candidates, a reference block is searched for in pixel units of the second motion vector resolution, The first motion vector resolution and the second motion vector resolution are different from each other, The first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate block and the temporal candidate block, or include different numbers of predicted motion vector candidates. An apparatus characterized by this. (Appendix 8) In a motion vector encoding apparatus, Generate a merge candidate list including at least one merge candidate related to the current block, use the motion vector of one candidate among the merge candidates included in the merge candidate list, and determine and encode the motion vector of the current block, The merge candidate list is characterized in that it includes motion vectors obtained by downscaling the motion vectors of the candidates included in the merge candidate list by a plurality of predetermined motion vector resolutions. An apparatus characterized by this. (Appendix 9) The downscaling is Instead of the pixel indicated by the motion vector of the minimum motion vector resolution, any one of the pixels located around the pixel indicated by the motion vector of the minimum motion vector resolution is selected based on the resolution of the motion vector of the current block, and adjusted to indicate the selected pixel. The apparatus according to Appendix 8, characterized by this. (Appendix 10) In a motion vector decoding apparatus, Using the current block's spatial candidate blocks and temporal candidate blocks, obtain predicted motion vector candidates with a plurality of predetermined motion vector resolutions, obtain information indicating the predicted motion vector of the current block among the predicted motion vector candidates, and obtain the residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block. An acquisition unit for acquiring; A decoding unit that restores the motion vector of the current block based on the residual motion vector, information indicating the predicted motion vector of the current block, and the motion vector resolution information of the current block. The apparatus according to claim 10, wherein the plurality of predetermined motion vector resolutions include resolutions in pixel units larger than the resolution in pixel units of one pixel. (Appendix 11) The predicted motion vector candidates with the plurality of predetermined motion vector resolutions are A set of first predicted motion vector candidates including one or more predicted motion vector candidates with a first motion vector resolution, and a set of second predicted motion vector candidates including one or more predicted motion vector candidates with a second motion vector resolution. The first motion vector resolution and the second motion vector resolution are different from each other. The apparatus according to Appendix 10, wherein the first predicted motion vector candidate set and the second predicted motion vector candidate set are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks, or include different numbers of predicted motion vector candidates. (Appendix 12) The decoding unit is The apparatus according to Appendix 10, wherein when the pixel unit of the motion vector resolution of the current block is larger than the pixel unit of the minimum motion vector resolution, the residual motion vector is upscaled by the minimum motion vector resolution and restored. (Appendix 13) The current block is a current coding unit that constitutes a video. For each coding unit, the motion vector resolution is determined to be the same. If there is a prediction unit predicted to be in the AMVP (advanced motion vector prediction) mode within the current coding unit, The acquisition unit acquires, from the bitstream, information indicating the motion vector resolution of the prediction unit predicted to be in the AMVP mode once as information indicating the motion vector resolution of the current block, according to the apparatus described in appended note 10. (Appended note 14) In a motion vector decoding apparatus, generate a merge candidate list including at least one merge candidate related to the current block, use the motion vector of one candidate among the merge candidates included in the merge candidate list to determine and decode the motion vector of the current block, The merge candidate list is characterized in that it includes motion vectors obtained by downscaling the motion vectors of the candidates included in the merge candidate list by a plurality of predetermined motion vector resolutions. (Appended note 15) In a motion vector decoding apparatus, using the spatial candidate block and the temporal candidate block of the current block, obtain prediction motion vector candidates of a plurality of predetermined motion vector resolutions, obtain information indicating the prediction motion vector of the current block among the prediction motion vector candidates, and obtain a residual motion vector between the motion vector of the current block and the prediction motion vector of the current block, and information indicating the motion vector resolution of the current block; and an acquisition unit a decoding unit that restores the motion vector of the current block based on the residual motion vector, the information indicating the prediction motion vector of the current block, and the motion vector resolution information of the current block, The prediction motion vector candidates of the plurality of predetermined motion vector resolutions are A set of first predicted motion vector candidates including predicted motion vector candidates of 1 or more of a first motion vector resolution, and a set of second predicted motion vector candidates including predicted motion vector candidates of 1 or more of a second motion vector resolution, wherein the first motion vector resolution and the second motion vector resolution are different resolutions from each other, wherein the first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate block and the temporal candidate block, or include different numbers of predicted motion vector candidates.

Prior Art Documents

Patent Documents

[0353]

Patent Document 1

Patent Document 2

Claims

1. In the method for decoding a motion vector, obtaining a predicted motion vector of a current block and a residual motion vector of the current block; obtaining a shift value for the residual motion vector based on a resolution corresponding to the current block among a plurality of resolutions; upscaling the residual motion vector by performing a left shift using the shift value; restoring a motion vector of the current block based on the upscaled residual motion vector and the predicted motion vector; the shift value is one of a plurality of integer values, When the information on the current block is first information, the number of the plurality of resolutions is a first number, and when the information on the current block is second information, the number of the plurality of resolutions is a second number; the first number of multiple resolutions includes a resolution greater than one; The first number of multiple resolutions and the second number of multiple resolutions include at least one resolution that is different from each other.

2. In the method for encoding a motion vector, obtaining a shift value for a residual motion vector of the current block based on a resolution corresponding to the current block among a plurality of resolutions; obtaining a predicted motion vector of the current block; obtaining a residual motion vector of the current block based on the motion vector of the current block and the predicted motion vector of the current block; downscaling a residual motion vector of the current block by performing a right shift using the shift value, and encoding the downscaled residual motion vector to generate a bitstream, the shift value is one of a plurality of integer values, When the information related to the current block is first information, the number of the plurality of resolutions is a first number, and when the information related to the current block is second information, the number of the plurality of resolutions is a second number; the first number of multiple resolutions includes a resolution greater than one; The first number of multiple resolutions and the second number of multiple resolutions include at least one resolution that is different from each other.

3. A method for transmitting a bitstream produced by the encoding method of claim 2.

Citation Information

Patent Citations

  • Image processing method, image processing unit and data storage medium

    JP1999239352A

  • Encoding apparatus, decoding apparatus, image processing apparatus, and method and program for them

    JP2003319400A

  • Moving picture signal coding method, decoding method, coding apparatus, and decoding apparatus

    JP2006187025A

  • Video coding using adaptive motion vector resolution

    US20130003849A1

  • Image processing device and method

    WO2010101064A1