Motion vector decoding method and encoding method
By employing multiple resolution techniques for motion vector prediction and encoding, the method addresses inefficiencies in existing video codecs, improving encoding and decoding complexity and compression rates.
Patent Information
- Application Number
- JP2025178222
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-10-31
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-27
AI Technical Summary
Existing video encoding and decoding methods face challenges in accurately predicting and encoding motion vectors, leading to inefficiencies in complexity and compression rates, particularly in codecs like H.264 AVC and HEVC.
A method and apparatus for predicting and encoding motion vectors using multiple resolutions, including pixel-unit resolutions greater than one pixel unit, by utilizing spatial and temporal candidate blocks to determine optimal motion vectors and resolutions, and encoding residual vectors efficiently.
This approach reduces the complexity of motion vector encoding and decoding processes, improving compression efficiency by adaptively determining motion vector resolutions and encoding residual vectors, thereby enhancing video encoding performance.
Smart Images

Figure 2026012838000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video encoding method and a video decoding method, and more particularly to a method and apparatus for predicting and encoding a motion vector of a video image, and a method and apparatus for predicting and decoding a motion vector of a video image. [Background technology]
[0002] In codecs such as H.264 AVC (advanced video coding) and HEVC (high efficiency video coding), to predict the motion vector of a current block, the motion vector of a previously coded block adjacent to the current block or a block at the same position in a previously coded picture can be used as the motion vector prediction of the current block.
[0003] In video encoding and decoding methods, in order to encode video, a picture is divided into macroblocks, and each macroblock can be predictively encoded using inter-prediction or intra-prediction.
[0004] Inter-prediction is a method for compressing video by removing temporal redundancy between pictures, and a typical example is motion estimation coding. Motion estimation coding predicts each block of a current picture using at least one reference picture. A predetermined evaluation function is used to search for the reference block that is most similar to the current block within a predetermined search range.
[0005] The current block is predicted based on the reference block, and the predicted block generated as a result of the prediction is subtracted from the current block to generate a residual block, which is then coded. In order to perform the prediction more accurately, interpolation is performed on the search range of the reference picture to generate sub-pixels in pixel units smaller than the integer pel unit, and inter-prediction can be performed based on the generated sub-pixels. Summary of the Invention
[0006] According to an embodiment, a motion vector decoding device and a motion vector encoding device and method thereof can determine an optimal predicted motion vector and a motion vector resolution, and efficiently encode or decode an image, thereby reducing the complexity of the device.
[0007] On the other hand, the technical problems and effects of the present invention are not limited to the features mentioned above, and other technical problems not mentioned or different from those mentioned above will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]
[0008] [Figure 1A] 1 is a block diagram of a motion vector encoding device according to an embodiment; [Figure 1B] 1 is a flowchart of a motion vector encoding method according to one embodiment. [Figure 2A] 1 is a block diagram of a motion vector decoding device according to one embodiment; [Figure 2B] 1 is a flowchart of a motion vector decoding method according to one embodiment. [Figure 3A] 1 is a diagram illustrating interpolation for performing motion compensation based on various resolutions; [Figure 3B] 10 is a diagram showing motion vector resolutions in units of 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixels. [Figure 4A] 10 is a diagram illustrating candidate blocks of a current block for obtaining candidate motion vector predictors. [Figure 4B]1 is a diagram illustrating a process of generating a motion vector predictor candidate according to an embodiment; [Figure 5A] 1 is a diagram illustrating coding units and prediction units according to an embodiment. [Figure 5B] FIG. 5B illustrates a portion of a prediction_unit syntax according to an embodiment for transmitting adaptively determined motion vector resolution. [Figure 5C] 10 is a diagram illustrating a portion of a prediction_unit syntax according to another embodiment for transmitting adaptively determined motion vector resolution. [Figure 5D] 10 is a diagram illustrating a portion of a prediction_unit syntax according to another embodiment for transmitting adaptively determined motion vector resolution. [Figure 6A] 1 illustrates one embodiment of constructing a merge candidate list using multiple resolutions. [Figure 6B] 10 illustrates another embodiment of constructing a merge candidate list using multiple resolutions. [Figure 7A] 10 is a diagram showing pixels indicated by two motion vectors with different resolutions; [Figure 7B] 1 is a diagram showing pixels constituting a picture enlarged by four times and motion vectors of different resolutions. [Figure 8] 1 is a block diagram illustrating a video encoding device based on a coding unit having a tree structure according to an embodiment of the present invention; [Figure 9] 1 is a block diagram illustrating a video decoding device based on a coding unit with a tree structure according to an embodiment; [Figure 10] 1 is a diagram illustrating a concept of a coding unit according to an embodiment; [Figure 11] 1 is a block diagram illustrating a coding unit-based video encoder according to one embodiment. [Figure 12] 1 is a block diagram illustrating a coding unit-based video decoder according to one embodiment. [Figure 13] 1 is a diagram illustrating coding units and partitions for each depth according to an embodiment; [Figure 14] 1 is a diagram illustrating a relationship between coding units and transform units according to one embodiment. [Figure 15] 1 is a diagram illustrating encoding information according to an embodiment; [Figure 16] 1 is a diagram illustrating coding units for each depth according to an embodiment; [Figure 17] 1 is a diagram illustrating a relationship between a coding unit, a prediction unit, and a transform unit according to an embodiment. [Figure 18] 1 is a diagram illustrating a relationship between a coding unit, a prediction unit, and a transform unit according to an embodiment. [Figure 19] 1 is a diagram illustrating a relationship between a coding unit, a prediction unit, and a transform unit according to an embodiment. [Figure 20] 10 is a diagram illustrating the relationship between coding units, prediction units, and transform units according to coding mode information in Table 1. [Figure 21] 1 is a diagram illustrating the physical structure of a disk on which a program is stored, according to one embodiment. [Figure 22] 1 is a diagram illustrating a disk drive for recording and reading a program using a disk. [Figure 23] 1 is a diagram illustrating the overall structure of a content supply system for providing a content distribution service. [Figure 24] 1 is a diagram illustrating an external structure of a mobile phone to which a video encoding method and a video decoding method of the present invention are applied, according to an embodiment; [Figure 25] 2 is a diagram illustrating the internal structure of the mobile phone. [Figure 26] 1 is a diagram illustrating a digital broadcasting system to which a communication system according to an embodiment is applied; [Figure 27]1 is a diagram illustrating a network structure of a cloud computing system using a video encoding device and a video decoding device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] In one embodiment, a motion vector encoding device includes: a prediction unit that uses a spatial candidate block and a temporal candidate block of a current block to obtain predicted motion vector candidates of predetermined multiple motion vector resolutions; and a coding unit that uses the predicted motion vector candidates to determine a predicted motion vector of the current block, a motion vector of the current block, and a motion vector resolution of the current block; and a coding unit that codes information indicating the predicted motion vector of the current block, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block, wherein the predetermined multiple motion vector resolutions include pixel-unit resolutions greater than a resolution of 1 pixel unit.
[0010] The prediction unit uses a first set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates to search for a reference block in pixel units of the first motion vector resolution, and uses a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates to search for a reference block in pixel units of the second motion vector resolution, wherein the first motion vector resolution and the second motion vector resolution are different resolutions, and the first set of motion vector predictor candidates and the second set of motion vector predictor candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate block and the temporal candidate block.
[0011] The prediction unit uses a first set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates to search for a reference block in pixel units of the first motion vector resolution, and uses a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates to search for a reference block in pixel units of the second motion vector resolution, wherein the first motion vector resolution and the second motion vector resolution are different resolutions, and the first set of motion vector predictor candidates and the second set of motion vector predictor candidates include different numbers of motion vector predictor candidates.
[0012] The encoding unit is characterized in that, if the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the residual motion vector is down-scaled and encoded according to the resolution of the motion vector of the current block.
[0013] The current block is a current coding unit constituting an image, and the motion vector resolution is determined to be the same for each coding unit. If a prediction unit predicted as an AMVP (advanced motion vector prediction) mode exists within the current coding unit, the encoding unit encodes information indicating the motion vector resolution of the prediction unit predicted as the AMVP mode once as information indicating the motion vector resolution of the current block.
[0014] The current block is a current coding unit that constitutes an image, and the motion vector resolution is determined to be the same for each prediction unit. If there is a prediction unit predicted as AMVP mode within the current coding unit, the encoding unit encodes information indicating the motion vector resolution for each prediction unit predicted as AMVP mode within the current block as information indicating the motion vector resolution of the current block.
[0015] In one embodiment, the motion vector encoding device includes a prediction unit that uses a spatial candidate block and a temporal candidate block of a current block to obtain candidate predictors of a plurality of motion vector resolutions, and determines a predicted motion vector of the current block, a motion vector of the current block, and a motion vector resolution of the current block using the candidate predictors; and an encoding unit that encodes information indicating the predicted motion vector of the current block, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block, wherein the prediction unit encodes one or more candidate predictors of the motion vectors selected from the candidate predictors. a first set of motion vector predictor candidates including motion vector predictor candidates, and searches for a reference block in pixel units of the first motion vector resolution; a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates, and searches for a reference block in pixel units of the second motion vector resolution; the first motion vector resolution and the second motion vector resolution are different resolutions, and the first set of motion vector predictor candidate and the second set of motion vector predictor candidate are obtained from different candidate blocks included in the spatial candidate block and the temporal candidate block, or include different numbers of motion vector predictor candidate.
[0016] In one embodiment, a motion vector encoding device generates a merge candidate list including at least one merge candidate related to a current block, determines and encodes a motion vector of the current block using a motion vector of one of the merge candidates included in the merge candidate list, and the merge candidate list includes motion vectors obtained by downscaling the motion vectors of the candidates included in the merge candidate list by a predetermined number of motion vector resolutions.
[0017] The downscaling is characterized in that, instead of the pixel indicated by the motion vector of the minimum motion vector resolution, one of the pixels located around the pixel indicated by the motion vector of the minimum motion vector resolution is selected based on the resolution of the motion vector of the current block, and adjusted to indicate the selected pixel.
[0018] In one embodiment, a motion vector decoding device includes an acquisition unit that uses spatial candidate blocks and temporal candidate blocks of a current block to acquire predicted motion vector candidates of predetermined multiple motion vector resolutions, acquires information indicating the predicted motion vector of the current block from the predicted motion vector candidates, and acquires a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block; and a decoding unit that restores the motion vector of the current block based on the residual motion vector, the information indicating the predicted motion vector of the current block, and the motion vector resolution information of the current block, wherein the predetermined multiple motion vector resolutions include pixel-unit resolutions greater than a resolution of 1 pixel unit.
[0019] The predicted motion vector candidates for the predetermined plurality of motion vector resolutions include a first set of predicted motion vector candidates including one or more predicted motion vector candidates for a first motion vector resolution, and a second set of predicted motion vector candidates including one or more predicted motion vector candidates for a second motion vector resolution, wherein the first motion vector resolution and the second motion vector resolution are different resolutions, and the first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks, or include different numbers of predicted motion vector candidates.
[0020] The decoding unit may restore the residual motion vector by upscaling it according to the minimum motion vector resolution if the pixel unit of the motion vector resolution of the current block is greater than the pixel unit of the minimum motion vector resolution.
[0021] The current block is a current coding unit constituting an image, and the motion vector resolution is determined to be the same for each coding unit. If a prediction unit predicted as AMVP mode exists within the current coding unit, the acquisition unit acquires information indicating the motion vector resolution of the prediction unit predicted as AMVP mode once from the bitstream as information indicating the motion vector resolution of the current block.
[0022] In one embodiment, a motion vector decoding device generates a merge candidate list including at least one merge candidate related to a current block, determines and decodes a motion vector for the current block using a motion vector of one of the merge candidates included in the merge candidate list, and the merge candidate list includes motion vectors obtained by downscaling the motion vectors of the candidates included in the merge candidate list by a predetermined number of motion vector resolutions.
[0023] An apparatus for decoding a motion vector according to an embodiment includes: an acquisition unit that acquires candidate predicted motion vectors of a plurality of predetermined motion vector resolutions using spatial candidate blocks and temporal candidate blocks of a current block; acquires information indicating a predicted motion vector of the current block from the candidate predicted motion vectors; acquires a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block; and a decoding unit that restores the motion vector of the current block based on the residual motion vector, the information indicating the predicted motion vector of the current block, and the motion vector resolution information of the current block. The motion vector predictor candidate set for the predetermined plurality of motion vector resolutions includes a first set of motion vector predictor candidates including one or more motion vector predictor candidate sets for a first motion vector resolution and a second set of motion vector predictor candidate sets including one or more motion vector predictor candidate sets for a second motion vector resolution, wherein the first motion vector resolution and the second motion vector resolution are different resolutions, and the first set of motion vector predictor candidate sets and the second set of motion vector predictor candidate sets are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks, or include different numbers of motion vector predictor candidate sets.
[0024] According to one embodiment, a computer-readable recording medium is provided that stores a program for causing a computer to execute the motion vector decoding method.
[0025] 1A to 7B, a video encoding apparatus and a video decoding apparatus, and a motion vector resolution encoding apparatus and decoding apparatus for the same, and a method thereof, according to an embodiment are proposed. Hereinafter, the video encoding apparatus and the method thereof may include a motion vector encoding apparatus and a motion vector encoding method, respectively, which will be described later. Also, the video decoding apparatus and the method thereof may include a motion vector decoding apparatus and a motion vector decoding method, respectively, which will be described later.
[0026] Also, a video encoding technique and a video decoding technique based on a tree-structured coding unit according to an embodiment applicable to the previously proposed video encoding method and video decoding method are disclosed with reference to Figures 8 to 20. Also, an embodiment applicable to the previously proposed video encoding method and video decoding method is disclosed with reference to Figures 21 to 27.
[0027] Hereinafter, "image" can refer to still images or moving images of a video, that is, the video itself.
[0028] Hereinafter, a "sample" refers to data assigned to a sampling position in an image and to data to be processed. For example, in a spatial domain image, a pixel is also a sample.
[0029] Hereinafter, the term "current block" refers to a block of a coding unit or a prediction unit of a current image to be coded or decoded.
[0030] Throughout this specification, when a part is described as "including" or "consisting of" a certain component, it does not mean that it excludes other components and may further include other components, unless specifically stated to the contrary. Furthermore, the term "module" used in this specification refers to software or a hardware component such as an FPGA or ASIC, and a "module" can perform a certain function. However, "module" is not limited to software or hardware. A "module" may be configured to reside on an addressable recording medium or to execute one or more processors. Thus, by way of example, "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided by a component or "module" may be combined into fewer components and "modules" or further separated into additional components and "modules."
[0031] First, referring to FIGS. 1A to 7B, an apparatus and method for encoding motion vectors for encoding video, and an apparatus and method for decoding motion vectors for decoding video, according to an embodiment, are disclosed.
[0032] FIG. 1A shows a block diagram of a motion vector encoding device 10 according to one embodiment.
[0033] In video encoding, inter-prediction refers to a prediction method that uses the similarity between a current image and another image. A reference area similar to a current area of the current image is detected from a reference image restored before the current image, the coordinate distance between the current area and the reference area is expressed as a motion vector, and the difference in pixel values between the current area and the reference area is expressed as residual data. Therefore, by performing inter-prediction on the current area, instead of directly outputting image information of the current area, an index indicating a reference image, a motion vector, and residual data are output, thereby improving encoding / decoding efficiency.
[0034] The motion vector encoding device 10 according to an embodiment may encode a motion vector used for inter-prediction for each block of each image of a video. The block type may be a square, a rectangle, or any other geometric shape. The block is not limited to a data unit of a certain size. According to an embodiment, a block may be a maximum coding unit, a coding unit, a prediction unit, a transform unit, or the like among coding units according to a tree structure. A video encoding / decoding method based on a coding unit according to a tree structure will be described below with reference to FIGS. 8 to 20.
[0035] The motion vector encoding device 10 may include a prediction unit 11 and an encoding unit 13.
[0036] The motion vector encoding device 10 according to an embodiment may encode a motion vector for inter prediction for each image block of a video.
[0037] The motion vector encoding device 10 can determine the motion vector of the current block by referring to the motion vector of a block different from the current block for motion vector prediction, block merging (PU merging), or advanced motion vector prediction (AMVP).
[0038] According to an embodiment, the motion vector encoding device 10 may determine a motion vector for a current block by referring to motion vectors of other blocks that are temporally or spatially adjacent to the current block. The motion vector encoding device 10 may determine prediction candidates including motion vectors of candidate blocks that are also reference targets for the motion vector of the current block. The motion vector encoding device 10 may determine a motion vector for the current block by referring to one motion vector selected from the prediction candidates.
[0039] According to one embodiment, the motion vector encoding device 10 divides a coding unit into prediction units, which are basic units of prediction divided from the coding unit, and searches for a prediction block that is most similar to the current coding unit in a reference picture adjacent to the current picture through motion estimation, and can determine motion parameters that indicate motion information between the current block and the prediction block.
[0040] According to an embodiment, the division of a prediction unit starts from a coding unit and is divided only once without dividing into a quadtree format. For example, one coding unit is divided into multiple prediction units, and the prediction units generated by the division are not further divided.
[0041] For motion estimation, the motion vector encoding device 10 can increase the resolution of the reference picture by up to n times (n is an integer) horizontally and up to n times vertically, and determine the motion vector of the current block with an accuracy of 1 / n pixel position, where n is referred to as the minimum motion vector resolution of the picture.
[0042] For example, if the pixel unit of the minimum motion vector resolution is 1 / 4 pixel, the motion vector encoding device 10 may increase the picture resolution by 4 times vertically and horizontally to determine a motion vector with an accuracy of up to 1 / 4 pixel position. However, depending on the characteristics of the image, determining a motion vector in 1 / 4 pixel units may not be sufficient, and conversely, determining a motion vector in 1 / 4 pixel units may be less efficient than determining a motion vector at 1 / 2 pixel positions. Therefore, the motion vector encoding device 10 according to an embodiment may adaptively determine the resolution of the motion vector of the current block and encode the determined predicted motion vector, actual motion vector, and motion vector resolution.
[0043] According to an embodiment, the prediction unit 11 may determine an optimal motion vector for inter prediction of the current block using one of the motion vector prediction candidates.
[0044] The prediction unit 11 may obtain motion vector predictor candidates with a plurality of motion vector resolutions using spatial candidate blocks and temporal candidate blocks of the current block, and may determine a motion vector predictor for the current block, a motion vector for the current block, and a motion vector resolution for the current block using the motion vector predictor candidates. The spatial candidate blocks may include at least one neighboring block spatially adjacent to the current block. Furthermore, the temporal candidate blocks may include at least one block located at the same position as the current block and a neighboring block spatially adjacent to the block located at the same position in a reference picture having a picture order count (POC) different from that of the current block. According to an embodiment, the prediction unit 11 may determine a motion vector for the current block by directly copying, combining, or modifying at least one motion vector predictor candidate.
[0045] The prediction unit 11 may use the motion vector predictor candidate to determine a motion vector predictor for the current block, a motion vector for the current block, and a motion vector resolution for the current block. The predetermined multiple motion vector resolutions may include pixel-unit resolutions greater than a one-pixel resolution. That is, the predetermined multiple motion vector resolutions may include resolutions of two-pixel units, three-pixel units, four-pixel units, etc. However, the predetermined multiple motion vector resolutions do not necessarily include resolutions of one pixel unit or greater, and may also be composed of only resolutions of one pixel unit or less.
[0046] The prediction unit 11 may search for a reference block in pixel units of a first motion vector resolution using a first set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates, and may search for a reference block in pixel units of the second motion vector resolution using a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates. The first and second motion vector resolutions may be different resolutions. The first and second motion vector predictor candidate sets are obtained from different candidate blocks among the spatial and temporal candidate blocks. The first and second motion vector predictor candidate sets may include different numbers of motion vector predictor candidates.
[0047] A method in which the prediction unit 11 uses candidate motion vector predictors with a plurality of predetermined motion vector resolutions to determine the motion vector predictor of the current block, the motion vector of the current block, and the motion vector resolution of the current block will be described later with reference to FIG. 4B.
[0048] The encoding unit 13 may encode information indicating a predicted motion vector of the current block, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block. The encoding unit 13 may encode the motion vector of the current block using fewer bits by using the residual motion vector between the actual motion vector and the predicted motion vector, thereby improving the compression rate of video encoding. The encoding unit 13 may encode an index indicating the motion vector resolution of the current block, as described below. In addition, the encoding unit 13 may downscale and encode the residual motion vector based on the difference between the minimum motion vector resolution and the motion vector resolution of the current block.
[0049] FIG. 1B shows a flowchart of a motion vector encoding method according to one embodiment.
[0050] In step 12, the motion vector encoding device 10 according to an embodiment may obtain motion vector predictor candidates of predetermined multiple motion vector resolutions by using motion vectors of spatial candidate blocks and temporal candidate blocks of the current block. The predetermined multiple motion vector resolutions may include pixel resolutions greater than one pixel resolution.
[0051] In step 14, the motion vector encoding device 10 according to one embodiment can use the predicted motion vector candidate obtained in step 12 to determine the predicted motion vector of the current block, the motion vector of the current block, and the motion vector resolution of the current block.
[0052] In step 16, the motion vector encoding device 10 according to one embodiment may encode information indicating the predicted motion vector determined in step 14, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating the motion vector resolution of the current block.
[0053] FIG. 2A shows a block diagram of a motion vector decoding device according to one embodiment.
[0054] The motion vector decoder 20 can parse the received bitstream and determine a motion vector for inter-prediction of the current block.
[0055] The acquisition unit 21 may acquire motion vector predictor candidates with predetermined multiple motion vector resolutions using spatial candidate blocks and temporal candidate blocks of the current block. The spatial candidate blocks may include at least one neighboring block spatially adjacent to the current block. Furthermore, the temporal candidate blocks may include at least one block located at the same position as the current block and one neighboring block spatially adjacent to the block located at the same position in a reference picture having a POC different from that of the current block. The predetermined multiple motion vector resolutions may include pixel-unit resolutions greater than one pixel unit resolution. That is, the predetermined multiple motion vector resolutions may include resolutions such as two-pixel units, three-pixel units, and four-pixel units. However, the predetermined multiple motion vector resolutions do not necessarily include resolutions greater than one pixel unit, and may also include resolutions less than one pixel unit.
[0056] The motion vector predictor candidates for a plurality of predetermined motion vector resolutions may include a first set of motion vector predictor candidates including one or more motion vector predictor candidates at a first motion vector resolution and a second set of motion vector predictor candidates including one or more motion vector predictor candidates at a second motion vector resolution. The first motion vector resolution and the second motion vector resolution are different resolutions. The first set of motion vector predictor candidates and the second set of motion vector predictor candidates are obtained from different candidate blocks among candidate blocks included in spatial candidate blocks and temporal candidate blocks. Furthermore, the first set of motion vector predictor candidate and the second set of motion vector predictor candidate may include different numbers of motion vector predictor candidates.
[0057] The acquisition unit 21 acquires information indicating the predicted motion vector of the current block from among the predicted motion vector candidates from the received bitstream, and can acquire information related to the motion vector of the current block, the residual motion vector between the predicted motion vector of the current block, and the motion vector resolution of the current block.
[0058] The decoding unit 23 can reconstruct the motion vector of the current block based on the residual motion vector acquired by the acquisition unit 21, information indicating the predicted motion vector of the current block, and motion vector resolution information of the current block. The decoding unit 23 can reconstruct the received data related to the residual motion vector by upscaling it based on the difference between the minimum motion vector resolution and the motion vector resolution of the current block.
[0059] 5A to 5D, a method for the motion vector decoding device 20 according to various embodiments to parse a received bitstream and obtain a motion vector resolution for a current block will be described below.
[0060] FIG. 2B shows a flowchart of a motion vector decoding method according to one embodiment.
[0061] In step 22, the motion vector decoding device 20 according to an embodiment can obtain motion vector predictor candidates of predetermined multiple motion vector resolutions using spatial candidate blocks and temporal candidate blocks of the current block.
[0062] In operation 24, the motion vector decoding device 20 according to an embodiment may obtain, from the bitstream, information indicating a predicted motion vector of the current block from among candidate predicted motion vectors, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block. The predetermined plurality of motion vector resolutions may include pixel-unit resolutions greater than a pixel-unit resolution.
[0063] In step 26, the motion vector decoding device 20 according to one embodiment can reconstruct the motion vector of the current block based on the residual motion vector obtained in step 24, information indicating the predicted motion vector of the current block, and motion vector resolution information of the current block.
[0064] FIG. 3A illustrates interpolation for motion compensation based on various resolutions.
[0065] The motion vector encoding device 10 can determine motion vectors of a predetermined resolution of multiple motion vectors for inter-predicting the current block. k The resolution may include a pixel unit (k is an integer). If k is greater than 0, the motion vector may indicate only some pixels in the reference image. If k is less than 0, an n-tap (n is an integer) finite impulse response (FIR) filter is used for interpolation to generate sub-pixel units, and the generated sub-pixel units may be indicated. For example, the motion vector encoding device 10 may determine the minimum motion vector resolution to be 1 / 4 pixel units, and the resolutions of a predetermined number of motion vectors to be 1 / 4, 1 / 2, 1, and 2 pixel units.
[0066] For example, an n-tap FIR filter can be used to perform interpolation to generate sub-pixels (a to l) in half-pixel units. For vertical half-pixels, interpolation can be performed using integer pixel units A1, A2, A3, A4, A5, and A6 to generate sub-pixel a, and integer pixel units B1, B2, B3, B4, B5, and B6 to generate sub-pixel b. Sub-pixels c, d, e, and f can be generated in the same way.
[0067] The pixel values of the horizontal sub-pixels are calculated as follows: a = (A1-5 x A2 + 20 x A3 + 20 x A4-5 x A5 + A6) / 32, b = (B1-5 x B2 + 20 x B3 + 20 x B4-5 x B5 + B6) / 32. The pixel values of the sub-pixels c, d, e, and f are calculated in the same way.
[0068] Similar to the horizontal sub-pixels, the vertical sub-pixels can also be generated by interpolation using a 6-tap FIR filter: A1, B1, C1, D1, E1, and F1 are used to generate sub-pixel g, and A2, B2, C2, D2, E2, and F2 are used to generate sub-pixel h.
[0069] The pixel values of the vertical sub-pixels are calculated in the same way as the horizontal sub-pixel values, for example, g = (A1 - 5 x B1 + 20 x C1 + 20 x D1 - 5 x E1 + F1) / 32.
[0070] The diagonal half-pixel sub-pixel m is interpolated using other half-pixel sub-pixels. In other words, the pixel value of sub-pixel m is calculated as m=(a-5×b+20×c+20×d-5×e+f) / 32.
[0071] If sub-pixels are generated in half-pixel units, then quarter-pixel sub-pixels can be generated using integer-pixel pixels and half-pixel sub-pixels. Sub-pixels in quarter-pixel units can also be generated by interpolating two adjacent pixels. Alternatively, quarter-pixel sub-pixels can be generated by applying an interpolation filter directly to integer-pixel pixel values without using half-pixel sub-pixel values.
[0072] Although the interpolation filter described above is a 6-tap filter, the motion vector encoding device 10 can interpolate pictures using filters with other numbers of taps. For example, the interpolation filter may include a 4-tap, 7-tap, 8-tap, or 12-tap filter.
[0073] As shown in FIG. 3A, when interpolation is performed on a reference picture to generate sub-pixels in 1 / 2 pixel units and sub-pixels in 1 / 4 pixel units, the interpolated reference picture is compared with the current block, and a block with the smallest SAD (sum of absolute difference) or rate-distortion cost is searched for in 1 / 4 pixel units, and a motion vector with a resolution in 1 / 4 pixel units is determined.
[0074] Figure 3B shows the motion vector resolutions of 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel units when the minimum motion vector resolution is 1 / 4 pixel. (a), (b), (c), and (d) in Figure 3B show the coordinates (shown as black rectangles) of pixels that can be indicated by motion vectors with resolutions of 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel units, respectively, based on the coordinate (0,0).
[0075] In one embodiment, the motion vector encoding device 10 performs motion compensation in sub-pixel units, and can search for a block similar to the current block in a reference picture in sub-pixel units based on a motion vector determined at an integer pixel position.
[0076] For example, the motion vector encoding device 10 can determine a motion vector at an integer pixel position, double the resolution of the reference picture, and search for the most similar prediction block in the range of (-1,-1) to (1,1) based on the motion vector determined at the integer pixel position. Next, the resolution is further doubled, and at four times the resolution, the motion vector at a half pixel position is used as a reference to search for the most similar prediction block in the range of (-1,-1) to (1,1), thereby finally determining a motion vector at a quarter pixel resolution.
[0077] For example, if the motion vector for an integer pixel position is (-4,-3) relative to the coordinates (0,0), then at half pixel resolution the motion vector becomes (-8,-6), and if it moves by (0,-1), the final motion vector for half pixel resolution is determined to be (-8,-7).Also, the motion vector at quarter pixel resolution is changed to (-16,-14), and if it moves further by (-1,0), the final motion vector for quarter pixel resolution is determined to be (-17,-14).
[0078] In order to perform motion compensation in units of pixels greater than one pixel, the motion vector encoding device 10 according to an embodiment may search for a block similar to a current block in a reference picture based on pixel positions greater than one pixel position, based on a motion vector determined at an integer pixel position. Hereinafter, pixel positions greater than one pixel position (e.g., 2 pixels, 3 pixels, 4 pixels) are referred to as super pixels.
[0079] For example, if a motion vector at an integer pixel position is (-4,-8) based on the coordinates (0,0), the motion vector is determined as (-2,-4) in a 2-pixel resolution. Although more bits are consumed to encode a motion vector in quarter-pixel units than a motion vector in integer pixel units, accurate inter-prediction in quarter-pixel units can be performed, thereby reducing the number of bits consumed to encode a residual block.
[0080] However, if interpolation is performed in pixel units smaller than 1 / 4 pixel units, for example, 1 / 8 pixel units, to generate sub-pixels and estimate motion vectors in 1 / 8 pixel units based on them, excessively many bits will be consumed in encoding the motion vectors, and the encoding compression rate will actually decrease.
[0081] Furthermore, when there is a lot of noise in the video or there is little texture, the resolution can be set in super-pixel units to perform motion estimation, thereby improving the encoding compression rate.
[0082] FIG. 4A shows candidate blocks for a current block for obtaining candidate motion vector predictors.
[0083] The prediction unit 11 may obtain one or more candidate motion vector predictors for the current block to perform motion prediction for the current block with respect to a reference picture of the current block to be coded. To obtain the candidate motion vector predictors, the prediction unit 11 may obtain at least one of a spatial candidate block and a temporal candidate block for the current block.
[0084] When the current block is predicted by referring to a reference frame having a different POC, the prediction unit 11 can obtain predicted motion vector candidates using blocks located around the current block, co-located blocks belonging to reference frames temporally different from the current block (with different POC), and neighboring blocks of the co-located blocks.
[0085] For example, the spatial candidate blocks may include at least one of the left block A1 411, the top block B1 412, the top left block B2 413, the top right block B0 414, and the bottom left block A0 425, which are adjacent blocks of the current block 410. The temporal candidate blocks may include at least one of the co-located block 430 and the adjacent block H 431 of the co-located block 430, which belong to a reference frame having a different POC from that of the current block. The motion vector encoding device 10 may obtain motion vectors of the temporal candidate blocks and the spatial candidate blocks as candidate motion vector predictors.
[0086] The prediction unit 11 may obtain motion vector predictor candidates with a plurality of predetermined motion vector resolutions. Each motion vector predictor candidate may have a different resolution. That is, the prediction unit 11 may use a first motion vector predictor candidate among the motion vector predictor candidates to search for a reference block in pixel units of the first motion vector resolution, and may use a second motion vector predictor candidate to search for a reference block in pixel units of the second motion vector resolution. The first motion vector predictor candidate and the second motion vector predictor candidate are obtained using different blocks among blocks belonging to spatial candidate blocks and temporal candidate blocks.
[0087] The prediction unit 11 can determine different numbers and types of sets of candidate blocks (that is, sets of motion vector predictor candidates) depending on the resolution of the motion vector.
[0088] For example, if the minimum motion vector resolution is ¼ pixel and ¼ pixel, ½ pixel, 1 pixel, and 2 pixel units are used as predetermined multiple motion vector resolutions, the prediction unit 11 may generate a predetermined number of motion vector predictor candidates for each resolution. The motion vector predictor candidates for each resolution are also motion vectors of different candidate blocks. The prediction unit 11 uses different candidate blocks for each resolution to obtain motion vector predictor candidates, thereby increasing the probability of finding an optimal reference block and reducing the rate / distortion cost, thereby improving coding efficiency.
[0089] To determine the motion vector of the current block, the prediction unit 11 may determine a search start position in a reference picture using each motion vector predictor candidate, and search for an optimal reference block based on the resolution of each motion vector predictor candidate. That is, once the motion vector predictor candidate of the current block is obtained, the motion vector encoding device 10 may search for a reference block in pixel units of a predetermined resolution corresponding to each motion vector predictor candidate, compare the rate and distortion costs based on the difference between the motion vector of the current block and each motion vector predictor, and determine the motion vector predictor having the smallest cost.
[0090] The encoding unit 13 can encode a residual motion vector, which is a difference vector between the determined predicted motion vector and the actual motion vector of the current block, and information indicating the motion vector resolution used for inter-prediction of the current block.
[0091] The encoding unit 13 may determine and encode the residual motion vector as shown in Equation (1). MVx is the x-component of the actual motion vector of the current block, and MVy is the y-component of the actual motion vector of the current block. pMVx is the x-component of the predicted motion vector of the current block, and pMVy is the y-component of the predicted motion vector of the current block. MVDx is the x-component of the residual motion vector of the current block, and MVDy is the y-component of the residual motion vector of the current block.
[0092] MVDx=MVx-pMVx MVDy=MVy-pMVy (1) The decoding unit 23 can reconstruct the motion vector of the current block using the information indicating the predicted motion vector acquired from the bitstream and the residual motion vector. The decoding unit 23 can determine the final motion vector by adding the predicted motion vector and the residual motion vector as shown in Equation (2).
[0093] MVx=pMVx+MVDx MVy=pMCy+MVDy (2) If the minimum motion vector resolution is in sub-pixel units, motion vector encoding device 10 can multiply the predicted motion vector and the actual motion vector by an integer value to represent the motion vector as an integer value. If a predicted motion vector with a resolution of 1 / 4 pixel units starting from coordinates (0,0) represents coordinates (1 / 2,3 / 2) and the minimum motion vector resolution is in 1 / 4 pixel units, motion vector encoding device 10 can encode the vector (2,6), which is the value obtained by multiplying the predicted motion vector by the integer 4, as the predicted motion vector. If the minimum motion vector resolution is in 1 / 8 pixel units, motion vector encoding device 10 can multiply the predicted motion vector by the integer 8 to represent the vector (4,12) as the predicted motion vector.
[0094] FIG. 4B illustrates a process for generating motion vector predictor candidates according to an embodiment.
[0095] As described above, the prediction unit 11 can obtain motion vector predictor candidates with a plurality of predetermined motion vector resolutions.
[0096] For example, if the minimum motion vector resolution is ¼ pixel and multiple predetermined motion vector resolutions include ¼ pixel, ½ pixel, 1 pixel, and 2 pixel, the predictor 11 may generate a predetermined number of motion vector predictor candidates for each resolution. The predictor 11 may configure different sets 460, 470, 480, and 490 of motion vector predictor candidates according to the predetermined multiple motion vector resolutions. Each set 460, 470, 490 may include motion vector predictors derived from different candidate blocks or may include different numbers of motion vector predictors. By using different candidate blocks for each resolution, the predictor 11 may increase the probability of finding an optimal predictor block and reduce the rate-distortion cost, thereby improving coding efficiency.
[0097] The quarter-pixel resolution motion vector predictor candidate 460 is determined from two different temporal or spatial candidate blocks. For example, the predictor 11 may obtain the motion vector of the left block 411 of the current block and the motion vector of the top block 412 as quarter-pixel motion vector predictor candidate 460, and search for an optimal motion vector for each motion vector predictor candidate. That is, the predictor 11 may determine a search start position using the motion vector of the left block 411 of the current block, search for reference blocks in quarter-pixel units, and determine an optimal motion vector predictor and reference block. The predictor 11 may also determine a search start position using the motion vector of the top block 412, search for reference blocks in quarter-pixel units, and determine another optimal motion vector predictor and reference block.
[0098] The half-pixel resolution motion vector predictor candidate 470 is determined from one temporal or spatial candidate block. The half-pixel resolution motion vector predictor candidate is a different motion vector predictor candidate from the quarter-pixel resolution motion vector predictor candidate. For example, the prediction unit 11 may obtain the motion vector of the top right block 414 of the current block as the half-pixel motion vector predictor candidate 470. That is, the prediction unit 11 may determine a search start position using the motion vector of the top right block 414, search for reference blocks in half-pixel units, and determine another optimal motion vector predictor and reference block.
[0099] The motion vector predictor candidate 480 for pixel-by-pixel resolution is determined from one temporal or spatial candidate block. The motion vector predictor for pixel-by-pixel resolution is also a motion vector predictor different from the motion vector predictors 460 and 470 used for the quarter-pixel resolution and half-pixel resolution. For example, the prediction unit 11 can determine the motion vector of the temporal candidate block 430 as the one-pixel motion vector predictor candidate 480. That is, the prediction unit 11 can determine a search start position using the motion vector of the temporal candidate block 430, search for a reference block in pixel units, and determine another optimal motion vector predictor and reference block.
[0100] The predicted motion vector candidate 490 for a resolution in two pixel units is determined from one temporal or spatial candidate block. The predicted motion vector for a resolution in two pixel units is also a predicted motion vector different from the predicted motion vector used for other resolutions. For example, the prediction unit 11 can determine the motion vector of the bottom left block 425 as the two-pixel predicted motion vector 490. That is, the prediction unit 11 can determine a search start position using the motion vector of the bottom left block 425, search for reference blocks in two-pixel units, and determine another optimal predicted motion vector and reference block.
[0101] The prediction unit 11 compares the rate and distortion cost based on the motion vector of the current block and each predicted motion vector, and can finally determine the motion vector of the current block, one predicted motion vector, and one motion vector resolution 495.
[0102] When the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the motion vector encoding device 10 can downscale and encode the residual motion vector by the resolution of the motion vector of the current block. Also, when the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the motion vector decoding device 20 can upscale and restore the residual motion vector by the minimum motion vector resolution.
[0103] When the minimum motion vector resolution is 1 / 4 pixel units and the motion vector resolution of the current block is determined to be 1 / 2 pixel units, the encoding unit 13 can calculate the residual motion vector by adjusting the determined actual motion vector and predicted motion vector to 1 / 2 pixel units in order to reduce the size of the residual motion vector.
[0104] The encoding unit 13 may halve the sizes of the actual motion vector and the predicted motion vector, and may also halve the size of the residual motion vector. That is, if the minimum motion vector resolution is in 1 / 4 pixel units, the predicted motion vector (MVx, MVy), which is expressed by multiplying it by 4, may be further divided by 2 to express the predicted motion vector. For example, if the minimum motion vector resolution is in 1 / 4 pixel units and the predicted motion vector is (-24, -16), the motion vector with a resolution of 1 / 2 pixel units will be (-12, -8), and the motion vector with a resolution of 2 pixel units will be (-3, -2). The following equation (3) shows a process of halving the sizes of the actual motion vector (MVx, MVy) and the predicted motion vector (pMVx, pMCy) using a bit shift operation.
[0105] MVDx=(MVx)>>1-(pMVx)>>1 MVDy=(MVy)>>1-(pMCy)>>1 (3) The decoding unit 23 may determine the final motion vector (MVx, MVy) for the current block by adding the received residual motion vector (MVDx, MVDy) to the finally determined predicted motion vector (pMVx, pMCy). The decoding unit 23 may restore the motion vector for the current block by upscaling the predicted motion vector and the residual motion vector obtained by using a bit shift operation as shown in Equation 4.
[0106] MVx=(pMVx)<<1+(MVDx)<<1 MVy=(pMCy)<<1+(MVDy)<<1 (4) According to an embodiment, the encoding unit 13 may be configured to: n pixel, and for the current block, k When pixel-resolution motion vectors are determined, the residual motion vector can be determined using equation (5) as follows:
[0107] MVDx=(MVx)>>(k+n)-(pMVx)>>(k+n) MVDy=(MVy)>>(k+n)-(pMCy)>>(k+n) (5) The decoding unit 23 can determine the final motion vector of the current block as shown in Equation (6) by adding the received residual motion vector to the finally determined predicted motion vector.
[0108] MVx=(pMVx)<<(k+n)+(MVDx)<<(k+n) MVy=(pMCy)<<(k+n)+(MVDy)<<(k+n) (6) If the magnitude of the residual motion vector is reduced, the number of bits required to represent the residual motion vector is reduced, improving coding efficiency.
[0109] As described with reference to FIGS. 3A and 3B, kFor pixel-wise motion vectors, if k is less than 0, the reference picture undergoes interpolation to generate pixels at non-integer positions. Conversely, for motion vectors where k is greater than 0, not every pixel in the reference picture is interpolated. k Therefore, if the motion vector resolution of the current block is equal to or greater than one pixel unit, the decoding unit 23 can omit interpolation of the reference picture according to the motion vector resolution of the current block to be decoded.
[0110] FIG. 5A illustrates coding units and prediction units according to one embodiment.
[0111] If the current block is a current coding unit constituting an image, the same motion vector resolution for inter prediction is determined for each coding unit, and one or more prediction units predicted as AMVP mode exist within the current coding unit, the encoding unit 13 may encode information indicating the motion vector resolution of the prediction unit predicted as AMVP mode only once as information indicating the motion vector resolution of the current block, and transmit the encoded information to the motion vector decoding device 20. The acquiring unit 21 may acquire information indicating the motion vector resolution of the prediction unit predicted as AMVP mode only once from the bitstream as information indicating the motion vector resolution of the current block.
[0112] 5B shows a part of a prediction_unit syntax for transmitting adaptively determined motion vector resolution according to an embodiment. Fig. 5B shows syntax defining an operation of the motion vector decoding device 20 according to an embodiment to obtain information indicating the motion vector resolution of the current block.
[0113] For example, as shown in FIG. 5A, if the current coding unit 560 is 2Nx2N in size and the prediction unit 563 is the same size (2Nx2N) as the coding unit 560 and is predicted using AMVP mode, the coding unit 13 codes the information indicating the motion vector resolution of the coding unit 560 once for the current coding unit 560, and the acquisition unit 21 can acquire the information indicating the motion vector resolution of the coding unit 560 once from the bitstream.
[0114] If the size of the current coding unit 570 is 2Nx2N and the current coding unit 570 is divided into two prediction units 573 and 577 of 2NxN size, since the prediction unit 573 is predicted as the merge mode, the encoding unit 13 does not transmit information indicating the motion vector resolution related to the prediction unit 573, whereas since the prediction unit 577 is predicted as the AMVP mode, the encoding unit 13 transmits the information indicating the motion vector resolution related to the prediction unit 573 once, and the acquiring unit 21 can acquire the information indicating the motion vector resolution related to the prediction unit 577 once from the bitstream. That is, the encoding unit 13 transmits the information indicating the motion vector resolution related to the prediction unit 577 once as information indicating the motion vector resolution of the current coding unit 570, and the decoding unit 23 can receive the information indicating the motion vector resolution related to the prediction unit 577 once as information indicating the motion vector resolution of the current coding unit 570. The information indicating the motion vector resolution may also be in the form of an index indicating one of a predetermined number of motion vectors, such as "cu_resolution_idx[x0][y0]".
[0115] 5B, the initial values of "parsedMVResolution 510," which is information indicating whether or not the motion vector resolution has been extracted for the current coding unit, and "mv_resolution_idx 512," which is information indicating the motion vector resolution of the current coding unit, can each be set to 0. First, in the case of prediction unit 573, since it is predicted as the merge mode, condition 513 is not satisfied, and therefore the acquisition unit 21 does not receive information "cu_resolution_idx[x0][y0]" indicating the motion vector resolution of the current prediction unit.
[0116] In the case of prediction unit 577, since it is predicted to be in AMVP mode, condition 513 is satisfied, and since "parsedMVResolution" has a value of 0, condition 514 is satisfied, and the acquisition unit 21 can receive "cu_resolution_idx[x0][y0] 516". Since "cu_resolution_idx[x0][y0]" has been received, "parsedMVResolution" is set to 1 (518). The received "cu_resolution_idx[x0][y0]" is saved in "mv_resolution_idx" (520).
[0117] If the size of the current coding unit 580 is 2Nx2N and it is divided into two prediction units 583 and 587 of 2NxN size, and if both prediction units 583 and 587 are predicted in AMVP mode, the coding unit 13 encodes and transmits information indicating the motion vector resolution of the current coding unit 580 once, and the acquisition unit 21 can receive information indicating the motion vector resolution of the coding unit 580 once from the bitstream.
[0118] 5B , in the case of prediction unit 583, condition 513 is satisfied and 'parsedMVResolution' has a value of 0, so condition 514 is satisfied and the acquisition unit 21 can acquire 'cu_resolution_idx[x0][y0] 516'. Since the decoding unit 23 has acquired 'cu_resolution_idx[x0][y0]', it sets 'parsedMVResolution' to 1 (518) and stores the acquired 'cu_resolution_idx[x0][y0]' in 'mv_resolution_idx' (520). However, in the case of prediction unit 587, 'parsedMVResolution' already has a value of 1, so condition statement 514 cannot be satisfied and the acquisition unit 21 does not acquire 'cu_resolution_idx[x0][y0] 516'. In other words, since the acquisition unit 21 has already received information indicating the motion vector resolution related to the current coding unit 580 from the prediction unit 583, there is no need to acquire information indicating the motion vector resolution (the same as the motion vector resolution of the prediction unit 583) from the prediction unit 587.
[0119] If the size of the current coding unit 590 is 2Nx2N and it is divided into two prediction units 593 and 597 of 2NxN size, and if both prediction units 593 and 597 are predicted in merge mode, then conditional statement 513 is not satisfied, and therefore the acquisition unit 21 does not acquire information indicating the motion vector resolution for the current coding unit 590.
[0120] According to another embodiment, the encoding unit 13 may transmit, as information indicating the motion vector resolution of the current block for each prediction unit predicted in AMVP mode within the current block, information indicating the motion vector resolution for each prediction unit predicted in AMVP mode within the current block to the motion vector decoding device 20. The acquiring unit 21 may acquire, from the bitstream, information indicating the motion vector resolution for each prediction unit within the current block as information indicating the motion vector resolution of the current block.
[0121] 5C illustrates a portion of a prediction_unit syntax for transmitting adaptively determined motion vector resolution according to another embodiment, which defines an operation of the motion vector decoding device 20 according to another embodiment to obtain information indicating the motion vector resolution of the current block.
[0122] The syntax of Figure 5C differs from the syntax of Figure 5B in that there is no 'parsedMVResolution 510' indicating whether or not a motion vector resolution has been extracted for the current coding unit, so the acquisition unit 21 acquires information 'cu_resolution_idx[x0,y0]' indicating the motion vector resolution for each prediction unit present in the current coding unit (524).
[0123] For example, referring again to FIG. 5A, if the size of the current coding unit 580 is 2Nx2N and is divided into two prediction units 583 and 587 of 2NxN size, and both prediction units 583 and 587 are predicted in AMVP mode, and the motion vector resolution of prediction unit 583 is 1 / 4 pixel and the motion vector resolution of prediction unit 587 is 2 pixel, the acquisition unit 21 acquires information indicating the two motion vector resolutions for one coding unit 580 (i.e., the motion vector resolution 1 / 2 of prediction unit 583 and the motion vector resolution 1 / 4 of prediction unit 587) (524).
[0124] In another embodiment, when the encoding unit 13 adaptively determines the motion vector resolution for the current prediction unit regardless of the prediction mode of the prediction unit, the encoding unit 13 encodes and transmits information indicating the motion vector resolution for each prediction unit regardless of the prediction mode, and the acquisition unit 11 can acquire the information indicating the motion vector resolution from the bitstream for each prediction unit regardless of the prediction mode.
[0125] 5D illustrates a portion of a prediction_unit syntax according to another embodiment for transmitting adaptively determined motion vector resolution, which defines an operation of the motion vector decoding device 20 according to another embodiment to obtain information indicating the motion vector resolution of the current block.
[0126] While the syntax described in Figures 5B and 5C assumes that the adaptive motion vector resolution determination method is applied only when the prediction mode is AMVP mode, Figure 5D shows an example of syntax assuming that the adaptive motion vector resolution determination method is applied regardless of the prediction mode. Referring to the syntax of Figure 5D, the acquisition unit 21 receives 'cu_resolution_idx[x0][y0]' for each prediction unit regardless of the prediction mode of the prediction unit (534).
[0127] The index 'cu_resolution_idx[x0][y0]' indicating the motion vector resolution described with reference to Figures 5B to 5D is coded as a unary or fixed length before transmission. Also, when only two motion vector resolutions are used, 'cu_resolution_idx[x0][y0]' can also be flag-format data.
[0128] The motion vector encoding device 10 can adaptively configure multiple predetermined motion vector resolutions used for encoding on a slice-by-slice or block-by-block basis. Furthermore, the motion vector decoding device 10 can adaptively configure multiple predetermined motion vector resolutions used for decoding on a slice-by-slice or block-by-block basis. The multiple predetermined motion vector resolutions adaptively configured on a slice-by-slice or block-by-block basis can be used as a motion vector resolution candidate set. That is, the motion vector encoding device 10 and the motion vector decoding device 20 can configure different types and numbers of motion vector candidates for the current block based on information about neighboring blocks that have already been encoded or decoded.
[0129] For example, when the minimum motion vector resolution is ¼ pixel, the motion vector encoding device 10 or the motion vector decoding device 20 may use ¼ pixel, ½ pixel, 1 pixel, and 2 pixel resolutions as a set of motion vector resolution candidates fixed uniformly for all images. Instead of using a set of motion vector resolution candidates fixed uniformly for all images, the motion vector encoding device 10 or the motion vector decoding device 20 may use ⅛ pixel, ¼ pixel, and ½ pixel resolutions as a set of motion vector resolution candidates for the current block when the motion vector resolution of a neighboring block already coded is small, and may use ½ pixel, 1 pixel, and 2 pixel resolutions as a set of motion vector resolution candidates for the current block when the motion vector resolution of a neighboring block is large. The motion vector encoding device 10 or the motion vector decoding device 20 may also be configured to vary the type and number of motion vector resolution candidates for each slice or each block based on the size of the motion vector and other information.
[0130] The types and number of resolutions constituting the motion vector resolution candidate set may be always set and used in exactly the same way by the motion vector encoding device 10 and the motion vector decoding device 20, or may be inferred in the same way based on information about neighboring blocks and other information. Alternatively, information about the motion vector resolution candidate set used by the motion vector encoding device 10 may be coded into a bitstream and explicitly transmitted to the motion vector decoding device 20.
[0131] FIG. 6A illustrates one embodiment of constructing a merge candidate list using multiple resolutions.
[0132] In order to reduce the amount of data related to motion information transmitted for each prediction unit, the motion vector encoding device 10 may use a merge mode in which motion information of a current block is set based on motion information of spatial / temporal neighboring blocks. The motion vector encoding device 10 may configure the same merge candidate list for predicting motion information in the encoding device and the decoding device, and transmit candidate selection information in the list to the decoding device, thereby effectively reducing the amount of motion-related data.
[0133] When the current block has a prediction unit that is predicted using merge mode, the motion vector decoding device 20 constructs a merge candidate list for the current block in the same manner as the motion vector encoding device 10, obtains candidate selection information in the list from the bitstream, and decodes the motion vector of the current block.
[0134] The merge candidate list may include spatial candidates based on motion information of spatially surrounding blocks and temporal candidates based on motion information of temporally surrounding blocks. The motion vector encoding device 10 and the motion vector decoding device 20 may include spatial candidates and temporal candidates of a plurality of predetermined motion vector resolutions in a merge candidate list in a predetermined order.
[0135] The motion vector resolution determination method described with reference to Figures 3A to 4B is not only applicable when the prediction unit of the current block is encoded using AMVP mode, but also applies when using a prediction mode (e.g., merge mode) in which one predicted motion vector among the candidate predicted motion vectors is immediately used as the final motion vector without transmitting a residual motion vector.
[0136] That is, the motion vector encoding device 10 and the motion vector decoding device 20 may adjust the motion vector predictor candidate to be suitable for a predetermined plurality of motion vector resolutions and determine the adjusted motion vector predictor candidate as the motion vector of the current block. That is, the motion vectors of the candidate blocks included in the merge candidate list may include motion vectors obtained by downscaling a motion vector of the minimum motion vector resolution by a predetermined plurality of motion vector resolutions. The downscaling method will be described later with reference to Figures 7A and 7B.
[0137] For example, assume that the minimum motion vector resolution is 1 / 4 pixel, the plurality of predetermined motion vector resolution candidates are 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 2 pixel, and the merge candidate list is configured as (A1, B1, B0, A0, B2, co-located block). The acquisition unit 21 of the motion vector decoding device 20 configures a 1 / 4 resolution predicted motion vector candidate as shown in FIG. 6A, and adjusts the 1 / 4 resolution predicted motion vector to 1 / 2 pixel, 1 pixel, and 2 pixel resolutions, thereby sequentially acquiring a merge candidate list of the plurality of resolutions in the order of resolution.
[0138] FIG. 6B illustrates another embodiment for constructing a merge candidate list using multiple resolutions.
[0139] The motion vector encoding device 10 and the motion vector decoding device 20 can also construct predicted motion vector candidates at a resolution of 1 / 4 pixel units, as shown in Figure 6B, adjust each predicted motion vector candidate at a resolution of 1 / 4 pixel units to a resolution of 1 / 2 pixel, 1 pixel, and 2 pixel units, and sequentially obtain a merge candidate list according to the order of the predicted motion vector candidates.
[0140] When there is a prediction unit for which the current block is predicted using a merge mode, the motion vector decoding device 20 can determine a motion vector for the current block based on the merge candidate lists for multiple resolutions acquired by the acquisition unit 21 and information related to the merge candidate index acquired from the bitstream.
[0141] FIG. 7A shows pixels indicated by two motion vectors with different resolutions.
[0142] In order to adjust a high-resolution motion vector to a corresponding low-resolution motion vector, the motion vector encoding device 10 can adjust the motion vector so that it points to a neighboring pixel instead of the pixel pointed to by the existing high-resolution motion vector. Selecting one of the neighboring pixels is called rounding.
[0143] For example, to adjust a quarter-pixel resolution motion vector indicating (19,27) based on coordinates (0,0) to a full-pixel resolution motion vector, the quarter-pixel resolution motion vector (19,27) is divided by the integer 4, resulting in rounding during the division. For ease of explanation, it is assumed below that the motion vectors for each resolution start from coordinates (0,0) and indicate coordinates (x,y) (x and y are integers).
[0144] 7A, the resolution of the smallest motion vector is 1 / 4 pixel, and to adjust the 1 / 4 pixel resolution motion vector 715 to a full pixel resolution motion vector, four integer pixels 720, 730, 740, and 750 around pixel 710 indicated by the 1 / 4 pixel motion vector 715 also become candidate pixels indicated by the corresponding full pixel motion vectors 725, 735, 745, and 755. That is, if the value of coordinate 710 is (19,27), then coordinate 1020 becomes (7,24), coordinate 730 becomes (16,28), coordinate 740 becomes (20,28), and coordinate 750 becomes (20,24).
[0145] When adjusting the quarter-pixel resolution motion vector 710 to a corresponding full-pixel resolution motion vector, the motion vector encoding device 10 according to an embodiment can determine that the quarter-pixel resolution motion vector 710 points to the top right integer pixel 740. That is, if the quarter-pixel resolution motion vector starts from coordinates (0,0) and points to coordinates (19,27), the corresponding full-pixel resolution motion vector starts from coordinates (0,0) and points to coordinates (20,28), and the final full-pixel resolution motion vector becomes (5,7).
[0146] In one embodiment, the motion vector encoding device 10 adjusts a high-resolution motion vector to a low-resolution motion vector so that the adjusted low-resolution motion vector always points to the upper right edge of the pixel indicated by the high-resolution motion vector. In another embodiment, the motion vector encoding device 10 adjusts the adjusted low-resolution motion vector so that the adjusted low-resolution motion vector always points to the upper left edge, lower left edge, or lower right edge of the pixel indicated by the high-resolution motion vector.
[0147] The motion vector encoding device 10 can select a different pixel to be pointed to by the corresponding low-resolution motion vector from among the four pixels located around the pixel pointed to by the high-resolution motion vector, namely the upper left corner, the upper right corner, the lower left corner, and the lower right corner, depending on the resolution of the motion vector of the current block.
[0148] For example, referring to FIG. 7B, the half-pixel motion vector can be adjusted to point to pixel 1080 at the top left of pixel 1060 pointed to by the quarter-pixel motion vector, the full-pixel motion vector can be adjusted to point to pixel 1070 at the top right of the pixel pointed to by the quarter-pixel motion vector, and the two-pixel motion vector can be adjusted to point to pixel 1090 at the bottom right of the pixel pointed to by the quarter-pixel motion vector.
[0149] When the motion vector encoding device 10 is configured to point to one of the surrounding pixels instead of the pixel pointed to by an existing high-resolution motion vector, it can determine the position of the pixel pointed to based on at least one of the resolution, 1 / 4-pixel motion vector candidates, information on surrounding blocks, encoding information, and an arbitrary pattern.
[0150] Meanwhile, for the sake of convenience, in Figures 3A to 7B, only the operations performed by the motion vector encoding device 10 are described, and the operations of the motion vector decoding device 20 are omitted, or only the operations performed by the motion vector decoding device 20 are described, and the operations of the motion vector encoding device 10 are omitted. However, it will be easily understood by a person skilled in the art to which this embodiment pertains that each motion vector encoding device 10 and motion vector decoding device 20 also performs operations corresponding to those of each motion vector decoding device 20 and motion vector encoding device 10.
[0151] Hereinafter, a video encoding method and apparatus therefor, and a video decoding method and apparatus therefor based on a tree-structured coding unit and a transform unit according to an embodiment will be disclosed with reference to Figures 8 to 20. The motion vector encoding device 10 described with reference to Figures 1A to 7B may be included in the video encoding device 800. That is, the motion vector encoding device 10 may encode information indicating a predicted motion vector, a residual motion vector, and information indicating a motion vector resolution for performing inter-prediction on an image to be encoded by the video encoding device 800 using the methods described with reference to Figures 1A to 7B.
[0152] FIG. 8 illustrates a block diagram of a video encoding device 800 based on a tree-structured coding unit according to one embodiment of the present invention.
[0153] A video encoding device 800 with video prediction based on a coding unit with a tree structure according to an embodiment includes a coding unit determination unit 820 and an output unit 830. Hereinafter, for convenience of description, the video encoding device 800 with video prediction based on a coding unit with a tree structure according to an embodiment will be abbreviated to "video encoding device 800."
[0154] The coding unit determination unit 820 may partition the current picture based on a maximum coding unit, which is a coding unit of the largest size for the current picture of the video. If the current picture is larger than the maximum coding unit, the video data of the current picture is divided into at least one maximum coding unit. According to an embodiment, the maximum coding unit may be a data unit of size 32x32, 64x64, 128x128, 256x256, etc., or may be a square data unit whose vertical and horizontal dimensions are a power of 2.
[0155] According to an embodiment, a coding unit is characterized by a maximum size and a depth. The depth indicates the number of times a coding unit is spatially divided from the maximum coding unit, and as the depth increases, the coding units for each depth are divided from the maximum coding unit to the minimum coding unit. The depth of the maximum coding unit is defined as the highest depth, and the minimum coding unit is defined as the lowest coding unit. As the depth of the maximum coding unit increases, the size of the coding units for each depth decreases, so a coding unit of a higher depth may include multiple coding units of lower depths.
[0156] As described above, the image data of the current picture may be divided into maximum coding units according to the maximum size of the coding unit, and each maximum coding unit may include coding units divided by depth. Since the maximum coding units according to an embodiment are divided by depth, image data of the spatial domain included in the maximum coding units may be hierarchically classified by depth.
[0157] There are preset maximum depths that limit the total number of times the height and width of the largest coding unit can be hierarchically divided, and maximum coding unit sizes.
[0158] The coding unit determination unit 820 encodes at least one divided region obtained by dividing the region of the largest coding unit for each depth, and determines a depth at which a final coding result is output for each of the at least one divided region. That is, the coding unit determination unit 820 encodes video data in coding units for each depth for each largest coding unit of the current picture, selects a depth at which a minimum coding error occurs, and determines it as a final depth. The determined final depth and video data for each largest coding unit are output to the output unit 830.
[0159] The video data in the maximum coding unit is coded based on coding units for each depth with at least one depth equal to or less than the maximum depth, and the coding results based on the coding units for each depth are compared. After comparing the coding errors of the coding units for each depth, the depth with the smallest coding error is selected. At least one final depth is determined for each maximum coding unit.
[0160] As the depth of the maximum coding unit increases, the coding unit is divided into layers, and the number of coding units increases. Even for coding units of the same depth included in one maximum coding unit, the coding error for each piece of data is measured to determine whether to divide it into sub-depths. Therefore, even for data included in one maximum coding unit, the coding error for each depth varies depending on the position, so the final depth is determined differently depending on the position. Therefore, one or more final depths are set for one maximum coding unit, and the data of the maximum coding unit is partitioned by coding units of one or more final depths.
[0161] Therefore, according to one embodiment, the coding unit determination unit 820 determines a tree-structured coding unit included in the current largest coding unit. According to one embodiment, the "tree-structured coding unit" includes a coding unit of a depth determined as the final depth among all depth-specific coding units included in the current largest coding unit. The coding unit of the final depth is determined hierarchically according to depth within the same region within the largest coding unit, and is determined independently for other regions. Similarly, the final depth for the current region is determined independently of the final depths for other regions.
[0162] According to one embodiment, the maximum depth is an index related to the number of divisions from the largest coding unit to the smallest coding unit. According to one embodiment, the first maximum depth may indicate the total number of divisions from the largest coding unit to the smallest coding unit. According to one embodiment, the second maximum depth may indicate the total number of depth levels from the largest coding unit to the smallest coding unit. For example, if the depth of the largest coding unit is 0, the depth of a coding unit obtained by dividing the largest coding unit once is set to 1, and the depth of a coding unit obtained by dividing the largest coding unit twice is set to 2. In this case, if the coding unit obtained by dividing the largest coding unit four times is the smallest coding unit, depth levels of 0, 1, 2, 3, and 4 exist, so the first maximum depth is set to 4 and the second maximum depth is set to 5.
[0163] Predictive coding and transformation are performed for the maximum coding unit. Predictive coding and transformation are also performed for each maximum coding unit and for each depth less than the maximum depth based on the depth-specific coding unit.
[0164] Since the number of coding units for each depth increases each time the maximum coding unit is divided by depth, coding including predictive coding and transform must be performed on all coding units for each depth generated as the depth increases. For convenience of explanation, predictive coding and transform will be described below based on a coding unit of a current depth among at least one maximum coding unit.
[0165] The video encoding device 800 according to an embodiment may select various sizes or shapes of data units for encoding video data. The video data may be encoded through steps such as predictive encoding, transform, and entropy encoding, and the same data unit may be used throughout all steps, or the data unit may be changed for each step.
[0166] For example, the video encoding device 800 can select not only a coding unit for encoding video data, but also a data unit different from the coding unit to perform predictive encoding of the video data of the coding unit.
[0167] For predictive coding of the maximum coding unit, predictive coding is performed based on a coding unit of the final depth, i.e., a coding unit that is not further divided, according to an embodiment. Hereinafter, a coding unit that is the basis for predictive coding and is not further divided will be referred to as a "prediction unit." A partition into which the prediction unit is divided may include the prediction unit and data units into which at least one of the height and width of the prediction unit is divided. The partition is a data unit into which a prediction unit of a coding unit is divided, and the prediction unit is also a partition of the same size as the coding unit.
[0168] For example, if a coding unit of size 2Nx2N (where N is a positive integer) is not further divided, it becomes a prediction unit of size 2Nx2N, and the partition size may be 2Nx2N, 2NxN, Nx2N, NxN, etc. Partition modes according to one embodiment may selectively include not only symmetric partitions in which the height or width of the prediction unit is divided at a symmetric ratio, but also partitions divided at an asymmetric ratio such as 1:n or n:1, partitions divided in a geometric shape, partitions of an arbitrary shape, etc.
[0169] The prediction mode of a prediction unit may be at least one of intra mode, inter mode, and skip mode. For example, intra mode and inter mode are performed on partitions of 2Nx2N, 2NxN, Nx2N, and NxN sizes. Also, skip mode is performed only on partitions of 2Nx2N size. Each prediction unit within a coding unit is independently coded, and a prediction mode with the smallest coding error is selected.
[0170] In addition, the video encoding device 800 according to an embodiment may convert video data of a coding unit based on not only a coding unit for encoding video data but also a data unit different from the coding unit. To convert a coding unit, conversion is performed based on a transform unit that is smaller than or equal to the coding unit. For example, the transform unit may include a data unit for an intra mode and a transform unit for an inter mode.
[0171] In one embodiment, in a manner similar to the tree-structured coding unit, the transform units within the coding unit are also recursively divided into smaller transform units, and the residual data of the coding unit is partitioned by the tree-structured transform units according to the transform depth.
[0172] According to an embodiment, the height and width of a coding unit are divided, and a transformation depth indicating the number of divisions required to reach the transformation unit is set for the transformation unit. For example, if the size of the transformation unit of a current coding unit having a size of 2Nx2N is 2Nx2N, the transformation depth is set to 0; if the size of the transformation unit is NxN, the transformation depth is set to 1; and if the size of the transformation unit is N / 2xN / 2, the transformation depth is set to 2. That is, for the transformation unit, a tree-structured transformation unit is set depending on the transformation depth.
[0173] The depth-based division information requires not only depth but also prediction-related information and transform-related information. Therefore, the coding unit determination unit 820 may determine not only the depth at which the minimum coding error occurs, but also a partition mode for dividing the prediction unit into partitions, a prediction mode for each prediction unit, a size of the transform unit for transform, etc.
[0174] A method for determining coding units and prediction units / partitions based on a tree structure of the largest coding unit, and a transform unit according to an embodiment will be described in detail with reference to FIGS. 17 to 19. FIG.
[0175] The coding unit determination unit 820 may measure the coding error of the coding unit for each depth using a rate-distortion optimization technique based on a Lagrangian multiplier.
[0176] The output unit 830 outputs the maximum coding unit of image data coded based on at least one depth determined by the coding unit determination unit 820 and depth-based division information in the form of a bitstream.
[0177] The encoded video data is also the result of encoding video residual data.
[0178] The depth-based partition information may include depth information, partition mode information of a prediction unit, prediction mode information, partition information of a transform unit, and the like.
[0179] The final depth information is defined using depth-specific division information indicating whether to encode using a coding unit of a lower depth instead of encoding using the current depth. If the current depth of the current coding unit is depth, the current coding unit is encoded using the coding unit of the current depth, and therefore the division information of the current depth is defined so that it is not further divided into lower depths. Conversely, if the current depth of the current coding unit is not depth, encoding using a coding unit of a lower depth must be attempted, and therefore the division information of the current depth is defined so that it is divided into coding units of lower depths.
[0180] If the current depth is not a depth, coding is performed on the coding units divided into coding units of lower depths. Since there are one or more coding units of lower depths within the coding unit of the current depth, coding is performed repeatedly for each coding unit of lower depth, and recursive coding is performed for each coding unit of the same depth.
[0181] Since a tree-structured coding unit is determined within one maximum coding unit and at least one piece of partition information must be determined for each depth coding unit, at least one piece of partition information is determined for one maximum coding unit. Also, since data of the maximum coding unit is hierarchically partitioned according to depth and the depth varies depending on the position, depth and partition information are set for the data.
[0182] Therefore, the output unit 830 according to an embodiment may allocate coding information related to the depth and coding mode to at least one of the coding unit, the prediction unit, and the smallest unit included in the largest coding unit.
[0183] According to one embodiment, the minimum unit is a square data unit having a size obtained by dividing the minimum coding unit, which is the lowest depth, into four. According to one embodiment, the minimum unit is also a square data unit of the largest size included in all coding units, prediction units, partition units, and transform units included in the maximum coding unit.
[0184] For example, the coding information output through the output unit 830 is classified into coding information for coding units according to depth and coding information for prediction units. The coding information for coding units according to depth may include prediction mode information and partition size information. The coding information transmitted for each prediction unit may include information related to an estimation direction in inter mode, information related to a reference picture index in inter mode, information related to a motion vector, information related to a chroma component in intra mode, information related to an interpolation method in intra mode, etc.
[0185] Information regarding the maximum size of a coding unit defined for each picture, slice, or GOP, and information regarding the maximum depth are inserted into a bitstream header, a sequence parameter set, or a picture parameter set.
[0186] In addition, information regarding the maximum size of the transform unit allowed for the current video and information regarding the minimum size of the transform unit are also output via a bitstream header, a sequence parameter set, a picture parameter set, etc. The output unit 830 may encode and output reference information, prediction information, slice type information, etc. related to prediction.
[0187] According to the simplest embodiment of the video encoding device 800, a coding unit for each depth is a coding unit having a size that is half the height and width of a coding unit for a next higher depth. That is, if the size of a coding unit for a current depth is 2Nx2N, the size of a coding unit for a lower depth is NxN. Also, a current coding unit of 2Nx2N size includes up to four coding units for a lower depth, each of which is NxN.
[0188] Therefore, the video encoding device 800 may determine a coding unit of an optimal shape and size for each maximum coding unit based on the size and maximum depth of the maximum coding unit determined in consideration of the characteristics of the current picture, and may construct the coding units according to a tree structure. Also, since each maximum coding unit can be coded using various prediction modes, conversion methods, etc., the optimal coding mode is determined in consideration of the image characteristics of the coding units of various image sizes.
[0189] Therefore, if an image with a very high resolution or a large amount of data is encoded using the existing macroblock unit, the number of macroblocks per picture will be excessively large. As a result, the amount of compression information generated for each macroblock will also increase, which increases the transmission burden of the compression information and tends to reduce data compression efficiency. Therefore, a video encoding device according to an embodiment can increase the maximum size of a coding unit in consideration of the size of an image and adjust the coding unit in consideration of image characteristics, thereby improving image compression efficiency.
[0190] FIG. 9 illustrates a block diagram of a video decoding device 900 based on a tree-structured coding unit, according to one embodiment.
[0191] The motion vector decoding device 20 described with reference to Figures 2A to 7B may be included in the video decoding device 900. That is, the motion vector decoding device 20 may receive and parse information indicating a predicted motion vector, a residual motion vector, and information indicating a motion vector resolution for performing inter prediction of an image decoded by the video decoding device 900 from a bitstream related to coded video, and may reconstruct a motion vector based on the parsed information.
[0192] A video decoding device 900 with video prediction based on a coding unit with a tree structure according to an embodiment includes a receiving unit 910, a video data and coding information extracting unit 920, and a video data decoding unit 930. Hereinafter, for convenience of explanation, the video decoding device 900 with video prediction based on a coding unit with a tree structure according to an embodiment will be referred to simply as "video decoding device 900."
[0193] The definitions of various terms such as coding unit, depth, prediction unit, transform unit, and various partition information for the decoding operation of the video decoding device 900 according to one embodiment are the same as those described with reference to FIG. 8 and the video encoding device 800.
[0194] The receiving unit 910 receives and parses a bitstream related to coded video. The video data and coding information extracting unit 920 extracts coded video data for each coding unit according to a tree structure for each maximum coding unit from the parsed bitstream and outputs the extracted video data to the video data decoding unit 930. The video data and coding information extracting unit 920 may extract information related to the maximum size of the coding unit of the current picture from a header, sequence parameter set, or picture parameter set related to the current picture.
[0195] The video data and coding information extraction unit 920 also extracts final depth and partition information related to the tree-structured coding units for each maximum coding unit from the parsed bitstream. The extracted final depth and partition information are output to the video data decoding unit 930. That is, the video data of the bitstream is divided into maximum coding units, and the video data decoding unit 930 decodes the video data for each maximum coding unit.
[0196] The maximum depth and partition information for each coding unit may be set for one or more pieces of depth information, and the partition information for each depth may include partition mode information of the corresponding coding unit, prediction mode information, partition information of a transform unit, etc. Also, the partition information for each depth may be extracted as the depth information.
[0197] The depth and partition information for each maximum coding unit extracted by the video data and coding information extraction unit 920 is depth and partition information determined by repeatedly encoding each coding unit for each maximum coding unit depth at the encoding end to generate a minimum coding error, as in the video encoding device 800 according to an embodiment of the present invention. Therefore, the video decoding device 900 can restore the image by decoding the data using an encoding method that generates a minimum coding error.
[0198] According to an embodiment, coding information related to depth and coding mode is assigned to a predetermined data unit among the coding unit, prediction unit, and minimum unit, so that the video data and coding information extraction unit 920 can extract depth and partition information for each predetermined data unit. If the depth and partition information of the maximum coding unit is recorded for each predetermined data unit, predetermined data units having the same depth and partition information are inferred as data units included in the same maximum coding unit.
[0199] The video data decoder 930 decodes video data of each largest coding unit based on the depth and partition information for each largest coding unit to restore a current picture. That is, the video data decoder 930 may decode video data encoded for each coding unit based on the partition mode, prediction mode, and transform unit determined for each coding unit based on the tree structure included in the largest coding unit. The decoding process may include a prediction process including intra prediction and motion compensation, and an inverse transform process.
[0200] The video data decoder 930 may perform intra prediction or motion compensation for each coding unit according to the partition and prediction mode information of the prediction unit of the depth-based coding unit.
[0201] In addition, the video data decoder 930 may read transform unit information according to a tree structure for each coding unit to perform inverse transform for each coding unit, and perform inverse transform based on the transform unit for each coding unit. Through the inverse transform, pixel values in the spatial domain of the coding unit are restored.
[0202] The video data decoder 930 may determine the depth of the current largest coding unit using the depth-specific partition information. If the partition information indicates that no further partitioning will occur at the current depth, the current depth is the depth. Therefore, the video data decoder 930 may decode the coding unit of the current depth for the video data of the current largest coding unit using the partition mode, prediction mode, and transform unit size information of the prediction unit.
[0203] That is, the coding information set for a predetermined data unit from among the coding unit, prediction unit, and minimum unit is observed, and data units having coding information including the same partition information are collected and regarded as one data unit to be decoded in the same coding mode by the video data decoder 930. For each coding unit determined in this way, information related to the coding mode is obtained, and decoding of the current coding unit is performed.
[0204] FIG. 10 illustrates the concept of a coding unit according to one embodiment.
[0205] Examples of coding units, where the size of a coding unit is expressed as width x height, may include coding units of size 64x64, to 32x32, 16x16, and 8x8. A 64x64 coding unit may be partitioned into partitions of sizes 64x64, 64x32, 32x64, and 32x32, a 32x32 coding unit may be partitioned into partitions of sizes 32x32, 32x16, 16x32, and 16x16, a 16x16 coding unit may be partitioned into partitions of sizes 16x16, 16x8, 8x16, and 8x8, and an 8x8 coding unit may be partitioned into partitions of sizes 8x8, 8x4, 4x8, and 4x4.
[0206] For video data 1010, the resolution is set to 1920x1080, the maximum size of the coding unit is set to 64, and the maximum depth is set to 2. For video data 1020, the resolution is set to 1920x1080, the maximum size of the coding unit is set to 64, and the maximum depth is set to 3. For video data 1030, the resolution is set to 352x288, the maximum size of the coding unit is set to 16, and the maximum depth is set to 1. The maximum depth shown in Figure 10 indicates the total number of divisions from the maximum coding unit to the minimum coding unit.
[0207] When the resolution is high or the amount of data is large, it is desirable to have a relatively large maximum encoding size, not only to improve encoding efficiency but also to accurately reflect video characteristics. Therefore, the maximum encoding size of the video data 1010 and 1020, which have higher resolution than the video data 1030, is selected to be 64.
[0208] Since the maximum depth of the video data 1010 is 2, the coding units 1015 of the video data 1010 may be divided twice from the maximum coding unit with a major axis size of 64 to a depth of two layers, and may include coding units with major axis sizes of 32 and 16. Meanwhile, since the maximum depth of the video data 1030 is 1, the coding units 1035 of the video data 1030 may be divided once from the maximum coding unit with a major axis size of 16 to a depth of one layer, and may include coding units with major axis sizes of 8.
[0209] Since the maximum depth of the video data 1020 is 3, the coding units 1025 of the video data 1020 may be divided three times from the maximum coding unit with a major axis size of 64, and may include coding units with major axis sizes of 32, 16, and 8, which are three levels deeper than the maximum coding unit with a major axis size of 64. The deeper the depth, the better the ability to express detailed information.
[0210] FIG. 11 illustrates a block diagram of a coding unit-based video encoder 1100 according to one embodiment.
[0211] The video encoder 1100 according to an embodiment performs the same operations as those performed by the picture encoder 1520 of the video encoder 800 to encode video data. That is, the intra predictor 1120 performs intra prediction for each prediction unit of intra-mode coding units of the current image 1105, and the inter predictor 1115 performs inter prediction for each prediction unit of inter-mode coding units using the current image 1105 and a reference image acquired from the reconstructed picture buffer 1110. The current image 1105 is divided into maximum coding units and then sequentially encoded. At this time, encoding is performed on coding units obtained by dividing the maximum coding unit into a tree structure.
[0212] Residue data is generated by removing prediction data for a coding unit of each mode output from the intra prediction unit 1120 or the inter prediction unit 1115 from data for a coding unit to be encoded of the current image 1105. The residue data is output as transform coefficients quantized for each transform unit through the transform unit 1125 and the quantization unit 1130. The quantized transform coefficients are restored to spatial domain residue data through the inverse quantization unit 1145 and the inverse transform unit 1150. The restored spatial domain residue data is added to prediction data for a coding unit of each mode output from the intra prediction unit 1120 or the inter prediction unit 1115 to restore spatial domain data for the coding unit of the current image 1105. The restored spatial domain data is generated as a restored image through the deblocking unit 1155 and the SAO performing unit 1160. The generated restored image is stored in the restored picture buffer 1110. The reconstructed image stored in the reconstructed picture buffer 1110 is used as a reference image for inter-prediction of other images. The transform coefficients quantized by the transform unit 1125 and the quantization unit 1130 are output as a bitstream 1140 via an entropy coding unit 1135.
[0213] Because the video encoding unit 1100 according to one embodiment is applied to the video encoding device 800, the components of the video encoding unit 1100, such as the inter prediction unit 1115, intra prediction unit 1120, transform unit 1125, quantization unit 1130, entropy encoding unit 1135, inverse quantization unit 1145, inverse transform unit 1150, deblocking unit 1155 and SAO performing unit 1160, can perform operations based on each coding unit among the tree-structured coding units for each maximum coding unit.
[0214] In particular, the intra prediction unit 1120 and the inter prediction unit 1115 determine the partition mode and prediction mode of each coding unit among the tree-structured coding units taking into account the maximum size and maximum depth of the current largest coding unit, and the transform unit 1125 can determine whether to divide the transform units according to a quadtree within each coding unit among the tree-structured coding units.
[0215] FIG. 12 illustrates a block diagram of a coding unit-based video decoder 1200 according to one embodiment.
[0216] The entropy decoding unit 1215 parses the coded video data to be decoded and coding information required for decoding from the bitstream 1205. The coded video data is quantized transform coefficients, and the inverse quantization unit 1220 and the inverse transform unit 1225 restore residue data from the quantized transform coefficients.
[0217] The intra prediction unit 1240 performs intra prediction for each prediction unit for intra-mode coding units, and the inter prediction unit 1235 performs inter prediction for each prediction unit for inter-mode coding units of the current picture using reference pictures acquired from the reconstructed picture buffer 1230.
[0218] By adding the predicted data for the coding unit of each mode, which has passed through the intra prediction unit 1240 or the inter prediction unit 1235, and the residue data, spatial domain data for the coding unit of the current image 1105 is restored, and the restored spatial domain data is output as a restored image 1260 via the deblocking unit 1245 and the SAO performing unit 1250. In addition, the restored image stored in the restored picture buffer 1230 is output as a reference image.
[0219] In order for the picture decoder 930 of the video decoding device 900 to decode video data, the steps following the entropy decoder 1215 of the video decoder 1200 according to an embodiment are performed.
[0220] Because the video decoding unit 1200 is applied to the video decoding device 900 according to one embodiment, the components of the video decoding unit 1200, such as the entropy decoding unit 1215, the inverse quantization unit 1220, the inverse transform unit 1225, the intra prediction unit 1240, the inter prediction unit 1235, the deblocking unit 1245 and the SAO performing unit 1250, can perform operations based on each coding unit among the tree-structured coding units for each maximum coding unit.
[0221] In particular, the intra prediction unit 1240 and the inter prediction unit 1235 determine the partition mode and prediction mode for each coding unit among the coding units based on the tree structure, and the inverse transform unit 1225 can determine whether to divide the transform unit based on the quadtree structure for each coding unit.
[0222] FIG. 13 illustrates coding units and partitions by depth according to one embodiment.
[0223] The video encoding device 800 and the video decoding device 900 according to an embodiment use hierarchical coding units to take into account image characteristics. The maximum height, width, and depth of the coding unit are adaptively determined according to image characteristics and may be variously set according to user requests. The size of the coding unit for each depth is determined according to a preset maximum size of the coding unit.
[0224] The coding unit hierarchical structure 1300 according to one embodiment illustrates a case where the maximum height and width of the coding units are 64 and the maximum depth is 3. In this case, the maximum depth indicates the total number of divisions from the maximum coding unit to the minimum coding unit. As the depth increases along the vertical axis of the coding unit hierarchical structure 1300 according to one embodiment, the height and width of the coding units for each depth are divided. Furthermore, along the horizontal axis of the coding unit hierarchical structure 1300, prediction units and partitions that are the basis for predictive coding of each coding unit for each depth are illustrated.
[0225] That is, coding unit 1310 is the largest coding unit in coding unit hierarchical structure 1300, has a depth of 0, and has a coding unit size, i.e., height and width, of 64x64. Depth increases along the vertical axis, with depth 1 coding unit 1320 having a size of 32x32, depth 2 coding unit 1330 having a size of 16x16, and depth 3 coding unit 1340 having a size of 8x8. Depth 3 coding unit 1340 having a size of 8x8 is the smallest coding unit.
[0226] The prediction units and partitions of the coding units are arranged along the horizontal axis for each depth. That is, if a 64x64 coding unit 1310 at depth 0 is a prediction unit, the prediction unit is divided into a 64x64 partition 1310, a 64x32 partition 1312, a 32x64 partition 1314, and a 32x32 partition 1316 included in the 64x64 coding unit 1310.
[0227] Similarly, the prediction unit of a coding unit 1320 of size 32x32 at depth 1 is divided into a partition 1320 of size 32x32, a partition 1322 of size 32x16, a partition 1324 of size 16x32, and a partition 1326 of size 16x16, all of which are included in the coding unit 1320 of size 32x32.
[0228] Similarly, the prediction unit of a coding unit 1330 of size 16x16 at depth 2 is divided into a partition 1330 of size 16x16, a partition 1332 of size 16x8, a partition 1334 of size 8x16, and a partition 1336 of size 8x8 contained in the coding unit 1330 of size 16x16.
[0229] Similarly, the prediction unit of a coding unit 1340 of size 8x8 at depth 3 is divided into a partition 1340 of size 8x8, a partition 1342 of size 8x4, a partition 1344 of size 4x8, and a partition 1346 of size 4x4 contained in the coding unit 1340 of size 8x8.
[0230] In order to determine the depth of the largest coding unit 1310, the coding unit determination unit 820 of the video encoding device 800 according to one embodiment must perform encoding for each coding unit of each depth included in the largest coding unit 1310.
[0231] The number of coding units for each depth to contain data of the same range and size increases as the depth increases. For example, data containing one coding unit for depth 1 requires four coding units for depth 2. Therefore, to compare the coding results of the same data by depth, it must be coded using one coding unit for depth 1 and four coding units for depth 2.
[0232] For each depth-specific coding, coding is performed for each prediction unit of the depth-specific coding unit along the horizontal axis of the coding unit hierarchical structure 1300, and a representative coding error, which is the minimum coding error at that depth, is selected. Also, as the depth increases along the vertical axis of the coding unit hierarchical structure 1300, coding is performed for each depth, and the representative coding errors for each depth are compared to find the minimum coding error. In the largest coding unit 1310, the depth and partition at which the minimum coding error occurs are selected as the depth and partition mode of the largest coding unit 1310.
[0233] FIG. 14 illustrates the relationship between coding units and transform units, according to one embodiment.
[0234] The video encoding device 800 or the video decoding device 900 according to an embodiment encodes or decodes video using coding units that are smaller than or equal to the maximum coding unit for each maximum coding unit. During the encoding process, the size of the transform unit for transform is selected based on a data unit that is not larger than each coding unit.
[0235] For example, in the video encoding device 800 according to an embodiment or the video decoding device 900 according to an embodiment, when the current coding unit 1410 has a size of 64x64, a transform unit 1420 having a size of 32x32 is used for transforming.
[0236] Furthermore, data of the 64x64 size coding unit 1410 is transformed and coded using transform units of sizes smaller than 64x64, 32x32, 16x16, 8x8, and 4x4, and then the transform unit with the smallest error from the original is selected.
[0237] FIG. 15 illustrates the encoding information according to one embodiment.
[0238] The output unit 830 of the video encoding device 800 according to one embodiment may encode and transmit information 1500 related to the partition mode, information 1510 related to the prediction mode, and information 1520 related to the transform unit size as partition information for each coding unit of each depth.
[0239] The partition mode information 1500 indicates information regarding the type of partitions into which a prediction unit of the current coding unit is divided as a data unit for predictive coding of the current coding unit. For example, a current coding unit CU_0 having a size of 2Nx2N is used by being divided into one of a partition 1502 having a size of 2Nx2N, a partition 1504 having a size of 2NxN, a partition 1506 having a size of NxN, and a partition 1508 having a size of NxN. In this case, the partition mode information 1500 of the current coding unit is set to indicate one of the partition 1502 having a size of 2Nx2N, the partition 1504 having a size of 2NxN, the partition 1506 having a size of NxN, and the partition 1508 having a size of NxN.
[0240] The information about prediction modes 1510 indicates a prediction mode for each partition. For example, the information about prediction modes 1510 indicates that the partition indicated by the information about partition modes 1500 is to be predictively coded in one of intra mode 1512, inter mode 1514, and skip mode 1516.
[0241] The information about the transform unit size 1520 indicates the transform unit based on which the current coding unit is transformed, for example, the transform unit may be one of a first intra transform unit size 1522, a second intra transform unit size 1524, a first inter transform unit size 1526, and a second inter transform unit size 1528.
[0242] The image data and coding information extraction unit 1610 of the video decoding device 900 according to one embodiment can extract information 1500 related to the partition mode, information 1510 related to the prediction mode, and information 1520 related to the transformation unit size for each depth-based coding unit and use them for decoding.
[0243] FIG. 16 illustrates coding units for each depth according to an embodiment.
[0244] To indicate a change in depth, partition information is used, which indicates whether a coding unit of a current depth is divided into coding units of a lower depth.
[0245] A prediction unit 1610 for predictive coding of a depth 0 and 2N_0x2N_0 size coding unit 1600 may include a 2N_0x2N_0 size partition mode 1612, a 2N_0xN_0 size partition mode 1614, an N_0x2N_0 size partition mode 1616, and an N_0xN_0 size partition mode 1618. Although only partitions 1612, 1614, 1616, and 1618 in which the prediction unit is divided into symmetric ratios are illustrated, as mentioned above, the partition modes are not limited thereto and may include asymmetric partitions, arbitrary partitions, geometric partitions, etc.
[0246] For each partition mode, predictive coding is iteratively performed on one 2N_0x2N_0 size partition, two 2N_0xN_0 size partitions, two N_0x2N_0 size partitions, or four N_0xN_0 size partitions. For partitions of size 2N_0x2N_0, size N_0x2N_0, size 2N_0xN_0, and size N_0xN_0, predictive coding is performed in intra mode and inter mode. In skip mode, predictive coding is performed only on partitions of size 2N_0x2N_0.
[0247] If the coding error from one of the partition modes 1612, 1614, 1616 of sizes 2N_0x2N_0, 2N_0xN_0 and N_0x2N_0 is minimal, then there is no need to further subdivide.
[0248] If the encoding error with the partition mode 1618 of size N_0xN_0 is the smallest, then the depth 0 is changed to 1 and divided 1620), and encoding is iteratively performed on the coding unit 1630 of the partition mode of depth 2 and size N_0xN_0 to search for the smallest encoding error.
[0249] A prediction unit 1640 for predictive coding of a coding unit 1630 of depth 1 and size 2N_1x2N_1 (=N_0xN_0) may include a partition mode of size 2N_1x2N_1 1642, a partition mode of size 2N_1xN_1 1644, a partition mode of size N_1x2N_1 1646, and a partition mode of size N_1xN_1 1648.
[0250] Also, if the encoding error with the partition mode 1648 of size N_1xN_1 is the smallest, the depth 1 is changed to depth 2 and the division is performed 1650), and encoding is repeatedly performed on the coding unit 1660 of depth 2 and size N_2xN_2 to search for the smallest encoding error.
[0251] When the maximum depth is d, coding units by depth are set up to depth d-1, and partition information is set up to depth d-2. That is, when partitioning 1670 is performed from depth d-2 and coding is performed up to depth d-1, a prediction unit 1690 for predictive coding of a coding unit 1680 of depth d-1 and size 2N_(d-1)x2N_(d-1) may include a partition mode 1692 of size 2N_(d-1)x2N_(d-1), a partition mode 1694 of size 2N_(d-1)xN_(d-1), a partition mode 1696 of size N_(d-1)x2N_(d-1), and a partition mode 1698 of size N_(d-1)xN_(d-1).
[0252] Among the partition modes, encoding is performed iteratively via predictive coding for one partition of size 2N_(d-1)x2N_(d-1), two partitions of size 2N_(d-1)xN_(d-1), two partitions of size N_(d-1)x2N_(d-1), and four partitions of size N_(d-1)xN_(d-1), and the partition mode that generates the minimum encoding error is searched for.
[0253] Even if the coding error due to the partition mode 1698 of size N_(d-1)xN_(d-1) is minimum, since the maximum depth is d, the coding unit CU_(d-1) of depth d-1 does not undergo any further partitioning process to lower depths, and the depth related to the current maximum coding unit 1600 is determined to be depth d-1, and the partition mode is determined to be N_(d-1)xN_(d-1). Also, since the maximum depth is d, partition information is not set for the coding unit 1652 of depth d-1.
[0254] The data unit 1699 is a "smallest unit" related to the current largest coding unit. According to one embodiment, the smallest unit is a square data unit having a size obtained by dividing the smallest coding unit, which is the lowest depth, into four. Through this iterative coding process, the video encoding device 800 according to one embodiment compares coding errors for each depth of the coding unit 1600, selects the depth at which the smallest coding error occurs, determines the depth, and sets the corresponding partition mode and prediction mode as the coding mode for the depth.
[0255] In this way, the minimum coding error for each depth of all depths 0, 1, ..., d-1, d is compared, and the depth with the minimum error is selected and determined as the depth. The depth, partition mode, and prediction mode of the prediction unit are coded and transmitted as partition information. Also, since the coding unit must be partitioned from depth 0 to depth , only the partition information of the depth is set to '0', and the partition information for each depth excluding the depth must be set to '1'.
[0256] The video data and coding information extraction unit 920 of the video decoding device 900 according to an embodiment may extract information related to a depth and a prediction unit related to the coding unit 1600 and use the extracted information for decoding the coding unit 1612. The video decoding device 900 according to an embodiment may use depth-specific partition information to identify a depth having partition information of “0” as a depth and use partition information related to the depth for decoding.
[0257] Figures 17, 18 and 19 illustrate the relationship between coding units, prediction units and transform units according to one embodiment.
[0258] The coding units 1710 are coding units for each depth determined by the video encoding device 800 according to an embodiment with respect to the maximum coding unit. The prediction units 1760 are partitions of the prediction units of the coding units for each depth among the coding units 1710, and the transform units 1770 are transform units of the coding units for each depth.
[0259] Assuming that the depth of the maximum coding unit for depth-specific coding units 1710 is 0, coding units 1712 and 1754 have a depth of 1, coding units 1714, 1716, 1718, 1728, 1750, and 1752 have a depth of 2, coding units 1720, 1722, 1724, 1726, 1730, 1732, and 1748 have a depth of 3, and coding units 1740, 1742, 1744, and 1746 have a depth of 4.
[0260] In the prediction unit 1760, some partitions 1714, 1716, 1722, 1732, 1748, 1750, 1752, and 1754 are formed by dividing the coding unit. That is, partitions 1714, 1722, 1750, and 1754 are in a 2NxN partition mode, partitions 1716, 1748, and 1752 are in an Nx2N partition mode, and partition 1732 is in an NxN partition mode. The prediction units and partitions of the depth-specific coding unit 1710 are smaller than or the same as the respective coding units.
[0261] In the transform unit 1770, the video data of a transform unit 1752 is transformed or inversely transformed using a data unit of a size smaller than the coding unit. Also, the transform units 1714, 1716, 1722, 1732, 1748, 1750, 1752, and 1754 are data units of different sizes or shapes compared to the corresponding prediction units and partitions in the prediction unit 1760. That is, the video encoding device 800 according to an embodiment and the video decoding device 900 according to another embodiment perform intra prediction / motion estimation / motion compensation operations and transform / inverse transform operations related to the same coding unit based on different data units.
[0262] Thus, for each largest coding unit, coding units of a hierarchical structure for each region are recursively coded, and an optimal coding unit is determined, thereby constructing a coding unit having a recursive tree structure. The coding information may include partition information, partition mode information, prediction mode information, and transform unit size information related to the coding unit. Table 1 below shows an example that can be set in the video encoding device 800 and the video decoding device 900 according to an embodiment.
[0263] [Table 1] The output unit 830 of the video encoding device 800 according to one embodiment outputs encoding information related to a coding unit based on a tree structure, and the encoding information extraction unit 920 of the video decoding device 900 according to one embodiment can extract encoding information related to a coding unit based on a tree structure from a received bitstream.
[0264] The partition information indicates whether the current coding unit is divided into coding units of lower depths. If the partition information of the current depth d is 0, the depth at which the current coding unit is not further divided into lower coding units is the depth, and therefore, partition mode information, prediction mode, and transform unit size information are defined for the depth. If further division is required according to the partition information, each of the four divided coding units of lower depths must be coded independently.
[0265] The prediction mode can be represented by one of intra mode, inter mode, and skip mode. The intra mode and inter mode are defined for all partition modes, while the skip mode is only defined for the 2Nx2N partition mode.
[0266] The partition mode information may indicate symmetric partition modes 2Nx2N, 2NxN, Nx2N, and NxN in which the height or width of the prediction unit is divided at a symmetric ratio, and asymmetric partition modes 2NxnU, 2NxnD, nLx2N, and nRx2N in which the height or width of the prediction unit is divided at an asymmetric ratio. The asymmetric partition modes 2NxnU and 2NxnD indicate that the height is divided at a ratio of 1:3 and 3:1, respectively, and the asymmetric partition modes nLx2N and nRx2N indicate that the width is divided at a ratio of 1:3 and 3:1, respectively.
[0267] The transform unit size is set to two sizes in intra mode and two sizes in inter mode. That is, if the transform unit split information is 0, the size of the transform unit is set to 2Nx2N, the size of the current coding unit. If the transform unit split information is 1, the transform unit of the size into which the current coding unit is divided is set. Also, if the partition mode related to the current coding unit of size 2Nx2N is a symmetric partition mode, the size of the transform unit is set to NxN, and if it is an asymmetric partition mode, the size is set to N / 2xN / 2.
[0268] According to an embodiment, coding information of a coding unit having a tree structure is assigned to at least one of a depth coding unit, a prediction unit, and an atomic unit. The depth coding unit may include one or more prediction units and atomic units having the same coding information.
[0269] Therefore, by checking the coding information held by each of adjacent data units, it can be determined whether they are included in a coding unit of the same depth. Also, by using the coding information held by each data unit, it is possible to determine the coding unit of the corresponding depth, so that the depth distribution within the maximum coding unit can be inferred.
[0270] Therefore, in this case, when the current coding unit makes a prediction by referring to a neighboring data unit, coding information of a data unit in a depth-dependent coding unit adjacent to the current coding unit is directly referenced and used.
[0271] In another embodiment, when predictive coding is performed on a current coding unit by referring to a neighboring coding unit, the neighboring coding unit may also be referenced by searching for data adjacent to the current coding unit within the depth-specific coding unit using coding information of the neighboring depth-specific coding unit.
[0272] FIG. 20 illustrates the relationship between the coding unit, prediction unit, and transform unit according to the coding mode information in Table 1.
[0273] The maximum coding unit 2000 includes depth coding units 2002, 2004, 2006, 2012, 2014, 2016, and 2018. Among them, one coding unit 2018 is a depth coding unit, and therefore its partition information is set to 0. The partition mode information of the coding unit 2018 of size 2Nx2N is set to one of partition modes 2Nx2N 2022, 2NxN 2024, Nx2N 2026, NxN 2028, 2NxnU 2032, 2NxnD 2034, nLx2N 2036, and nRx2N 2038.
[0274] The transform unit partition information (TU size flag) is a type of transform index, and the size of the transform unit corresponding to the transform index changes depending on the prediction unit type or partition mode of the coding unit.
[0275] For example, when the partition mode information is set to one of symmetric partition modes 2Nx2N 2022, 2NxN 2024, Nx2N 2026 and NxN 2028, if the transform unit division information is 0, a transform unit 2042 of size 2Nx2N is set, and if the transform unit division information is 1, a transform unit 2044 of size NxN is set.
[0276] When the partition mode information is set to one of the asymmetric partition modes 2NxnU 2032, 2NxnD 2034, nLx2N 2036, and nRx2N 2038, if the transform unit division information (TU size flag) is 0, a transform unit 2052 of size 2Nx2N is set, and if the transform unit division information is 1, a transform unit 2054 of size N / 2xN / 2 is set.
[0277] The transform unit division information (TU size flag) described with reference to Fig. 20 is a flag having a value of 0 or 1, but the transform unit division information according to an embodiment is not limited to a 1-bit flag and may be increased to 0, 1, 2, 3, etc. depending on the setting, thereby dividing the transform units hierarchically. The transform unit division information is used as an embodiment of a transform index.
[0278] In this case, the size of the transform units actually used can be expressed by using the transform unit partition information according to an embodiment together with the maximum size and minimum size of the transform units. The video encoding device 800 according to an embodiment can encode the maximum transform unit size information, the minimum transform unit size information, and the maximum transform unit partition information. The encoded maximum transform unit size information, the minimum transform unit size information, and the maximum transform unit partition information are inserted into the SPS. The video decoding device 900 according to an embodiment can use the maximum transform unit size information, the minimum transform unit size information, and the maximum transform unit partition information for video decoding.
[0279] For example, (a) if the current coding unit is 64x64 in size and the maximum transform unit size is 32x32, (a-1) when the transform unit split information is 0, the size of the transform unit is set to 32x32, (a-2) when the transform unit split information is 1, the size of the transform unit is set to 16x16, and (a-3) when the transform unit split information is 2, the size of the transform unit is set to 8x8.
[0280] As another example, (b) if the current coding unit is 32x32 in size and the minimum transform unit size is 32x32, (b-1) when the transform unit split information is 0, the size of the transform unit is set to 32x32, and since the size of the transform unit cannot be smaller than 32x32, no further transform unit split information is set.
[0281] As yet another example, (c) if the current coding unit is 64x64 in size and the maximum transform unit partition information is 1, the transform unit partition information is 0 or 1, and no other transform unit partition information is set.
[0282] Therefore, when the maximum transform unit division information is defined as "MaxTransformSizeIndex", the minimum transform unit size is defined as "MinTransformSize", and the transform unit size when the transform unit division information is 0 is defined as "RootTuSize", the minimum transform unit size possible for the current coding unit, "CurrMinTuSize", is defined as shown in the following equation (I).
[0283] CurrMinTuSize =max(MinTransformSize, RootTuSize / (2^MaxTransformSizeIndex)) (I) 'RootTuSize', which is the transform unit size when the transform unit partition information is 0, can indicate the maximum transform unit size that can be adopted by the system, compared with 'CurrMinTuSize', the minimum transform unit size possible for the current coding unit. That is, according to formula (I), 'RootTuSize / (2^MaxTransformSizeIndex)' is the transform unit size obtained by dividing 'RootTuSize', which is the transform unit size when the transform unit partition information is 0, by the number of times corresponding to the maximum transform unit partition information, and 'MinTransformSize' is the minimum transform unit size, so the smaller value of these is also the minimum transform unit size 'CurrMinTuSize' possible for the current coding unit.
[0284] The maximum transform unit size "RootTuSize" according to one embodiment varies depending on the prediction mode.
[0285] For example, if the current prediction mode is an inter mode, 'RootTuSize' is determined by the following equation (II): In equation (II), 'MaxTransformSize' indicates the maximum transform unit size, and 'PUSize' indicates the current prediction unit size.
[0286] RootTuSize=min(MaxTransformSize, PUSize) (II) That is, if the current prediction mode is inter mode, "RootTuSize", which is the transform unit size when the transform unit split information is 0, is set to the smaller value of the maximum transform unit size and the current prediction unit size.
[0287] If the prediction mode of the current partition unit is the intra mode, 'RootTuSize' is determined by the following equation (III): 'PartitionSize' indicates the size of the current partition unit.
[0288] RootTuSize=min(MaxTransformSize, PartitionSize) (III) That is, if the current prediction mode is the intra mode, "RootTuSize", which is the transform unit size when the transform unit split information is 0, is set to the smaller value of the maximum transform unit size and the current partition unit size.
[0289] However, it should be noted that the current maximum transform unit size "RootTuSize" according to one embodiment, which varies depending on the partition-based prediction mode, is only one embodiment, and the factors determining the current maximum transform unit size are not limited to these.
[0290] According to the video encoding technique based on the tree-structured coding unit described with reference to Figures 8 to 20, spatial domain video data is encoded for each tree-structured coding unit, and according to the video decoding technique based on the tree-structured coding unit, the spatial domain video data is restored while decoding is performed for each maximum coding unit, and pictures and video sequences are restored. The restored video is played back by a playback device, stored on a recording medium, or transmitted over a network.
[0291] Meanwhile, the above-described embodiments of the present invention can be written as a computer-executable program and implemented in a general-purpose digital computer that runs the program using a computer-readable recording medium, including magnetic recording media (e.g., ROM (read-only memory), floppy disk, hard disk, etc.) and optically readable media (e.g., CD-ROM (compact disc read-only memory), DVD (digital versatile disc), etc.).
[0292] For ease of explanation, the video encoding method and / or video encoding method previously described with reference to Figures 1A to 20 will be referred to as the "video encoding method of the present invention," and the video decoding method and / or video decoding method previously described with reference to Figures 1A to 20 will be referred to as the "video decoding method of the present invention."
[0293] 1A to 20, video encoding device 800, or a video encoding device configured with video encoding unit 1100 will be referred to as the "video encoding device of the present invention." Also, video decoding device 900, or a video decoding device configured with video decoding unit 1200, will be referred to as the "video decoding device of the present invention."
[0294] An embodiment in which the computer-readable recording medium on which the program is stored is a disk 21000 according to one embodiment will be described in detail below.
[0295] 21 illustrates the physical structure of a disk 26000 storing a program according to one embodiment. The disk 26000 described as a recording medium may also be a hard drive, a CD-ROM disk, a Blu-ray (registered trademark) disk, or a DVD disk. The disk 26000 is composed of a number of concentric tracks Tr, and the tracks Tr are divided into a predetermined number of sectors Se along the circumferential direction. Programs for implementing the quantization parameter determination method, video encoding method, and video decoding method described above are allocated and stored in specific areas of the disk 26000 storing the program according to the embodiment.
[0296] A computer system implemented using a recording medium storing a program for implementing the above-described video encoding and decoding methods will now be described with reference to FIG.
[0297] 22 illustrates a disk drive 26800 for recording and reading a program using a disk 26000. The computer system 26700 can store a program for implementing at least one of the video encoding method and the video decoding method of the present invention on the disk 26000 using the disk drive 26800. In order to execute the program stored on the disk 26000 on the computer system 26700, the disk drive 26800 reads the program from the disk 26000 and transmits the program to the computer system 26700.
[0298] In addition to the disk 26000 illustrated in Figures 21 and 22, a program for implementing at least one of the video encoding method and video decoding method of the present invention is also stored on a memory card, a ROM cassette, or an SSD (solid state drive).
[0299] A system to which the video encoding method and video decoding method according to the above embodiment are applied will now be described.
[0300] 23 illustrates the overall structure of a content supply system 11000 for providing a content distribution service. The service area of the communication system is divided into cells of a predetermined size, and radio base stations 11700, 11800, 11900, and 12000, which serve as base stations, are installed in each cell.
[0301] The content delivery system 11000 includes a number of independent devices, such as a computer 12100, a personal digital assistant (PDA) 12200, a camera 12600, and a mobile phone 12500, which are connected to the Internet 11100 via an Internet service provider 11200, a communication network 11400, and wireless base stations 11700, 11800, 11900, and 12000.
[0302] However, the content supply system 11000 is not limited to the structure shown in Fig. 23, and devices may be selectively connected. Independent devices may also be directly connected to the communication network 11400 without going through the wireless base stations 11700, 11800, 11900, and 12000.
[0303] The video camera 12300 is an imaging device capable of capturing video images, such as a digital video camera. The mobile phone 12500 may employ at least one communication method among various protocols, such as a personal digital communications (PDC) method, a code division multiple access (CDMA) method, a wideband code division multiple access (W-CDMA) method, a global system for mobile communications (GSM) method, and a personal handyphone system (PHS) method.
[0304] The video camera 12300 is connected to the streaming server 11300 via a wireless base station 11900 and a communication network 11400. The streaming server 11300 can stream content transmitted by a user using the video camera 12300 in real-time broadcast. The content received from the video camera 12300 is encoded by the video camera 12300 or the streaming server 11300. The video data captured by the video camera 12300 is also transmitted to the streaming server 11300 via the computer 12100.
[0305] Video data captured by the camera 12600 is also transmitted to the streaming server 11300 via the computer 12100. The camera 12600 is an imaging device capable of capturing both still and video images, like a digital camera. The video data received from the camera 12600 is encoded by the camera 12600 or the computer 12100. Software for video encoding and video decoding is stored on a computer-readable recording medium such as a CD-ROM disk, floppy disk, hard disk drive, SSD, or memory card that can be accessed by the computer 12100.
[0306] Also, if video is taken by a camera mounted on the mobile phone 12500, the video data is received from the mobile phone 12500.
[0307] The video data is encoded by an LSI (large scale integrated circuit) system installed in the video camera 12300, mobile phone 12500 or camera 12600.
[0308] In a content supply system 11000 according to one embodiment, content recorded by a user using a video camera 12300, a camera 12600, a mobile phone 12500, or other imaging device, such as on-site recording content of a concert, is encoded and transmitted to a streaming server 11300. The streaming server 11300 can stream the content data to other clients that have requested the content data.
[0309] The client is a device capable of decoding encoded content data, such as a computer 12100, a PDA 12200, a video camera 12300, or a mobile phone 12500. Thus, the content delivery system 11000 allows the client to receive and play encoded content data. The content delivery system 11000 also allows the client to receive, decode, and play encoded content data in real time, enabling personal broadcasting.
[0310] The video encoding device and video decoding device of the present invention are applied to the encoding and decoding operations of the independent devices included in the content supply system 11000.
[0311] 24 and 25, an embodiment of the mobile phone 12500 in the content supply system 11000 will be described in detail.
[0312] 24 illustrates the external structure of a mobile phone 12500 to which the video encoding and decoding methods of the present invention are applied, according to one embodiment. The mobile phone 12500 is a smartphone whose functions are not limited and whose functions can be changed or expanded substantially through application programs.
[0313] The mobile phone 12500 includes a built-in antenna 12510 for exchanging RF signals with the wireless base station 12000, and a display screen 12520, such as an LCD (liquid crystal display) screen or an OLED (organic light emitting diodes) screen, for displaying images captured by a camera 12530 or images received and decoded by the antenna 12510. The smartphone 12510 includes an operation panel 12540 including control buttons and a touch panel. If the display screen 12520 is a touch screen, the operation panel 12540 further includes a touch-sensing panel of the display screen 12520. The smartphone 12510 includes a speaker 12580 or other form of audio output unit for outputting voice and sound, and a microphone 12550 or other form of audio input unit for inputting voice and sound. The smartphone 12510 further includes a camera 12530, such as a CCD camera, for capturing video and still images. Smartphone 12510 may also include storage medium 12570 for storing encoded or decoded data, such as video or still images captured by camera 12530, received by email, or otherwise acquired, and slot 12560 for inserting storage medium 12570 into mobile phone 12500. Storage medium 12570 may also be an SD card or other form of flash memory, such as an EEPROM (electrically erasable and programmable read only memory) housed in a plastic case.
[0314] 25 illustrates the internal structure of the mobile phone 12500. In order to coordinately control each part of the mobile phone 12500, which is composed of the display screen 12520 and the operation panel 12540, the power supply circuit 12700, the operation input control unit 12640, the video encoding unit 12720, the camera interface 12630, the LCD control unit 12620, the video decoding unit 12690, the multiplexer / demultiplexer (MUX / DEMUX) 12680, the recording / reading unit 12670, the modulation / demodulation unit 12660, and the sound processing unit 12650 are connected to the central control unit 12710 via a synchronization bus 12730.
[0315] When the user operates the power button to change the state from "power off" to "power on," the power supply circuit 12700 supplies power from the battery pack to each part of the mobile phone 12500, setting the mobile phone 12500 into operating mode.
[0316] The central control unit 12710 includes a CPU (central processing unit), a ROM, and a RAM (random access memory).
[0317] In the process in which the mobile phone 12500 transmits communication data to the outside, a digital signal is generated in the mobile phone 12500 under the control of the central control unit 12710. For example, a digital audio signal is generated in the audio processing unit 12650, a digital video signal is generated in the video encoding unit 12720, and message text data is generated via the operation panel 12540 and the operation input control unit 12640. When the digital signal is transmitted to the modulation / demodulation unit 12660 under the control of the central control unit 12710, the modulation / demodulation unit 12660 modulates the frequency band of the digital signal, and the communication circuit 12610 performs D / A conversion (digital-analog conversion) and frequency conversion on the band-modulated digital audio signal. The transmission signal output from the communication circuit 12610 is sent to the voice communication base station or radio base station 12000 via the antenna 12510.
[0318] For example, when the mobile phone 12500 is in a call mode, an acoustic signal acquired by the microphone 12550 is converted into a digital acoustic signal by the acoustic processing unit 12650 under the control of the central control unit 12710. The generated digital acoustic signal is converted into a transmission signal via the modulation / demodulation unit 12660 and the communication circuit 12610 and is sent out via the antenna 12510.
[0319] When a text message such as an e-mail is transmitted in the data communication mode, the text data of the message is input using the operation panel 12540, and the text data is transmitted to the central control unit 12610 via the operation input control unit 12640. Under the control of the central control unit 12610, the text data is converted into a transmission signal via the modulation / demodulation unit 12660 and the communication circuit 12610, and is sent to the wireless base station 12000 via the antenna 12510.
[0320] In order to transmit video data in the data communication mode, video data captured by the camera 12530 is provided to the video encoding unit 12720 via the camera interface 12630. The video data captured by the camera 12530 is immediately displayed on the display screen 12520 via the camera interface 12630 and the LCD control unit 12620.
[0321] The structure of the video encoding unit 12720 corresponds to the structure of the video encoding device of the present invention described above. The video encoding unit 12720 can encode video data provided from the camera 12530 according to the video encoding method of the present invention described above, convert the encoded video data into compression-encoded video data, and output the encoded video data to the multiplexing / demultiplexing unit 12680. During recording by the camera 12530, an audio signal acquired by the microphone 12550 of the mobile phone 12500 is also converted into digital audio data via the audio processing unit 12650, and the digital audio data is transmitted to the multiplexing / demultiplexing unit 12680.
[0322] The multiplexing / demultiplexing unit 12680 multiplexes the coded video data provided from the image coding unit 12720 together with the audio data provided from the audio processing unit 12650. The multiplexed data is converted into a transmission signal via the modulation / demodulation unit 12660 and the communication circuit 12610 and is sent out via the antenna 12510.
[0323] When the mobile phone 12500 receives communication data from the outside, the signal received via the antenna 12510 is converted into a digital signal through frequency recovery processing and analog-to-digital conversion (A / D) processing. The modulation / demodulation unit 12660 demodulates the frequency band of the digital signal. The band-demodulated digital signal is transmitted to the video decoding unit 12690, the audio processing unit 12650, or the LCD control unit 12620 depending on the type.
[0324] When the mobile phone 12500 is in call mode, it amplifies a signal received via the antenna 12510 and generates a digital audio signal through frequency conversion and A / D (analog-to-digital) conversion processing. The received digital audio signal is converted into an analog audio signal through the modulation / demodulation unit 12660 and audio processing unit 12650 under the control of the central control unit 12710, and the analog audio signal is output via the speaker 12580.
[0325] In data communication mode, when data of a video file accessed from an Internet website is received, the signal received from the radio base station 12000 via the antenna 12510 is processed by the modulation / demodulation unit 12660, which outputs multiplexed data, and the multiplexed data is transmitted to the multiplexing / demultiplexing unit 12680.
[0326] To decode the multiplexed data received via antenna 12510, multiplexer / demultiplexer 12680 demultiplexes the multiplexed data and separates the encoded video data stream from the encoded audio data stream. A synchronization bus 12730 provides the encoded video data stream to video decoder 12690 and the encoded audio data stream to audio processor 12650.
[0327] The structure of the video decoding unit 12690 corresponds to the structure of the video decoding device of the present invention described above. The video decoding unit 12690 decodes encoded video data using the video decoding method of the present invention described above, generates restored video data, and provides the restored video data to the display screen 12520 via the LCD control unit 12620.
[0328] Thereby, the video data of the video file accessed from the Internet website is displayed on the display screen 12520. At the same time, the sound processing unit 12650 can also convert the audio data into an analog sound signal and provide the analog sound signal to the speaker 12580. Thereby, the audio data included in the video file accessed from the Internet website is also played on the speaker 12580.
[0329] The mobile phone 12500, or other type of communication terminal, may be a transmitting / receiving terminal that includes both the video encoding device and the video decoding device of the present invention, a transmitting terminal that includes only the above-mentioned video encoding device of the present invention, or a receiving terminal that includes only the video decoding device of the present invention.
[0330] The communication system of the present invention is not limited to the structure described with reference to Fig. 25. For example, Fig. 26 illustrates a digital broadcasting system to which a communication system according to an embodiment is applied.
[0331] The digital broadcasting system according to the embodiment of FIG. 26 can receive digital broadcasts transmitted via a satellite network or a terrestrial network by using the video encoding device and video decoding device of the present invention.
[0332] Specifically, a broadcast station 12890 transmits a video data stream via radio waves to a communications or broadcast satellite 12900. The broadcast satellite 12900 transmits a broadcast signal that is received by a satellite receiver at the home via an antenna 12860. In each home, the encoded video stream is decoded and played by a TV receiver 12810, a set-top box 12870, or other device.
[0333] The video decoding device of the present invention is implemented in the playback device 12830, so that the playback device 12830 can read and decode the encoded video stream recorded on the recording medium 12820, such as a disk or memory card, and the restored video signal is then played back on, for example, a monitor 12840.
[0334] The video decoding device of the present invention is also installed in a set-top box 12870 connected to an antenna 12860 for satellite / terrestrial broadcasting or a cable antenna 12850 for cable TV reception. The output data of the set-top box 12870 is also reproduced on a TV monitor 12880.
[0335] As another example, instead of the set-top box 12870, the TV receiver 12810 itself may be equipped with the video decoding device of the present invention.
[0336] A vehicle 12920 equipped with an appropriate antenna 12910 can also receive signals transmitted from the satellite 12800 or the radio base station 11700. The decoded video is played on a display screen of a vehicle navigation system 12930 installed in the vehicle 12920.
[0337] The video signal is encoded by the video encoding device of the present invention and then recorded and stored on a recording medium. Specifically, the video signal is stored on a DVD disc 12960 by a DVD recorder, or on a hard disk by a hard disk recorder 12950. As another example, the video signal may be stored on an SD card 12970. If the hard disk recorder 12950 is equipped with the video decoding device of the present invention according to an embodiment, the video signal recorded on the DVD disc 12960, the SD card 12970, or another type of recording medium is played back on a monitor 12880.
[0338] The automobile navigation system 12930 may not include the camera 12530, the camera interface 12630, and the video encoder 12720 of Figure 25. For example, the computer 12100 and the TV receiver 12810 may also not include the camera 12530, the camera interface 12630, and the video encoder 12720 of Figure 25.
[0339] FIG. 27 illustrates a network structure of a cloud computing system utilizing a video encoding device and a video decoding device, according to one embodiment.
[0340] The cloud computing system of the present invention comprises a cloud computing server 14000, a user DB 14100, computing resources 14200 and user terminals.
[0341] The cloud computing system provides on-demand outsourcing services for computing resources via information and communication networks such as the Internet in response to requests from user terminals. In a cloud computing environment, service providers use virtualization technology to integrate computing resources from data centers in different physical locations to provide services needed by users. Service users do not have to install computing resources such as applications, storage, operating systems, and security on their own terminals, but can select and use services in a virtual space created through virtualization technology at the desired time and to the desired extent.
[0342] A user terminal of a specific service user connects to the cloud computing server 14100 via an information communication network including the Internet and a mobile communication network. The user terminal receives cloud computing services, particularly video playback services, from the cloud computing server 14100. The user terminal may be any electronic device that can connect to the Internet, such as a desktop PC 14300, a smart TV 14400, a smartphone 14500, a laptop computer 14600, a PMP (portable multimedia player) 14700, or a tablet PC 14800.
[0343] The cloud computing server 14100 can integrate multiple computing resources 14200 distributed across a cloud network and provide them to user terminals. The multiple computing resources 14200 may include various data services and data uploaded from user terminals. In this way, the cloud computing server 14100 integrates video databases distributed in various locations using virtualization technology and provides services requested by user terminals.
[0344] The user DB 14100 stores information about users who subscribe to the cloud computing service. Here, the user information may include login information and personal credit information such as address and name. The user information may also include an index of videos. Here, the index may include a list of videos that have completed playback, a list of videos currently being played, and the stop time of a video currently being played.
[0345] Information related to videos stored in the user DB 14100 is shared between user devices. Therefore, for example, when a playback request is received from the laptop computer 14600 and a certain video service is provided to the laptop computer 14600, the playback history of the certain video service is stored in the user DB 14100. When a playback request for the same video service is received from the smartphone 14500, the cloud computing server 14100 refers to the user DB 14100, searches for the certain video service, and plays it. When the smartphone 14500 receives a video data stream via the cloud computing server 14100, the operation of decoding the video data stream and playing the video is similar to the operation of the mobile phone 12500 described above with reference to FIG. 24.
[0346] The cloud computing server 14100 can also refer to the playback history of a predetermined video service stored in the user DB 14100. For example, the cloud computing server 14100 receives a playback request for a video stored in the user DB 14100 from a user terminal. If the video was previously being played, the cloud computing server 14100 selects whether to play the video from the beginning or from the point where it was previously stopped, and the streaming method varies depending on the selection made by the user terminal. For example, if the user terminal requests playback from the beginning, the cloud computing server 14100 streams the video to the user terminal from the first frame. On the other hand, if the terminal requests playback to continue from the point where it was previously stopped, the cloud computing server 14100 streams the video to the user terminal from the frame where it was stopped.
[0347] In this case, the user terminal may include the video decoding device of the present invention described with reference to Figures 1A to 20. As another example, the user terminal may include the video encoding device of the present invention described with reference to Figures 1A to 20. Furthermore, the user terminal may include both the video encoding device and the video decoding device of the present invention described with reference to Figures 1A to 20.
[0348] An embodiment in which the video encoding method and video decoding method, and the video encoding device and video decoding device described with reference to Figures 1A to 20 are utilized has been described with reference to Figures 21 to 27. However, an embodiment in which the video encoding method and video decoding method described with reference to Figures 1A to 20 are stored on a recording medium or the video encoding device and video decoding device are implemented in a device is not limited to the embodiment of Figures 21 to 27.
[0349] The methods, processes, apparatus, products, and / or systems according to the present invention are simple, cost-effective, uncomplicated, and highly versatile and accurate. Furthermore, by incorporating well-known components into the processes, apparatus, products, and systems according to the present invention, they can be readily utilized, while also enabling efficient and economical manufacture, application, and utilization. Another important aspect of the present invention is that it meets current trends calling for cost reduction, system simplification, and performance improvement. These beneficial aspects of the present invention will, at the very least, advance the state of the art.
[0350] While the present invention has been described in connection with certain preferred embodiments, other alternatives, variations, and modifications of the present invention will be apparent to those skilled in the art in light of the foregoing description. Therefore, the appended claims are intended to encompass all such alternatives, variations, and modifications. Accordingly, all content described in the specification and drawings should be interpreted in an illustrative and non-limiting sense. The methods, processes, apparatus, products, and / or systems according to the present invention are simple, cost-effective, uncomplicated, and highly versatile and precise. Furthermore, by incorporating well-known components into the processes, apparatus, products, and systems according to the present invention, they can be readily utilized, enabling efficient and economical manufacture, application, and utilization. Another important aspect of the present invention is that it meets current trends calling for cost reduction, system simplification, and performance improvement. These advantageous aspects of the present invention will ultimately advance the state of the art.
[0351] While the present invention has been described with reference to certain preferred embodiments, other alternatives, variations, and modifications of the present invention will be apparent to those skilled in the art in light of the foregoing description. Therefore, the claims are intended to encompass all such alternatives, variations, and modifications. Accordingly, all content described in the specification and drawings should be interpreted in an illustrative and non-limiting sense.
[0352] The following is a list of exemplary means according to the embodiment. (Appendix 1) In a motion vector encoding device, a prediction unit that obtains motion vector predictor candidates of a plurality of predetermined motion vector resolutions by using spatial candidate blocks and temporal candidate blocks of the current block, and determines a motion vector predictor of the current block, a motion vector of the current block, and a motion vector resolution of the current block by using the motion vector predictor candidates; an encoding unit that encodes information indicating a predicted motion vector of the current block, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block, The apparatus, wherein the predetermined multiple motion vector resolutions include pixel resolutions greater than one pixel resolution. (Appendix 2) The prediction unit searching for a reference block in pixel units of the first motion vector resolution using a first set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates; searching for a reference block in pixel units of the second motion vector resolution using a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates; the first motion vector resolution and the second motion vector resolution are different from each other; The device described in Supplementary Note 1, characterized in that the first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and the temporal candidate blocks. (Appendix 3) The prediction unit searching for a reference block in pixel units of the first motion vector resolution using a first set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates; searching for a reference block in pixel units of the second motion vector resolution using a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates; the first motion vector resolution and the second motion vector resolution are different from each other; The device according to Supplementary Note 1, wherein the first set of motion vector predictor candidates and the second set of motion vector predictor candidate include different numbers of motion vector predictor candidates. (Appendix 4) The encoding unit The device described in Supplementary Note 1, characterized in that if the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the residual motion vector is downscaled and encoded by the resolution of the motion vector of the current block. (Appendix 5) If the current block is a current coding unit constituting an image, the motion vector resolution is determined to be the same for each coding unit, and there is a prediction unit predicted as an AMVP (advanced motion vector prediction) mode in the current coding unit, The device described in Supplementary Note 1, characterized in that the encoding unit encodes information indicating the AMVP mode and the motion vector resolution of the predicted prediction unit once as information indicating the motion vector resolution of the current block. (Appendix 6) If the current block is a current coding unit constituting a video, and the motion vector resolution is determined to be the same for each prediction unit, and there is a prediction unit predicted as an AMVP mode in the current coding unit, The encoding unit encodes information indicating the motion vector resolution of the current block for each AMVP mode and predicted prediction unit present in the current block as information indicating the motion vector resolution of the current block. (Appendix 7) In a motion vector encoding device, a prediction unit that obtains motion vector predictor candidates of a plurality of predetermined motion vector resolutions by using spatial candidate blocks and temporal candidate blocks of the current block, and determines a motion vector predictor of the current block, a motion vector of the current block, and a motion vector resolution of the current block by using the motion vector predictor candidates; an encoding unit that encodes information indicating a predicted motion vector of the current block, a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block, The prediction unit searching for a reference block in pixel units of the first motion vector resolution using a first set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates; searching for a reference block in pixel units of the second motion vector resolution using a second set of motion vector predictor candidates including one or more motion vector predictor candidates selected from the motion vector predictor candidates; the first motion vector resolution and the second motion vector resolution are different from each other; The first set of candidate predictor motion vectors and the second set of candidate predictor motion vectors are obtained from different candidate blocks among the spatial candidate blocks and the temporal candidate blocks, or include different numbers of candidate predictor motion vectors. (Appendix 8) In a motion vector encoding device, generating a merge candidate list including at least one merge candidate related to a current block; determining and encoding a motion vector of the current block using a motion vector of one of the merge candidates included in the merge candidate list; The merge candidate list includes motion vectors obtained by downscaling motion vectors of candidates included in the merge candidate list by a predetermined multiple of motion vector resolution. (Appendix 9) The downscaling may be The device described in Supplementary Note 8, characterized in that instead of the pixel indicated by the motion vector of the minimum motion vector resolution, one of the pixels located around the pixel indicated by the motion vector of the minimum motion vector resolution is selected based on the resolution of the motion vector of the current block, and adjusted to indicate the selected pixel. (Appendix 10) In a motion vector decoding device, an acquisition unit that acquires candidate predicted motion vectors of a plurality of predetermined motion vector resolutions using spatial candidate blocks and temporal candidate blocks of a current block, acquires information indicating a predicted motion vector of the current block from the candidate predicted motion vectors, and acquires a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block; a decoding unit that reconstructs a motion vector of the current block based on the residual motion vector, information indicating a predicted motion vector of the current block, and motion vector resolution information of the current block, The apparatus, wherein the predetermined multiple motion vector resolutions include pixel resolutions greater than one pixel resolution. (Appendix 11) The motion vector predictor candidates of the predetermined plurality of motion vector resolutions are a first set of motion vector predictor candidates including one or more motion vector predictor candidates of a first motion vector resolution, and a second set of motion vector predictor candidates including one or more motion vector predictor candidates of a second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other; The device described in Supplementary Note 10, characterized in that the first set of predicted motion vector candidates and the second set of predicted motion vector candidates are obtained from different candidate blocks among the candidate blocks included in the spatial candidate blocks and temporal candidate blocks, or include different numbers of predicted motion vector candidates. (Appendix 12) The decoding unit The device described in Supplementary Note 10, characterized in that if the pixel unit of the resolution of the motion vector of the current block is larger than the pixel unit of the minimum motion vector resolution, the residual motion vector is restored by upscaling it by the minimum motion vector resolution. (Appendix 13) The current block is a current coding unit constituting an image, and the motion vector resolution is determined to be the same for each coding unit. If there is a prediction unit predicted as an AMVP (advanced motion vector prediction) mode in the current coding unit, The device described in Supplementary Note 10, characterized in that the acquisition unit acquires information indicating the AMVP mode and the motion vector resolution of the predicted prediction unit once from the bitstream as information indicating the motion vector resolution of the current block. (Appendix 14) In a motion vector decoding device, generating a merge candidate list including at least one merge candidate related to a current block; determining and decoding a motion vector of the current block using a motion vector of one of the merge candidates included in the merge candidate list; The merge candidate list includes motion vectors obtained by downscaling motion vectors of candidates included in the merge candidate list by a predetermined multiple of motion vector resolution. (Appendix 15) In a motion vector decoding device, an acquisition unit that acquires candidate predicted motion vectors of a plurality of predetermined motion vector resolutions using spatial candidate blocks and temporal candidate blocks of a current block, acquires information indicating a predicted motion vector of the current block from the candidate predicted motion vectors, and acquires a residual motion vector between the motion vector of the current block and the predicted motion vector of the current block, and information indicating a motion vector resolution of the current block; a decoding unit that reconstructs a motion vector of the current block based on the residual motion vector, information indicating a predicted motion vector of the current block, and motion vector resolution information of the current block, The motion vector predictor candidates of the predetermined plurality of motion vector resolutions are a first set of motion vector predictor candidates including one or more motion vector predictor candidates of a first motion vector resolution, and a second set of motion vector predictor candidates including one or more motion vector predictor candidates of a second motion vector resolution, the first motion vector resolution and the second motion vector resolution are different from each other; The first set of candidate predictor motion vectors and the second set of candidate predictor motion vectors are obtained from different candidate blocks among the spatial candidate blocks and the temporal candidate blocks, or include different numbers of candidate predictor motion vectors. [Prior art documents] [Patent documents]
[0353] [Patent Document 1] Japanese Patent Application Publication No. 7-240927 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-187025
Claims
1. In the method for decoding a motion vector, obtaining a predicted motion vector of a current block and a residual motion vector of the current block; obtaining a shift value for the residual motion vector based on a resolution corresponding to the current block among a plurality of resolutions; upscaling the residual motion vector by performing a left shift using the shift value; restoring a motion vector of the current block based on the upscaled residual motion vector and the predicted motion vector; the shift value is one of a plurality of integer values, When the information about the current block is first information, the number of the plurality of resolutions is a first number, and when the information about the current block is second information, the number of the plurality of resolutions is a second number; the first number of resolutions includes a resolution greater than one; the first number of resolutions and the second number of resolutions include at least one resolution that is different from each other; The current block is obtained by dividing a higher-order block.
2. In the motion vector encoding method, obtaining a shift value for a residual motion vector of the current block based on a resolution corresponding to the current block among a plurality of resolutions; obtaining a predicted motion vector of the current block; obtaining a residual motion vector of the current block based on the motion vector of the current block and a predicted motion vector of the current block; downscaling the residual motion vector of the current block by performing a right shift using the shift value, and encoding the downscaled residual motion vector to generate a bitstream; the shift value is one of a plurality of integer values, When the information about the current block is first information, the number of the plurality of resolutions is a first number, and when the information about the current block is second information, the number of the plurality of resolutions is a second number; the first number of resolutions includes a resolution greater than one; the first number of resolutions and the second number of resolutions include at least one resolution that is different from each other; The current block is obtained by dividing a higher-order block.
3. A method for transmitting a bitstream generated by the encoding method of claim 2.
Citation Information
Patent Citations
Image processing method, image processing unit and data storage medium
JP1999239352A
Encoding apparatus, decoding apparatus, image processing apparatus, and method and program for them
JP2003319400A
Moving picture signal coding method, decoding method, coding apparatus, and decoding apparatus
JP2006187025A
Motion vector decoding method and encoding method
JP7764653B2
Image processing device and method
WO2010101064A1