Method of decoding or encoding a video and method of transmitting video data

CN116320476BActive Publication Date: 2026-09-22KT CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310355381.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-12-22
Filing Date
2017-12-15
Publication Date
2026-09-22
Estimated Expiration
2037-12-15

AI Technical Summary

Technical Problem

因此,在通过使用介质例如常规的有线和无线宽带网络传输图像数据时,或者在通过使用常规的存储介质存储图像数据时,传输和存储的成本增加了

Benefits of technology

[0023]根据本发明,可以对编码/解码目标块执行高效的帧间预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116320476B_ABST
    Figure CN116320476B_ABST
Patent Text Reader

Abstract

A method of decoding or encoding a video and a method of transmitting video data, including: determining a current block motion vector precision and a motion vector difference value; scaling the motion vector difference value with the motion vector precision; generating a current block motion vector candidate list; obtaining a motion vector predictor based on the motion vector candidate list and first information notifying and specifying one of the motion vector candidates from a bitstream; obtaining a current block motion vector with the motion vector predictor and the scaled motion vector difference value; determining the current block motion vector precision when the motion vector precision is determined from a motion vector precision set including a plurality of motion vector precision candidates, based on index information specifying one of the plurality of motion vector precision candidates parsed from the bitstream; determining the current block motion vector precision when the current block motion vector precision is not determined using the motion vector precision set, without parsing the index information from the bitstream; and determining whether the motion vector precision set is used based on a flag notified for the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 201780079795.1, filed on June 21, 2019, entitled "Video Signal Processing Method and Apparatus". The international filing date of the parent application is December 15, 2017, with international application number PCT / KR2017 / 014869, and priority date is December 22, 2016. Technical Field

[0002] The present invention relates to methods and apparatus for processing video signals, and particularly to methods for decoding or encoding video and methods for transmitting video data. Background Technology

[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, has increased across various application areas. However, the data volume of higher resolution and quality image data increases compared to regular image data. Therefore, the costs of transmission and storage increase when transmitting image data using media such as conventional wired and wireless broadband networks, or when storing image data using conventional storage media. To address these issues arising from the increasing resolution and quality of image data, efficient image encoding / decoding techniques can be utilized.

[0004] Image compression techniques encompass various methods, including: inter-frame prediction techniques that predict pixel values ​​included in the current image based on previous or subsequent images; intra-frame prediction techniques that predict pixel values ​​included in the current image using pixel information from the current image; and entropy coding techniques that assign short codes to frequently occurring values ​​and long codes to less frequently occurring values. Image data can be effectively compressed using such image compression techniques, and image data can be transmitted or stored.

[0005] Simultaneously, with the increasing demand for high-resolution images, the demand for stereoscopic image content as a new image service is also increasing. Video compression techniques for effectively delivering stereoscopic image content with high and ultra-high resolution are being discussed. Summary of the Invention

[0006] Technical issues

[0007] The purpose of this invention is to provide a method and apparatus for efficiently performing inter-frame prediction on target blocks during the encoding / decoding of video signals.

[0008] The object of this invention is to provide a method and apparatus for variably determining the accuracy of motion vectors when encoding / decoding video signals.

[0009] The object of the present invention is to provide a method and apparatus for compensating for differences in motion vector accuracy between blocks by comparing the motion vector accuracy of the blocks.

[0010] The technical objectives of this invention are not limited to the aforementioned technical problems. Furthermore, those skilled in the art will clearly understand from the following description other technical problems not mentioned.

[0011] Technical solution

[0012] The method and apparatus for decoding video signals according to the present invention can determine the motion vector precision of the current block, generate a motion vector candidate list for the current block, obtain a motion vector prediction value for the current block based on the motion vector candidate list, determine whether the precision of the motion vector prediction value is the same as the motion vector precision of the current block, scale the motion vector prediction value based on the motion vector precision of the current block if the precision of the motion vector prediction value is different from the motion vector precision of the current block, and obtain the motion vector of the current block using the scaled motion vector prediction value.

[0013] The method and apparatus for encoding video signals according to the present invention can determine the motion vector precision of the current block, generate a motion vector candidate list for the current block, obtain a motion vector prediction value for the current block based on the motion vector candidate list, determine whether the precision of the motion vector prediction value is the same as the motion vector precision of the current block, scale the motion vector prediction value based on the motion vector precision of the current block if the precision of the motion vector prediction value is different from the motion vector precision of the current block, and obtain the motion vector of the current block using the scaled motion vector prediction value.

[0014] In the method and apparatus for encoding / decoding video signals according to the present invention, the motion vector precision of the current block is determined based on a motion vector precision set including a plurality of motion vector precision candidates.

[0015] In the method and apparatus for encoding / decoding video signals according to the present invention, the motion vector precision of the current block is determined based on index information of one of a plurality of motion vector precision candidates.

[0016] In the method and apparatus for encoding / decoding video signals according to the present invention, the motion vector difference includes a prefix portion representing the integer part and a suffix portion representing the fractional part.

[0017] In the method and apparatus for encoding / decoding video signals according to the present invention, scaling is performed by shifting operations based on the scaling ratio between the accuracy of the motion vector prediction value and the motion vector accuracy of the current block.

[0018] According to one aspect of the present invention, a method for decoding video is provided, the method comprising: determining motion vector precision of a current block; obtaining motion vector difference of the current block; scaling the motion vector difference based on the motion vector precision of the current block; generating a motion vector candidate list for the current block, the motion vector candidate list including motion vector candidates, the motion vector candidates including spatial motion vector candidates and temporal motion vector candidates; obtaining a motion vector prediction value for the current block based on the motion vector candidate list and first information, wherein the first information is signaled from a bitstream and specifies one of the motion vector candidates in the motion vector candidate list; and using the motion vector prediction... The motion vector of the current block is obtained by using a value and a scaled difference in motion vectors, wherein when it is determined that the motion vector precision of the current block is determined from a set of motion vector precisions including multiple motion vector precision candidates, the motion vector precision of the current block is determined based on index information parsed from the bitstream, the index information specifying one of the multiple motion vector precision candidates, wherein when it is determined that the motion vector precision of the current block is determined without using the set of motion vector precisions, the motion vector precision of the current block is determined without parsing the index information from the bitstream, and wherein whether the set of motion vector precisions is used is determined based on a flag signaled for the current block.

[0019] According to one aspect of the present invention, a method for encoding video is provided, the method comprising: determining motion vector precision of a current block; obtaining motion vectors of the current block; generating a motion vector candidate list for the current block, the motion vector candidate list including motion vector candidates, the motion vector candidates including spatial motion vector candidates and temporal motion vector candidates; determining a motion vector prediction value of the current block based on one of the motion vector candidates in the motion vector candidate list; determining the motion vector prediction value of the current block based on the motion vector candidate list; deriving a motion vector difference based on the motion vectors and the motion vector prediction value; and encoding the scaled motion vector difference, wherein... The scaled motion vector difference described herein is obtained by scaling the motion vector difference based on the motion vector precision of the current block; wherein, first information specifying one of the motion vector candidates is encoded in the bitstream; wherein, when the motion vector precision of the current block is determined from a set of motion vector precisions including multiple motion vector precision candidates, index information specifying one of the multiple motion vector precision candidates is encoded into the bitstream; wherein, when the motion vector precision of the current block is determined without using the set of motion vector precisions, the encoding of the index information is skipped; and wherein, a flag indicating whether the set of motion vector precisions is used is encoded in the bitstream.

[0020] According to one aspect of the present invention, a method for transmitting video data is provided, comprising: determining the motion vector precision of a current block; obtaining the motion vector of the current block; generating a motion vector candidate list for the current block, the motion vector candidate list including motion vector candidates, the motion vector candidates including spatial motion vector candidates and temporal motion vector candidates; determining a motion vector prediction value for the current block based on one of the motion vector candidates in the motion vector candidate list; determining the motion vector prediction value for the current block based on the motion vector candidate list; deriving a motion vector difference based on the motion vector and the motion vector prediction value; and generating a bitstream including the video data by encoding the scaled motion vector difference, wherein the scaled motion vector difference... The vector interpolation is obtained by scaling the motion vector interpolation based on the motion vector precision of the current block; and a bitstream including the video data is transmitted, wherein first information specifying one of the motion vector candidates is encoded in the bitstream, wherein when it is determined that the motion vector precision of the current block is determined from a set of motion vector precisions including multiple motion vector precision candidates, index information specifying one of the multiple motion vector precision candidates is encoded into the bitstream, wherein when the motion vector precision of the current block is determined without using the set of motion vector precisions, the encoding of the index information is skipped, and wherein a flag indicating whether the set of motion vector precisions is used is encoded in the bitstream.

[0021] The features briefly outlined above are merely illustrative aspects of the invention as described in the detailed description below, and do not limit the scope of the invention.

[0022] Beneficial effects

[0023] According to the present invention, efficient inter-frame prediction can be performed on the encoded / decoded target block.

[0024] According to the present invention, the accuracy of motion vectors can be variably determined.

[0025] According to the present invention, motion vectors can be derived by compensating for the differences in motion vector resolutions between the blocks.

[0026] The effects achievable by this invention are not limited to those described above, and those skilled in the art will clearly understand other effects not mentioned below based on the following description. Attached Figure Description

[0027] Figure 1 This is a block diagram illustrating an apparatus for encoding video according to an embodiment of the present invention.

[0028] Figure 2This is a block diagram illustrating an apparatus for decoding video according to an embodiment of the present invention.

[0029] Figure 3 This is a diagram illustrating an example of hierarchical segmentation of coded blocks based on a tree structure according to an embodiment of the present invention.

[0030] Figure 4 This is a diagram illustrating a type of segmentation that allows binary tree-based segmentation according to an embodiment of the present invention.

[0031] Figure 5 This is a diagram illustrating an example of binary tree-based partitioning that allows only predetermined types of partitioning according to an embodiment of the present invention.

[0032] Figure 6 This is a diagram illustrating an example of encoding / decoding information related to the permissible number of binary tree splits according to an embodiment of the present invention.

[0033] Figure 7 This is a diagram illustrating a segmentation pattern applicable to coded blocks according to an embodiment of the present invention.

[0034] Figure 8 This is a flowchart illustrating an inter-frame prediction method according to an embodiment of the present invention.

[0035] Figure 9 This is a diagram illustrating the process of deriving motion information for the current block when a merge mode is applied to it.

[0036] Figure 10 This illustrates the process of exporting motion information for the current block when the AMVP mode is applied to it.

[0037] Figure 11 and Figure 12 The method for deriving motion vectors based on the accuracy of the current block's motion vectors is shown. Detailed Implementation

[0038] Various modifications can be made to this invention, and various embodiments of the invention exist. Examples of various embodiments will now be provided with reference to the accompanying drawings, and examples of various embodiments will be described in detail. However, the invention is not limited thereto, and the exemplary embodiments can be interpreted as including all modifications, equivalents, or alternatives within the technical concept and scope of the invention. In the described drawings, similar reference numerals refer to similar elements.

[0039] The terms "first," "second," etc., used in this specification may be used to describe various components, but these components are not to be construed as limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the invention, a "first" component may be referred to as a "second" component, and a "second" component may similarly be referred to as a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple terms.

[0040] It should be understood that in this specification, when an element is simply referred to as "connected to" or "coupled to" another element rather than "directly connected to" or "directly coupled to" another element, the element may be "directly connected to" or "directly coupled to" another element, or the element may be connected to or coupled to another element and there are other elements in between. Conversely, it should be understood that when an element is referred to as "directly coupled to" or "directly connected to" another element, there are no intermediate elements.

[0041] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Expressions used in the singular include expressions in the plural unless they have a distinct meaning in the context. It should be understood in this specification that terms such as “comprising,” “having,” etc., are intended to indicate the presence of features, numbers, steps, actions, elements, portions, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, portions, or combinations thereof may be present or added.

[0042] Preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. In the following drawings, the same constituent elements are indicated by the same reference numerals, and repeated descriptions of the same elements will be omitted.

[0043] Figure 1 This is a block diagram illustrating an apparatus for encoding video according to an embodiment of the present invention.

[0044] Reference Figure 1 The device 100 for encoding video may include: an image segmentation module 110, prediction modules 120 and 125, a transformation module 130, a quantization module 135, a rearrangement module 160, an entropy coding module 165, an inverse quantization module 140, an inverse transformation module 145, a filter module 150, and a memory 155.

[0045] Figure 1The constituent parts shown are illustrated independently to represent different functional characteristics within an apparatus for encoding video. Therefore, this does not imply that each constituent part is composed of a separate hardware or software unit. In other words, for convenience, each constituent part includes each of the listed constituent parts. Thus, at least two constituent parts of each constituent part can be combined to form a single constituent part, or a constituent part can be divided into multiple constituent parts to perform each function. Embodiments combining each constituent part and embodiments dividing a constituent part are also included within the scope of this invention without departing from its spirit.

[0046] Furthermore, some of the constituent parts may not be essential components for performing the basic functions of the invention, but rather optional components used only to improve the performance of the invention. The invention can be implemented by excluding components used to improve performance and including only those essential for achieving the essence of the invention. Structures that exclude optional components used only to improve performance and include only essential components are also included within the scope of the invention.

[0047] Image segmentation module 110 can segment an input image into one or more processing units. Here, the processing unit can be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). Image segmentation module 110 can segment an image into a combination of multiple coding units, prediction units, and transform units, and can encode the image by selecting a combination of coding units, prediction units, and transform units using a predetermined criterion (e.g., a cost function).

[0048] For example, an image can be segmented into multiple coding units. A recursive tree structure, such as a quadtree, can be used to segment the image into coding units. Coding units that are further segmented with an image or the largest coding unit as the root can be segmented in such a way that the number of child nodes corresponds to the number of coding units they were segmented into. Coding units that cannot be further segmented by a predetermined constraint are used as leaf nodes. That is, when it is assumed that only square segmentation is feasible for a coding unit, a coding unit can be segmented into at most four other coding units.

[0049] In the following, in embodiments of the present invention, a coding unit may refer to a unit that performs encoding or a unit that performs decoding.

[0050] A prediction unit can be divided into one of the partitions that are square or rectangular in shape and have the same size in a single coding unit, or a prediction unit can be divided into one of the partitions that have different shapes / sizes in a single coding unit.

[0051] When a prediction unit undergoing intra-frame prediction is generated based on a coding unit and the coding unit is not the smallest coding unit, intra-frame prediction can be performed without dividing the coding unit into multiple prediction units N×N.

[0052] Prediction modules 120 and 125 may include an inter-frame prediction module 120 that performs inter-frame prediction and an intra-frame prediction module 125 that performs intra-frame prediction. It can be determined whether inter-frame or intra-frame prediction is performed for a prediction unit, and detailed information based on each prediction method (e.g., intra-frame prediction mode, motion vectors, reference image, etc.) can be determined. Here, the processing unit undergoing prediction may be different from the processing unit that determines the prediction method and details for it. For example, the prediction method, prediction mode, etc., may be determined by the prediction unit, and the prediction may be performed by the transform unit. The residual value (residual block) between the generated prediction block and the original block can be input to the transform module 130. Furthermore, prediction mode information, motion vector information, etc., used for prediction can be encoded together with the residual value by the entropy coding module 165 and can be transmitted to the device for decoding the video. When using a specific coding mode, the original block can be encoded as is without generating a prediction block through prediction modules 120 and 125 and transmitted to the device for decoding the video.

[0053] The inter-frame prediction module 120 can predict prediction units based on information from at least one of the previous or subsequent images of the current image, or in some cases, it can predict prediction units based on information from some coded regions in the current image. The inter-frame prediction module 120 may include a reference image interpolation module, a motion prediction module, and a motion compensation module.

[0054] The reference image interpolation module can receive reference image information from the memory 155 and generate pixel information of integer pixels or smaller based on the reference image. In the case of luminance pixels, an 8-tap DCT-based interpolation filter with different filter coefficients can be used to generate pixel information of integer pixels or smaller in units of 1 / 4 pixels. In the case of chrominance signals, a 4-tap DCT-based interpolation filter with different filter coefficients can be used to generate pixel information of integer pixels or smaller in units of 1 / 8 pixels.

[0055] The motion prediction module can perform motion prediction based on a reference image interpolated by the reference image interpolation module. Various methods can be used to calculate motion vectors, such as Full Search-Based Block Matching (FBMA), Three-Step Search (TSS), and New Three-Step Search (NTS). Based on the interpolated pixels, the motion vector can have motion vector values ​​in units of 1 / 2 pixel or 1 / 4 pixel. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods can be used as motion prediction methods, such as skipping methods, merging methods, AMVP (Advanced Motion Vector Prediction) methods, and intra-block copying methods.

[0056] The intra-frame prediction module 125 can generate prediction units based on reference pixel information adjacent to the current block, which serves as pixel information in the current image. When the neighboring block of the current prediction unit is a block undergoing inter-frame prediction and therefore the reference pixel is a pixel undergoing inter-frame prediction, the reference pixel information of the neighboring block undergoing intra-frame prediction can be used to replace the reference pixel included in the block undergoing inter-frame prediction. That is, when a reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information.

[0057] Intra-frame prediction can include directional prediction modes that use reference pixel information depending on the prediction direction and non-directional prediction modes that do not use directional information when performing prediction. The mode used to predict luminance information can be different from the mode used to predict chrominance information, and to predict chrominance information, either the intra-frame prediction mode information used to predict luminance information or the predicted luminance signal information can be used.

[0058] When performing intra-prediction, if the size of the prediction unit is the same as the size of the transform unit, intra-prediction can be performed based on the pixels located to the left, upper left, and top of the prediction unit. However, when performing intra-prediction, if the size of the prediction unit is different from the size of the transform unit, intra-prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-prediction using N×N segmentation can be used only for the smallest coding unit.

[0059] In intra-frame prediction methods, prediction blocks can be generated after applying an AIS (Adaptive Intra-Frame Smoothing) filter to a reference pixel, depending on the prediction mode. The type of AIS filter applied to the reference pixel can vary. To perform intra-frame prediction, the intra-frame prediction mode of the current prediction unit can be predicted based on the intra-frame prediction modes of prediction units adjacent to it. When predicting the prediction mode of the current prediction unit using mode information predicted by neighboring prediction units, if the intra-frame prediction mode of the current prediction unit is the same as that of neighboring prediction units, predetermined flag information can be used to transmit information indicating that the prediction modes of the current prediction unit and those of neighboring prediction units are identical. If the prediction mode of the current prediction unit differs from that of neighboring prediction units, entropy coding can be performed to encode the prediction mode information of the current block.

[0060] Furthermore, residual blocks containing information about residual values, which are the differences between the predicted units and the original blocks of the predicted units, can be generated based on the predicted units generated by the prediction modules 120 and 125. The generated residual blocks can then be input into the transformation module 130.

[0061] Transform module 130 can transform the residual block using transformation methods such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and KLT. The residual block includes information about the residual values ​​between the original block and the prediction units generated by prediction modules 120 and 125. The choice between applying DCT, DST, or KLT to transform the residual block can be determined based on the intra-frame prediction mode information of the prediction units used to generate the residual block.

[0062] Quantization module 135 can quantize values ​​transformed to the frequency domain by transformation module 130. The quantization coefficients can vary depending on the importance of the block or image. The values ​​calculated by quantization module 135 can be provided to inverse quantization module 140 and rearrangement module 160.

[0063] The rearrangement module 160 can rearrange the coefficients of the quantized residual values.

[0064] The rearrangement module 160 can transform coefficients in two-dimensional block form into coefficients in one-dimensional vector form using a coefficient scanning method. For example, the rearrangement module 160 can use a zigzag scanning method to scan from DC coefficients to coefficients in the high-frequency domain to transform the coefficients into one-dimensional vector form. Depending on the size of the transform unit and the intra-frame prediction mode, a vertical scan (scanning coefficients in two-dimensional block form along the column direction) or a horizontal scan (scanning coefficients in two-dimensional block form along the row direction) can be used instead of a zigzag scan. That is, the choice between a zigzag scan, a vertical scan, and a horizontal scan can be determined based on the size of the transform unit and the intra-frame prediction mode.

[0065] Entropy coding module 165 can perform entropy coding based on the value calculated by rearrangement module 160. Entropy coding can use various coding methods, such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).

[0066] The entropy coding module 165 can encode various information from the rearrangement module 160 and the prediction modules 120 and 125. These various information include residual coefficient information and block type information of coding units, prediction mode information, segmentation unit information, prediction unit information, transform unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc.

[0067] The entropy coding module 165 can entropy code the coefficients of the coding units input from the rearrangement module 160.

[0068] The inverse quantization module 140 can perform inverse quantization on the value quantized by the quantization module 135, and the inverse transform module 145 can perform inverse transform on the value transformed by the transform module 130. The residual value generated by the inverse quantization module 140 and the inverse transform module 145 can be combined with the prediction units predicted by the motion estimation module, motion compensation module and intra-frame prediction module of the prediction modules 120 and 125 to generate a reconstruction block.

[0069] The filter module 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).

[0070] Deblocking filters remove block distortion caused by boundaries between blocks in a reconstructed image. To determine whether to perform deblocking, the pixels included in several rows or columns of a block can be the basis for deciding whether to apply a deblocking filter to the current block. When a deblocking filter is applied to a block, a strong or weak filter can be applied depending on the desired deblocking filtering intensity. Furthermore, horizontal and vertical filtering can be processed in parallel when applying a deblocking filter.

[0071] The offset correction module can correct the offset from the original image on a pixel-by-pixel basis in the deblocked image. To perform offset correction on a specific image, a method that considers the edge information of each pixel to apply the offset can be used, or the following method can be used: dividing the image's pixels into a predetermined number of regions, determining the regions to be offset, and applying the offset to the determined regions.

[0072] Adaptive Loop Filtering (ALF) can be performed based on values ​​obtained by comparing the filtered reconstructed image with the original image. Pixels included in the image can be divided into predetermined groups, the filter to be applied to each group can be determined, and filtering can be performed individually for each group. Information about whether ALF is applied and the luminance signal can be transmitted via the coding unit (CU). The shape and filter coefficients of the filter used for ALF can vary depending on each block. Furthermore, a filter of the same shape (fixed shape) for ALF can be applied regardless of the characteristics of the target block.

[0073] The memory 155 can store the reconstructed blocks or reconstructed images calculated by the filter module 150. The stored reconstructed blocks or reconstructed images can be provided to the prediction modules 120 and 125 during inter-frame prediction.

[0074] Figure 2 This is a block diagram illustrating an apparatus for decoding video according to an embodiment of the present invention.

[0075] Reference Figure 2 The device 200 for decoding video may include: an entropy decoding module 210, a rearrangement module 215, an inverse quantization module 220, an inverse transform module 225, a prediction module 230, 235, a filter module 240, and a memory 245.

[0076] When a video bitstream is input from a device used for encoding video, the input bitstream can be decoded by inverse processing of the device used for encoding video.

[0077] The entropy decoding module 210 can perform entropy decoding based on the inverse processing of entropy encoding performed by the entropy encoding module of the device for encoding video. For example, various methods can be applied corresponding to the method performed by the device for encoding video, such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).

[0078] The entropy decoding module 210 can decode information about intra-frame prediction and inter-frame prediction performed by the means for encoding video.

[0079] The rearrangement module 215 can rearrange the bitstream entropy decoded by the entropy decoding module 210 based on the rearrangement method used in the apparatus for encoding video. The rearrangement module can reconstruct and rearrange coefficients in one-dimensional vector form into coefficients in two-dimensional block form. The rearrangement module 215 can receive information related to the coefficient scan performed in the apparatus for encoding video, and can perform the rearrangement via a method that inversely scans the coefficients based on the scan order performed in the apparatus for encoding video.

[0080] The inverse quantization module 220 can perform inverse quantization based on the quantization parameters received from the means for encoding the video and the coefficients of the rearranged block.

[0081] The inverse transform module 225 can perform inverse transforms, namely inverse DCT, inverse DST, and inverse KLT. These are the inverse processes of DCT, DST, and KLT, which are performed by the transform module on the quantization results of the means for encoding video. The inverse transform can be performed based on the transform units determined by the means for encoding video. The inverse transform module 225 of the means for decoding video can selectively execute transform schemes (e.g., DCT, DST, KLT) based on multiple pieces of information such as the prediction method, the size of the current block, and the prediction direction.

[0082] Prediction modules 230 and 235 can generate prediction blocks based on information about prediction block generation received from entropy decoding module 210 and previously decoded block or image information received from memory 245.

[0083] As described above, similar to the operation of an apparatus for encoding video, when performing intra-frame prediction, if the size of the prediction unit is the same as the size of the transform unit, intra-frame prediction can be performed on the prediction unit based on the pixels located to its left, upper left, and top. When performing intra-frame prediction, if the size of the prediction unit is different from the size of the transform unit, intra-frame prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-frame prediction using N×N segmentation can be used only for the smallest coding unit.

[0084] Prediction modules 230 and 235 may include a prediction unit determination module, an inter-frame prediction module, and an intra-frame prediction module. The prediction unit determination module can receive various information from the entropy decoding module 210, such as prediction unit information, prediction mode information of the intra-frame prediction method, motion prediction information about the inter-frame prediction method, etc., and can divide the current coding unit into prediction units and determine whether to perform inter-frame prediction or intra-frame prediction on the prediction unit. By using the information required for inter-frame prediction of the current prediction unit received from the means for encoding video, the inter-frame prediction module 230 can perform inter-frame prediction on the current prediction unit based on information from at least one of the previous or subsequent images that include the current image of the current prediction unit. Alternatively, inter-frame prediction can be performed based on information from some pre-reconstructed regions in the current image that includes the current prediction unit.

[0085] To perform inter-frame prediction, it is possible to determine for the coding unit which of the following modes—skip mode, merge mode, AMVP mode, and inter-block copy mode—will be used as the motion prediction method for the prediction unit included in the coding unit.

[0086] Intra-prediction module 235 can generate prediction blocks based on pixel information in the current image. When the prediction unit is a prediction unit undergoing intra-prediction, intra-prediction can be performed based on intra-prediction mode information of the prediction unit received from the means for encoding video. Intra-prediction module 235 may include an adaptive intra-smoothing (AIS) filter, a reference pixel interpolation module, and a DC filter. The AIS filter performs filtering on the reference pixels of the current block and can determine whether to apply the filter based on the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode of the prediction unit and AIS filter information received from the means for encoding video. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.

[0087] When the prediction mode of the prediction unit is a prediction mode that performs intra-frame prediction based on pixel values ​​obtained by interpolating reference pixels, the reference pixel interpolation module can interpolate the reference pixels to generate reference pixels that are integers or smaller than integers. When the prediction mode of the current prediction unit is a prediction mode that generates prediction blocks without interpolating reference pixels, interpolation of reference pixels is not required. When the prediction mode of the current block is DC mode, the DC filter can generate prediction blocks through filtering.

[0088] The reconstructed block or reconstructed image can be provided to the filter module 240. The filter module 240 may include a deblocking filter, an offset correction module, and an ALF.

[0089] Information regarding whether a deblocking filter should be applied to a corresponding block or image can be received from the means for encoding video, as well as information about which filter, strong or weak, should be applied when applying the deblocking filter. The deblocking filter of the means for decoding video can receive this information from the means for encoding video and can perform deblocking filtering on the corresponding block.

[0090] The offset correction module can perform offset correction on the reconstructed image based on the type and offset value information of the offset correction applied to the image during encoding.

[0091] AFL can be applied to the coding unit based on information received from the device used to encode the video, such as whether ALF is applied and ALF coefficient information. ALF information can be provided as included in a specific parameter set.

[0092] The memory 245 can store reconstructed images or reconstructed blocks for use as reference images or reference blocks, and can provide the reconstructed images to the output module.

[0093] As described above, in embodiments of the present invention, for ease of explanation, the term "encoding unit" is used to refer to a unit used for encoding; however, the term "encoding unit" can also be used to refer to a unit that performs both decoding and encoding.

[0094] Furthermore, the current block can represent the target block to be encoded / decoded. And, depending on the encoding / decoding steps, the current block can represent a coding tree block (or coding tree unit), a coding block (or coding unit), a transform block (or transform unit), a prediction block (or prediction unit), etc.

[0095] Images can be encoded / decoded by dividing them into basic blocks of square or non-square shapes. These basic blocks are called coding tree units (CMUs). A CMU can be defined as the largest allowed coding unit within a sequence or slice. Information about the shape of the CMU, whether it is square or non-square, or its size, can be signaled via sequence parameter sets, image parameter sets, or slice headers. CMUs can be further divided into smaller partitions. For example, if the partition depth generated by dividing a CMU is 1, then the partition depth generated by dividing a CMU into partitions of depth 1 can be defined as 2. That is, the partition generated by dividing a CMU into partitions of depth k can be defined as having depth k+1.

[0096] A partition of arbitrary size generated by dividing the coding tree into units can be defined as a coding unit. A coding unit can be recursively divided or subdivided into basic units for performing prediction, quantization, transform, or loop filtering, etc. For example, a partition of arbitrary size generated by dividing the coding tree into units can be defined as a coding unit, or it can be defined as a transform unit or prediction unit, which is a basic unit for performing prediction, quantization, transform, or loop filtering, etc.

[0097] The segmentation of a coding tree unit or coding unit can be performed based on at least one of vertical lines and horizontal lines. Furthermore, the number of vertical or horizontal lines used to segment the coding tree unit or coding unit can be at least one or more. For example, a coding tree unit or coding unit can be divided into two partitions using one vertical line or one horizontal line, or into three partitions using two vertical lines or two horizontal lines. Alternatively, a coding tree unit or coding unit can be divided into four partitions with a length and width of 1 / 2 using one vertical line and one horizontal line.

[0098] When dividing a coding tree unit or coding unit into multiple partitions using at least one vertical line or at least one horizontal line, the partitions may have a uniform size or different sizes. Alternatively, any one partition may have a different size than the other partitions.

[0099] In the embodiments described below, it is assumed that the coding tree unit or coding unit is divided into a quadtree structure or a binary tree structure. However, it is also possible to use a greater number of vertical lines or a greater number of horizontal lines to divide the coding tree unit or coding unit.

[0100] Figure 3 This is a diagram illustrating an example of hierarchical segmentation of coded blocks based on a tree structure according to an embodiment of the present invention.

[0101] The input video signal is decoded in predetermined block units. The default unit for decoding the input video signal is a coding block. A coding block can be a unit that performs intra / inter-frame prediction, transform, and quantization. Furthermore, a prediction mode (e.g., intra-frame prediction mode or inter-frame prediction mode) is determined on a per-coding-block basis, and prediction blocks included in a coding block can share the determined prediction mode. A coding block can be a square or non-square block of any size in the range of 8×8 to 64×64, or it can be a square or non-square block of 128×128, 256×256, or larger.

[0102] Specifically, the coded blocks can be hierarchically segmented based on at least one of quadtrees and binary trees. Here, quadtree-based segmentation can mean dividing a 2N×2N coded block into four N×N coded blocks, and binary tree-based segmentation can mean dividing a coded block into two coded blocks. Even when performing binary tree-based segmentation, square-shaped coded blocks can exist at a lower depth.

[0103] Binary tree-based segmentation can be performed symmetrically or asymmetrically. Furthermore, the encoded blocks from binary tree-based segmentation can be square or non-square, such as rectangular shapes. For example, segmentation types that allow binary tree-based segmentation can include at least one of the following: symmetrical types of 2N×N (horizontal non-square coding units) or N×2N (vertical non-square coding units), and asymmetrical types of nL×2N, nR×2N, 2N×nU, or 2N×nD.

[0104] Binary tree-based segmentation can be restricted to either symmetric or asymmetric segmentation. In this case, constructing a coding tree unit using square blocks corresponds to quadtree CU segmentation, and constructing a coding tree unit using symmetric non-square blocks corresponds to binary tree segmentation. Constructing coding tree units using both square blocks and symmetric non-square blocks corresponds to quadtree CU segmentation and binary tree CU segmentation.

[0105] Binary tree-based partitioning can be performed on coded blocks that no longer require quadtree-based partitioning. Quadtree-based partitioning can be omitted from coded blocks that already require binary tree-based partitioning.

[0106] Furthermore, the segmentation at a lower depth can be determined based on the segmentation type at a higher depth. For example, if binary tree-based segmentation is allowed at two or more depths, then only the same type of binary tree segmentation as at higher depths can be allowed at lower depths. For instance, if binary tree-based segmentation at a higher depth is performed using a 2N×N type, then binary tree-based segmentation at a lower depth is also performed using a 2N×N type. Alternatively, if binary tree-based segmentation at a higher depth is performed using an N×2N type, then binary tree-based segmentation at a lower depth is also performed using an N×2N type.

[0107] In contrast, it is also possible to allow only types that are different from the binary tree partitioning types at lower depths.

[0108] It is possible to restrict the use of only specific types of binary tree-based segmentations for sequences, slices, coding tree units, or coding units. For example, for coding tree units, only 2N×N or N×2N type binary tree-based segmentations may be allowed. Available segmentation types can be predefined in the encoder or decoder. Alternatively, information about available segmentation types or unavailable segmentation types can be encoded and then signaled via a bitstream.

[0109] Figure 5 This is a diagram illustrating an example of how only specific types of binary tree-based partitioning are allowed. Figure 5 (a) shows an example that only allows binary tree-based partitioning of type N×2N, and Figure 5 (b) shows an example that only allows binary tree-based segmentation of type 2N×N. To achieve adaptive segmentation based on quadtrees or binary trees, the following information can be used: information indicating quadtree-based segmentation, information about the size / depth of the coded block that allows quadtree-based segmentation, information indicating binary tree-based segmentation, information about the size / depth of the coded block that allows binary tree-based segmentation, information about the size / depth of the coded block that does not allow binary tree-based segmentation, information about whether binary tree-based segmentation is performed in the vertical or horizontal direction, etc.

[0110] Additionally, for a coding tree unit or a specific coding unit, information can be obtained regarding the number of allowed binary tree splits, the depth of allowed binary tree splits, or the number of allowed binary tree split depths. This information can be encoded at the coding tree unit or coding unit level and sent to the decoder via a bitstream.

[0111] For example, the syntax "max_binary_depth_idx_minus1" indicating the maximum depth allowed for binary tree splits can be encoded / decoded via a bitstream. In this case, max_binary_depth_idx_minus1+1 can indicate the maximum depth allowed for binary tree splits.

[0112] Reference Figure 6 The example shown is in Figure 6 In the code, binary tree splits have already been performed for coding units of depth 2 and depth 3. Therefore, the bitstream can be encoded / decoded using at least one of the following: information indicating the number of times a binary tree split has been performed in the coding unit (i.e., 2 times), information indicating the maximum depth at which a binary tree split is allowed in the coding unit (i.e., depth 3), or information indicating the number of depths at which a binary tree split has been performed in the coding unit (i.e., 2 (depth 2 and depth 3)).

[0113] As another example, for each sequence or slice, at least one of the following information can be obtained: the number of allowed binary tree splits, the depth of allowed binary tree splits, or the number of depths of allowed binary tree splits. For example, this information can be encoded in units of sequences, images, or slices and transmitted via a bitstream. Therefore, at least one of the following—the number of binary tree splits in the first slice, the maximum depth of allowed binary tree splits in the first slice, or the number of depths at which binary tree splits are performed in the first slice—can differ from that in the second slice. For example, in the first slice, binary tree splits can be allowed at only one depth, while in the second slice, binary tree splits can be allowed at two depths.

[0114] As another example, the allowed number of binary tree splits, the allowed depth of binary tree splits, or the allowed depth of binary tree splits can be set differently based on the temporal ID of the slice or image. Here, the temporal ID is used to identify each of multiple video layers that have scalability in at least one of view, space, time, or quality.

[0115] like Figure 3 As shown, a first coding block 300 with a partition depth (splitting depth) of k can be divided into multiple second coding blocks based on a quadtree. For example, second coding blocks 310 to 340 can be square blocks with half the width and half the height of the first coding block, and the partition depth of the second coding blocks can be increased to k+1.

[0116] A second coded block 310 with a partition depth of k+1 can be divided into multiple third coded blocks with a partition depth of k+2. The partitioning of the second coded block 310 can be performed by selectively using either a quadtree or a binary tree, depending on the partitioning method. Here, the partitioning method can be determined based on at least one of information indicating quadtree-based partitioning and information indicating binary tree-based partitioning.

[0117] When segmenting the second coding block 310 based on a quadtree, the second coding block 310 can be divided into four third coding blocks 310a, each having half the width and half the height of the second coding block, and the partition depth of the third coding blocks 310a can be increased to k+2. In contrast, when segmenting the second coding block 310 based on a binary tree, the second coding block 310 can be divided into two third coding blocks. Here, each of the two third coding blocks can be a non-square block having half the width and half the height of the second coding block, and the partition depth can be increased to k+2. The second coding block can be determined as a horizontal or vertical non-square block depending on the segmentation direction, and the segmentation direction can be determined based on information about whether the binary tree-based segmentation is performed along the vertical or horizontal direction.

[0118] Meanwhile, the second coding block 310 can be determined as a leaf coding block that is no longer segmented based on a quadtree or binary tree. In this case, the leaf coding block can be used as a prediction block or a transform block.

[0119] Similar to the segmentation of the second coding block 310, the third coding block 310a can be determined as a leaf coding block, or it can be further segmented based on a quadtree or a binary tree.

[0120] Simultaneously, the third coding block 310b, based on binary tree segmentation, can be further segmented into vertical coding blocks 310b-2 or horizontal coding blocks 310b-3, and the partition depth of the relevant coding blocks can be increased to k+3. Alternatively, the third coding block 310b can be determined as a leaf coding block 310b-1 that is no longer segmented based on a binary tree. In this case, coding block 310b-1 can be used as a prediction block or a transform block. However, the above segmentation process can be performed restrictively based on at least one of the following information: information about the size / depth of coding blocks that allow quadtree-based segmentation, information about the size / depth of coding blocks that allow binary tree-based segmentation, and information about the size / depth of coding blocks that do not allow binary tree-based segmentation.

[0121] The number of candidates representing the size of a coded block can be limited to a predetermined number, or the size of the coded block within a predetermined unit can have a fixed value. For example, the size of a coded block in a sequence or image can be limited to 256×256, 128×128, or 32×32. Information indicating the size of the coded blocks in a sequence or image can be signaled via a sequence header or image header.

[0122] As a result of the partitioning based on quadtrees and binary trees, the coding unit can be represented as a square or rectangle of arbitrary size.

[0123] The coded block is encoded using at least one of skip mode, intra-frame prediction, inter-frame prediction, or skip method. Once the coded block is determined, the predicted block can be determined by predictive segmentation of the coded block. Predictive segmentation of the coded block can be performed using a segmentation mode (Part_mode) that indicates the segmentation type of the coded block. The size or shape of the predicted block can be determined based on the segmentation mode of the coded block. For example, the size of the predicted block determined based on the segmentation mode can be equal to or smaller than the size of the coded block.

[0124] Figure 7 This is a diagram showing the segmentation modes that can be applied to a coding block when encoding the coding block via inter-frame prediction.

[0125] When encoding a coded block using inter-frame prediction, one of eight segmentation modes can be applied to the coded block, such as... Figure 4 The example shown.

[0126] When encoding a coded block using intra-frame prediction, the partitioning mode PART_2Nx2N or the partitioning mode PART_NxN can be applied to the coded block.

[0127] When the coded block has a minimum size, PART_NxN can be applied. Here, the minimum size of the coded block can be predefined in the encoder and decoder. Alternatively, information about the minimum size of the coded block can be signaled via the bitstream. For example, the minimum size of the coded block can be signaled via the slice header, thus defining the minimum size of the coded block via slices.

[0128] Typically, prediction blocks can have sizes ranging from 64×64 to 4×4. However, when encoding blocks via inter-frame prediction, the prediction blocks can be restricted to a size other than 4×4 to reduce memory bandwidth when performing motion compensation.

[0129] Figure 8 This is a flowchart illustrating an inter-frame prediction method according to an embodiment of the present invention.

[0130] Reference Figure 8 S810 Determine the motion information of the current block. The motion information of the current block may include at least one of the motion vectors related to the current block, the reference image index of the current block, or the inter-frame prediction direction of the current block.

[0131] The motion information of the current block can be obtained based on at least one of the information signaled by a bit stream or the motion information of the adjacent blocks adjacent to the current block.

[0132] Figure 9 This is a diagram illustrating the process of deriving motion information for the current block when a merge mode is applied to it.

[0133] If the merge mode is applied to the current block, spatial merge candidates S910 can be derived based on the spatially adjacent blocks of the current block. Spatially adjacent blocks can include at least one of the blocks that are adjacent to the top, left, or corner (e.g., at least one of the top-left, top-right, or bottom-left corners) of the current block.

[0134] The motion information of spatial merging candidates can be set to be the same as the motion information of spatially adjacent blocks.

[0135] Temporal merge candidates S920 can be derived based on temporally adjacent blocks of the current block. A temporally adjacent block can mean a co-located block included in a collocated picture. The collocated picture has a different picture order count (POC) than the current picture that includes the current block. The collocated picture can be determined as a picture with a predefined index from a list of reference pictures, or it can be determined by an index signaled from the bitstream. A temporally adjacent block can be determined as a block in the collocated picture that has the same position and size as the current block, or a block adjacent to a block that has the same position and size as the current block. For example, a block in the collocated picture that includes the center coordinates of a block with the same position and size as the current block, or a block adjacent to the lower right boundary of the current block, can be determined as a temporally adjacent block.

[0136] Motion information for temporal merging candidates can be determined based on motion information from temporally adjacent blocks. For example, the motion vector of a temporal merging candidate can be determined based on the motion vector of a temporally adjacent block. Furthermore, the inter-frame prediction direction of the temporal merging candidate can be set to the same as the inter-frame prediction direction of the temporally adjacent blocks. However, the reference image index of the temporal merging candidate can have a fixed value. For example, the reference image index of the temporal merging candidate can be set to "0".

[0137] Subsequently, a merge candidate list S930, including spatial merge candidates and temporal merge candidates, can be generated. If the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, a combined merge candidate that combines two or more merge candidates or a merge candidate with zero motion vector (0,0) can be included in the merge candidate list.

[0138] When generating a list of merge candidates, at least one of the merge candidates included in the list can be specified based on the merge candidate index, S940.

[0139] The motion information of the current block can be set to be the same as the motion information of the merge candidate specified by the merge candidate index (S950). For example, when a spatial merge candidate is selected by the merge candidate index, the motion information of the current block can be set to be the same as the motion information of the spatially adjacent block. Alternatively, when a temporal merge candidate is selected by the merge candidate index, the motion information of the current block can be set to be the same as the motion information of the temporally adjacent block.

[0140] Figure 10 This illustrates the process of exporting motion information for the current block when the AMVP mode is applied to it.

[0141] When AMVP mode is applied to the current block, at least one of the inter-frame prediction direction or reference image index of the current block can be decoded based on the bitstream (S1010). In other words, when AMVP mode is applied, at least one of the inter-frame prediction direction or reference image index of the current block can be determined based on the encoded information from the bitstream.

[0142] Spatial motion vector candidates S1020 can be determined based on the motion vectors of spatially adjacent blocks of the current block. The spatial motion vector candidates can include at least one of a first spatial motion vector candidate derived from the top adjacent block of the current block and a second spatial motion vector candidate derived from the left adjacent block of the current block. Here, the top adjacent block can include at least one of the blocks adjacent to the top or upper right corner of the current block, and the left adjacent block of the current block can include at least one of the blocks adjacent to the left or lower left corner of the current block. The block adjacent to the upper left corner of the current block can be considered either a top adjacent block or a left adjacent block.

[0143] When the reference image between the current block and its spatial neighbors is different, the spatial motion vector can be obtained by scaling the motion vector of the spatial neighbors.

[0144] Candidate temporal motion vectors (S1030) can be determined based on the motion vectors of temporally adjacent blocks. If the reference images between the current block and its temporally adjacent blocks are different, the temporal motion vectors can be obtained by scaling the motion vectors of the temporally adjacent blocks.

[0145] A list of motion vector candidates, S1040, including spatial motion vector candidates and temporal motion vector candidates, can be generated.

[0146] When generating a list of motion vector candidates, at least one of the motion vector candidates included in the list can be specified based on information from at least one of the candidates in the list (S1050).

[0147] The motion vector candidate specified by the information is set as the motion vector prediction value for the current block. Furthermore, the motion vector S1060 for the current block is obtained by adding the motion vector difference to the motion vector prediction value. At this point, the motion vector difference can be parsed from the bitstream.

[0148] When motion information of the current block is obtained, motion compensation for the current block can be performed based on the obtained motion information S820. More specifically, motion compensation for the current block can be performed based on the inter-frame prediction direction, reference image index, and motion vector of the current block.

[0149] As in the example above, motion compensation for the current block can be performed based on the current block's motion information. In this case, the motion vector can have precision (or resolution) in whole-pixel units or fractional-pixel units.

[0150] Integer pixel units can include N integer pixels, such as integer pixels, 2 integer pixels, 4 integer pixels, etc. Here, N can be represented by a natural number of 1 or greater, specifically by a power of 2. An integer pixel can represent one pixel precision (i.e., one pixel unit), 2 integer pixels can represent twice the precision of one pixel (i.e., two pixel units), and 4 integer pixels can represent four times the precision of one pixel (i.e., four pixel units). Depending on the chosen integer pixel, motion vectors can be represented in N-pixel units, and motion compensation can be performed in N-pixel units.

[0151] Fractional pixel units can include 1 / N pixels, such as half a pixel, a quarter pixel, an eighth pixel, etc. Here, N can be represented by a natural number of 1 or greater, specifically by a power of 2. A half pixel can represent 1 / 2 precision of a pixel (i.e., a half-pixel unit), a quarter pixel can represent 1 / 4 precision of a pixel (i.e., a quarter-pixel unit), and an eighth pixel can represent 1 / 8 precision of a pixel (i.e., an eighth-pixel unit). Depending on the chosen decimal pixel unit, motion vectors can be represented in units of 1 / N pixels, and motion compensation can be performed in units of 1 / N pixels.

[0152] Figure 11 and Figure 12 This is a diagram illustrating the method for deriving motion vectors based on the accuracy of the motion vectors of the current block. Figure 11 The method for deriving motion vectors in AMVP mode is shown. Figure 12 The method for deriving motion vectors in the merge mode is shown.

[0153] First, the motion vector precision S1110 and S1210 of the current block can be determined.

[0154] Motion vector precision can be determined in units of sequences, images, slices, or predetermined blocks. Here, a predetermined block can represent a CTU, CU, PU, ​​or a block of a predetermined size / shape. A CTU can represent a CU of the maximum allowed size in the encoder / decoder. If the motion vector precision is determined at a level higher than the block level, such as at a sequence, image, or slice, motion compensation for the predetermined block can be performed based on the motion vector precision determined at that higher level. For example, motion compensation for a block included in a first slice can be performed using a motion vector where the precision is in whole pixel units, while motion compensation for a block included in a second slice can be performed using a motion vector where the precision is in quarter pixel units.

[0155] To determine the accuracy of motion vectors, information used to determine the accuracy can be signaled via a bitstream. This information can be an index, "mv_resolution_idx," specifying at least one of several motion vector accuracies. For example, Table 1 shows the motion vector accuracy based on mv_resolution_idx.

[0156] [Table 1]

[0157] 0 Quarter pixel unit 1 half-pixel unit 2 Integer pixel unit 3 One-eighth pixel unit 4 2 integer pixels 5 4 integer pixels

[0158] The examples shown in Table 1 are merely examples to which the present invention can be applied. The type and / or number of motion vector precision candidates that can be applied to a predetermined cell may differ from those shown in Table 1. The value and / or range of mv_resolution_idx may also differ depending on the type and / or number of motion vector precision candidates.

[0159] In another example, motion vector precision can be derived from cells spatially or temporally adjacent to a predetermined cell. Here, the predetermined cell can represent an image, slice, or block, and adjacent cells can represent images, slices, or blocks spatially or temporally adjacent to the predetermined cell. For example, the motion vector precision of the current block can be set to be equal to the motion vector precision of the blocks specified by index information in the spatially and / or temporally adjacent blocks.

[0160] As another example, the motion vector precision of the current block can be adaptively determined based on the motion information of the current block. For instance, the motion vector precision of the current block can be adaptively determined based on whether the time sequence or image sequence count of the reference image of the current block precedes the current image, whether the time sequence or image sequence count of the reference image of the current block follows the current image, or whether the reference image of the current block is the current image.

[0161] Some of the multiple motion vector accuracy candidates can be selectively used. For example, after defining a motion vector accuracy set that includes at least one motion vector accuracy candidate, the motion vector accuracy can be determined by including at least one motion vector accuracy candidate in the motion vector accuracy set.

[0162] Motion vector precision sets can be determined in units of sequences, images, slices, or blocks. Motion vector precision candidates included in the motion vector precision set can be predefined in the encoder and decoder. Alternatively, the motion vector precision set can be determined based on encoded information signaled via a bitstream. Here, the encoded information can relate to at least one of the types and / or numbers of motion vector precision candidates included in the motion vector resolution set. As another example, the motion vector precision set can be derived from units spatially or temporally adjacent to a predetermined unit. Here, a predetermined unit can represent an image, slice, or block, and an adjacent unit can represent an image, slice, or block spatially or temporally adjacent to the predetermined unit. For example, the motion vector precision set of a predetermined slice can be set equal to the motion vector precision set of slices spatially adjacent to the slice. Alternatively, depending on the dependency between slices, the motion vector precision set of an independent slice can be set as the motion vector precision set of a subordinate slice.

[0163] If a motion vector precision set is determined, at least one motion vector precision candidate included in the motion vector precision set can be identified as a motion vector precision. This can be achieved by signaling index information specifying at least one of the motion vector precision candidates included in the motion vector precision set via a bitstream. For example, the motion vector precision of the current block can be set to a candidate specified by the index information that is included in the motion vector precision set.

[0164] Whether to use a motion vector precision set can be adaptively determined based on the slice type, the size / shape of the current block, or the motion information of the current block (e.g., a reference image of the current block or the predicted orientation of the current block). Alternatively, information indicating whether to use a motion vector precision set can be signaled via a bitstream (e.g., a flag).

[0165] If a set of motion vector precision is determined at a level higher than the block level, such as at a sequence, image, or slice, the motion vector precision of a predetermined block can be derived from the set of motion vector precision determined at the higher level. For example, if a set of motion vector precision including quarter pixels and 2 whole pixels is defined at the image level, the blocks included in the image can be restricted to using at least one of quarter pixels or 2 whole pixels.

[0166] When multi-directional prediction is applied to the current block, the multiple motion vectors predicted from the multi-directional prediction can have different motion vector accuracies from each other. That is, the accuracy of any one of the motion vectors in the current block can be different from the accuracy of another motion vector. For example, when bidirectional prediction is applied to the current block, the accuracy of the forward motion vector mvL0 can be different from the accuracy of the backward motion vector mvL1. Even when multi-directional prediction with more than three directions is applied to the current block, at least one of the multiple motion vectors can have a different accuracy than another. Therefore, the information used to determine the accuracy of the motion vectors can be encoded / decoded for each prediction direction of the current block.

[0167] If the AMVP mode is applied to the current block and the motion vector precision of each block is variably determined, the precision of the motion vector predictions (or motion vector predictions, MVPs) derived from adjacent blocks can differ from the motion vector precision of the current block. To adjust the precision of the motion vector predictions to match the motion vector precision of the current block, the motion vector predictions can be scaled according to the motion vector precision of the current block (S1120). The motion vector predictions can be scaled according to the motion vector precision of the current block. The motion vectors of the current block can be derived by adding the motion vector difference (MVD) to the scaled motion vector predictions (S1130).

[0168] For example, if the pixel units of the motion vectors of adjacent blocks are quarter pixels and the pixel units of the motion vector of the current block are whole pixels, then the motion vector predictions derived from the adjacent blocks can be scaled in whole pixels, and a motion vector with whole pixel precision can be derived by adding the scaled motion vector predictions to the motion vector difference. For example, Equation 1 below shows an example of obtaining a motion vector by scaling the motion vector predictions in whole pixels.

[0169] [Equation 1]

[0170] mvLX[0]=((mvpLX[0]>>2)+mvdLX[0])<<2

[0171] mvLX[1]=((mvpLX[1]>>2)+mvdLX[1])<<2

[0172] In Equation 1, mvpLX represents the predicted value of the motion vector, and mvdLX represents the difference in the motion vector. In addition, mvLX[0], mvpLX[0] and mvdLX[0] represent the vertical motion vector components, and mvLX[1], mvpLX[1] and mvdLX[1] represent the horizontal motion vector components.

[0173] As another example, when the pixel units of the motion vectors of adjacent blocks are 2 whole pixels and the pixel units of the motion vector of the current block are quarter pixels, the motion vector predictions derived from the adjacent blocks can be scaled in quarter-pixel units, and a motion vector with quarter-pixel accuracy can be derived by adding the scaled motion vector predictions to the motion vector difference. For example, Equation 2 below shows an example of obtaining motion vectors when the current image is used as a reference image.

[0174] [Equation 2]

[0175] mvLX[0]=(mvpLX[0]>>3+mvdLX[0])<<3

[0176] mvLX[1]=(mvpLX[1]>>3+mvdLX[1])<<3

[0177] In Equations 1 and 2, the shift value used to scale the predicted motion vector value can be adaptively determined based on the scaling ratio between the motion vector accuracy of the current block and the motion vector accuracy of the adjacent blocks.

[0178] Unlike Figure 11 The example shown can also scale the motion vector generated by adding the motion vector prediction value and the motion vector difference according to the motion vector accuracy of the current block.

[0179] The motion vector difference can be encoded / decoded based on the motion vector precision of the current block. For example, when the motion vector precision of the current block is one-quarter of a pixel, the motion vector difference of the current block can be encoded / decoded in units of one-quarter of a pixel.

[0180] Regardless of the motion vector precision of the current block, the motion vector difference can be encoded / decoded in predetermined units. Here, the predetermined unit can be a fixed pixel unit predefined in the encoder and decoder (e.g., an integer pixel or a quarter pixel), or it can be a pixel unit determined at a higher level, such as at the image or slice level. When the motion vector precision of the current block differs from the precision of the motion vector difference, the motion vector of the current block can be derived by scaling the motion vector difference or by scaling the motion vector predicted value derived by adding the scaled motion vector prediction value to the motion vector difference. For example, when the motion vector precision of the current block is an integer pixel and the motion vector difference is encoded with a quarter pixel precision, as shown in Equation 1, the motion vector of the current block can be obtained by scaling the motion vector derived by adding the scaled motion vector predicted value to the motion vector difference.

[0181] Depending on the precision of the motion vector, different encoding / decoding methods can be determined for the motion vector difference. For example, if the resolution is in decimal pixels, the motion vector difference can be encoded / decoded by dividing it into a prefix and a suffix. The prefix can represent the integer part of the motion vector, and the suffix can represent the fractional part. For example, Equation 3 below shows an example of deriving the prefix "predfix_mvd" and the suffix "suffix_mvd".

[0182] [Equation 3]

[0183] prefix_mvd=MVD / N

[0184] suffix_mvd = MVD%N

[0185] In Equation 3, N can be a fixed value, or it can be a value that is variably determined based on the motion vector precision of the current block. For example, N can be proportional to the motion vector precision of the current block.

[0186] If the motion vector precision of the current block is two or more integer pixels, the value obtained by shifting the motion vector difference by N can be encoded. For example, if the motion vector precision of the current block is 2 integer pixels, half of the motion vector difference can be encoded / decoded. If the motion vector precision of the current block is 4 integer pixels, one-quarter of the motion vector difference can be encoded / decoded. In this case, the motion vector of the current block can be derived by scaling the decoded motion vector difference according to the motion vector precision of the current block.

[0187] If a merge mode or skip mode is applied to the current block and the motion vector precision of each block is variably determined, the motion vector precision of the current block may differ from that of the spatial / temporal merge candidate blocks. Therefore, the motion vectors of spatially / temporally adjacent blocks are scaled according to the motion vector precision of the current block (S1220), and the scaled motion vectors can be set as the motion information of the spatial / temporal merge candidate (S1230). For example, the motion vectors mvLX[0] and / or mvLX[1] of spatially / temporally adjacent blocks are scaled according to the motion vector precision of the current block to derive the scaled motion vectors mxLXscale[0] and / or mvLXscale[1], and the scaled motion vectors can be set as the motion vectors of the spatial / temporal merge candidate.

[0188] For example, when the motion vector precision of the adjacent block is a quarter pixel and the motion vector precision of the current block is an integer pixel, the motion vector of the adjacent block can be scaled as shown in Equation 4, and the scaled motion vector can be set as the motion vector of the spatial merging candidate.

[0189] [Equation 4]

[0190] mvLXscale[0]=(mvLX[0]>>2)<<2

[0191] mvLXscale[1]=(mvLX[1]>>2)<<2

[0192] In Equation 4, the shift value used to scale the motion vectors of adjacent blocks can be adaptively determined based on the scaling ratio between the motion vector accuracy of the current block and the motion vector accuracy of the adjacent blocks.

[0193] As another example, after selecting a merge candidate to merge with the current block (i.e., a merge candidate selected by the merge index), it can be checked whether its motion vector precision corresponds to the motion vector precision of the current block. If the motion vector precision of the selected merge candidate differs from that of the current block, the motion vector of the selected merge candidate can be scaled according to the motion vector precision of the current block.

[0194] The motion vector of the current block can be set to be equal to the motion vector of the merge candidate selected by the index information in the merge candidate (i.e., the scaled motion vector) S1240.

[0195] and Figure 12 Unlike the example shown, the motion vector precision of spatial / temporal neighboring blocks can be considered to determine the merging candidate for the current block. For example, whether a spatial / temporal neighboring block is suitable as a merging candidate can be determined based on whether the difference or scaling ratio between the motion vector precision of the spatial neighboring block and the motion vector precision of the current block is equal to or greater than a predetermined threshold. For instance, if the motion vector precision of a spatial merging candidate is 2 whole pixels and the motion vector precision of the current block is 1 / 4 pixel, it may mean that the correlation between the two blocks is not important. Therefore, spatial / temporal neighboring blocks whose precision difference from the motion vector precision of the current block is greater than a threshold can be set as not suitable as merging candidates. That is, a spatial / temporal neighboring block can only be used as a merging candidate if the difference between the motion vector precision of the spatial / temporal neighboring block and the motion vector precision of the current block is less than a threshold. Spatial / temporal neighboring blocks that are not suitable as merging candidates may not be added to the merging candidate list.

[0196] When the difference or scaling ratio between the motion vector precision of the current block and the motion vector precision of the adjacent blocks is less than or equal to a threshold but the two precisions are different from each other, the scaled motion vector can be set as the motion vector to be merged into the candidate motion vector, or it can be done according to the above reference. Figure 12 The described implementation method is used to scale the motion vectors of the merge candidates specified by the merge index.

[0197] The motion vector of the current block can be derived from the motion vectors of the merge candidates added to the merge candidate list. If the precision of the current block's motion vector differs from the precision of the motion vectors of the merge candidates added to the merge candidate list,

[0198] The motion vector precision difference can represent the difference between motion vector precisions, or it can represent the difference between corresponding values ​​for each motion vector precision. Here, the corresponding value can indicate the index value corresponding to the motion vector precisions shown in Table 1, or it can represent the value assigned to each motion vector precision shown in Table 2. For example, in Table 2, the corresponding value assigned to a quarter pixel is 2, and the corresponding value assigned to a whole pixel is 3, so the difference between the two precisions can be determined to be 2.

[0199] [Table 2]

[0200]

[0201]

[0202] The availability of temporally / spatially adjacent blocks can also be determined using the scaling factor of the motion vector precision instead of the difference in motion vector precision. Here, the scaling factor of the motion vector precision can represent the ratio between two motion vector precisions. For example, the scaling factor between a quarter pixel and a whole pixel can be defined as 4.

[0203] Although the above embodiments have been described based on a series of steps or flowcharts, they do not limit the timing of the invention and can be executed simultaneously or in different orders as needed. Furthermore, each of the components (e.g., units, modules, etc.) constituting the block diagrams in the above embodiments can be implemented by hardware devices or software and multiple components. Alternatively, multiple components can be implemented by a single hardware device or software combination. The above embodiments can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. A computer-readable recording medium can include one or a combination of program commands, data files, data structures, etc. Examples of computer-readable media include magnetic media (e.g., hard disks, floppy disks, and magnetic tapes), optical recording media (e.g., CD-ROMs and DVDs), magneto-optical media (e.g., optical-magnetic floppy disks), and hardware devices and media specifically configured to store and execute program instructions (e.g., ROMs, RAMs, flash memory), etc. Hardware devices can be configured to operate as one or more software modules to perform the processing according to the invention, and vice versa.

[0204] Industrial application

[0205] This invention can be applied to electronic devices capable of encoding / decoding video.

[0206] The present invention can also be implemented through the following embodiments.

[0207] Implementation Scheme 1. A method for decoding video, the method comprising:

[0208] Determine the accuracy of the motion vector for the current block;

[0209] Generate a candidate list of motion vectors for the current block;

[0210] The predicted motion vector value of the current block is obtained based on the motion vector candidate list;

[0211] Determine whether the accuracy of the predicted motion vector value is the same as the accuracy of the motion vector of the current block;

[0212] If the accuracy of the motion vector prediction value differs from the accuracy of the motion vector of the current block, the motion vector prediction value is scaled according to the accuracy of the motion vector of the current block; and the scaled motion vector prediction value is used to obtain the motion vector of the current block.

[0213] Implementation Scheme 2. The method according to Implementation Scheme 1, wherein the motion vector precision of the current block is determined based on a motion vector precision set including multiple motion vector precision candidates.

[0214] Implementation Scheme 3. The method according to Implementation Scheme 2, wherein the motion vector accuracy of the current block is determined based on index information of one of the plurality of motion vector accuracy candidates.

[0215] Implementation Scheme 4. The method according to Implementation Scheme 1, wherein the motion vector difference includes a prefix portion representing the integer part and a suffix portion representing the fractional part.

[0216] Implementation Scheme 5. The method according to Implementation Scheme 1, wherein the scaling is performed by a shift operation based on the scaling ratio between the accuracy of the motion vector prediction value and the accuracy of the motion vector of the current block.

[0217] Implementation Scheme 6. A method for encoding video, the method comprising:

[0218] Determine the accuracy of the motion vector for the current block;

[0219] Generate a candidate list of motion vectors for the current block;

[0220] The predicted motion vector value of the current block is obtained based on the motion vector candidate list;

[0221] Determine whether the accuracy of the predicted motion vector value is the same as the accuracy of the motion vector of the current block;

[0222] If the accuracy of the motion vector prediction value differs from the accuracy of the motion vector of the current block, the motion vector prediction value is scaled according to the accuracy of the motion vector of the current block; and the scaled motion vector prediction value is used to obtain the motion vector of the current block.

[0223] Implementation Scheme 7. The method according to Implementation Scheme 6, wherein the motion vector accuracy of the current block is determined based on a motion vector accuracy set including multiple motion vector accuracy candidates.

[0224] Implementation Scheme 8. The method according to Implementation Scheme 7, wherein the motion vector accuracy of the current block is determined based on index information of one of the plurality of motion vector accuracy candidates.

[0225] Implementation Scheme 9. The method according to Implementation Scheme 6, wherein the motion vector difference includes a prefix portion representing the integer part and a suffix portion representing the fractional part.

[0226] Implementation Scheme 10. The method according to Implementation Scheme 6, wherein the scaling is performed by a shift operation based on the scaling ratio between the accuracy of the motion vector prediction value and the accuracy of the motion vector of the current block.

[0227] Implementation Scheme 11. An apparatus for decoding video, the apparatus comprising:

[0228] An inter-frame prediction unit is configured to: determine the motion vector precision of the current block; generate a motion vector candidate list for the current block; obtain a motion vector prediction value for the current block based on the motion vector candidate list; determine whether the precision of the motion vector prediction value is the same as the motion vector precision of the current block; scale the motion vector prediction value based on the motion vector precision of the current block if the precision of the motion vector prediction value is different from the motion vector precision of the current block; and use the scaled motion vector prediction value to obtain the motion vector of the current block.

Claims

1. A method for decoding video, the method comprising: Determine the accuracy of the motion vector for the current block; Obtain the motion vector difference of the current block; The motion vector difference is scaled based on the motion vector accuracy of the current block; Generate a motion vector candidate list for the current block, the motion vector candidate list including motion vector candidates, the motion vector candidates including spatial motion vector candidates and temporal motion vector candidates; The motion vector prediction value of the current block is obtained based on the motion vector candidate list and the first index information, wherein the first index information specifies one of the motion vector candidates in the motion vector candidate list; as well as The motion vector of the current block is obtained using the predicted motion vector value and the scaled motion vector difference. Specifically, the motion vector precision of the current block is determined based on the flags decoded from the bitstream, indicating whether it is determined from a set of motion vector precision candidates. Wherein, if the flag indicates that the motion vector precision is determined from the set of motion vector precisions, the motion vector precision of the current block is determined based on second index information parsed from the bitstream, wherein the second index information specifies one of the plurality of motion vector precision candidates. Specifically, when it is determined that the motion vector precision of the current block is determined without using the motion vector precision set, the motion vector precision of the current block is determined without parsing the second index information from the bit stream.

2. The method according to claim 1, wherein, Based on the information parsed from the bit stream, the number or type of motion vector precision candidates in the motion vector precision set is determined.

3. A method for encoding video, the method comprising: Determine the accuracy of the motion vector for the current block; Obtain the motion vector of the current block; Generate a motion vector candidate list for the current block, the motion vector candidate list including motion vector candidates, the motion vector candidates including spatial motion vector candidates and temporal motion vector candidates; The predicted motion vector value of the current block is determined based on one of the motion vector candidates in the motion vector candidate list; The motion vector difference is derived based on the motion vector and the predicted value of the motion vector. as well as The scaled motion vector difference is encoded, wherein the scaled motion vector difference is obtained by scaling the motion vector difference based on the motion vector precision of the current block; Specifically, the first index information specifying one of the motion vector candidates is encoded in the bitstream. Specifically, a flag indicating whether the motion vector precision of the current block is determined from a set of motion vector precision candidates is encoded in the bitstream. Wherein, when the flag is encoded to indicate that the motion vector precision is determined from the set of motion vector precisions, second index information specifying one of the plurality of motion vector precision candidates is also encoded into the bitstream; and Specifically, when the flag is encoded to indicate that the motion vector precision of the current block is determined without using the motion vector precision set, the encoding of the second index information is skipped.

4. The method according to claim 3, wherein, Information used to determine the number or type of motion vector precision candidates in the motion vector precision set is encoded in the bitstream.

5. A method for sending video data, comprising: Determine the accuracy of the motion vector for the current block; Obtain the motion vector of the current block; Generate a motion vector candidate list for the current block, the motion vector candidate list including motion vector candidates, the motion vector candidates including spatial motion vector candidates and temporal motion vector candidates; The predicted motion vector value of the current block is determined based on one of the motion vector candidates in the motion vector candidate list; The motion vector difference is calculated based on the motion vector and the predicted value of the motion vector. A bitstream comprising the video data is generated by encoding a scaled motion vector difference, wherein the scaled motion vector difference is obtained by scaling the motion vector difference based on the motion vector precision of the current block; as well as Send a bitstream including the video data. Specifically, a first index information specifying one of the motion vector candidates is encoded in the bitstream. Specifically, a flag indicating whether the motion vector precision of the current block is determined from a set of motion vector precision candidates is encoded in the bitstream. Wherein, when the flag is encoded to indicate that the motion vector precision is determined from the set of motion vector precisions, second index information specifying one of the plurality of motion vector precision candidates is also encoded into the bitstream, and Specifically, when the flag is encoded to indicate that the motion vector precision is determined without using the motion vector precision set, the encoding of the second index information is skipped.

Citation Information

Patent Citations

  • Video coding and decoding method and apparatus using adaptive motion vector coding / encoding

    KR1020120080552A