Video decoding method, video encoding method, and video data transmission method

By deriving the spatial merging candidate list and motion information of the video signal, the problem of low inter-frame prediction efficiency in high-resolution video signal encoding/decoding is solved, achieving efficient parallel merging and intra-frame prediction, thus improving encoding/decoding efficiency.

CN116506597BActive Publication Date: 2026-05-05KT CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KT CORP
Filing Date
2017-08-03
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies have low inter-frame prediction efficiency when encoding/decoding high-resolution and high-quality video signals, and lack a method for merging candidate derivation based on a predetermined shape or size.

Method used

By exporting spatial merge candidates for the current block, a merge candidate list is generated, and motion information is obtained based on this list to perform motion compensation. This applies to square or non-square blocks and limits the number of merge candidates in the merge estimation region.

Benefits of technology

It achieves efficient inter-frame prediction and parallel merging processing, improves encoding/decoding efficiency, and supports intra-frame prediction and filter applications with multiple reference lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116506597B_ABST
    Figure CN116506597B_ABST
Patent Text Reader

Abstract

This application discloses a video decoding method, a video encoding method, and a video data transmission method. The video decoding method may include: determining the current block from the coded blocks according to tree-based block partitioning; generating a merging candidate list for the current block, the merging candidate list including spatial merging candidates and temporal merging candidates; obtaining motion information of the current block based on the merging candidate list; and obtaining prediction samples of the current block based on the motion information, wherein the current block is asymmetrically divided into two partitions based on a vertical line or a horizontal line, and wherein the two partitions share a merging candidate list used for inter-frame prediction of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of national application number 201780048928.9, international application date August 3, 2017, national entry date February 2, 2019, and invention title "Video Signal Processing Method and Apparatus". Technical Field

[0002] This invention relates to methods and apparatus for processing video signals. Background Technology

[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, has increased across various application areas. However, the data volume of image data with higher resolution and quality increases compared to regular image data. Therefore, the costs of transmission and storage increase when transmitting image data using media such as conventional wired and wireless broadband networks, or when storing image data using conventional storage media. To address these issues arising from the increasing resolution and quality of image data, efficient image encoding / decoding techniques can be utilized.

[0004] Image compression techniques encompass various methods, including: inter-frame prediction techniques that predict pixel values ​​included in the current image based on previous or subsequent images; intra-frame prediction techniques that predict pixel values ​​included in the current image using pixel information from the current image; and entropy coding techniques that assign short codes to frequently occurring values ​​and long codes to less frequently occurring values. Image data can be effectively compressed using such image compression techniques, and image data can be transmitted or stored.

[0005] Simultaneously, with the increasing demand for high-resolution images, the need for stereoscopic image content as a new image service has also increased. This research explores video compression techniques for effectively delivering stereoscopic image content with both high and ultra-high resolution. Summary of the Invention

[0006] Technical issues

[0007] The purpose of this invention is to provide a method and apparatus for efficiently performing inter-frame prediction on target blocks during the encoding / decoding of video signals.

[0008] The purpose of this invention is to provide a method and apparatus for deriving merging candidates based on blocks having a predetermined shape or a predetermined size when encoding / decoding video signals.

[0009] The purpose of this invention is to provide a method and apparatus for performing parallel merging in units of a predetermined shape or size when encoding / decoding video signals.

[0010] The technical objectives of this invention are not limited to the technical problems mentioned above. Furthermore, those skilled in the art will clearly understand from the following description other technical problems not mentioned.

[0011] Technical solutions

[0012] The method and apparatus for decoding video signals according to the present invention can: derive spatial merging candidates for the current block; generate a merging candidate list for the current block based on the spatial merging candidates; obtain motion information of the current block based on the merging candidate list; and perform motion compensation on the current block using the motion information.

[0013] The method and apparatus for encoding video signals according to the present invention can: derive spatial merging candidates for the current block; generate a merging candidate list for the current block based on the spatial merging candidates; obtain motion information of the current block based on the merging candidate list; and perform motion compensation on the current block using the motion information.

[0014] In the method and apparatus for encoding / decoding video signals according to the present invention, if the current block does not have a predefined shape or does not have a size equal to or greater than a predefined size, a spatial merging candidate for the current block can be derived based on a block having a predefined shape or a size equal to or greater than a predefined size, the current block being included.

[0015] In the method and apparatus for encoding / decoding video signals according to the present invention, the predefined shape may be a square shape.

[0016] In the method and apparatus for encoding / decoding video signals according to the present invention, the current block may have the same spatial merging candidate as the adjacent blocks included in the square-shaped block having the current block.

[0017] In the method and apparatus for encoding / decoding video signals according to the present invention, if the current block and spatial merging candidate are included in the same merging estimation region, it can be determined that the spatial merging candidate is unavailable.

[0018] In the method and apparatus for encoding / decoding video signals according to the present invention, the merging estimation region may have a square shape or a non-square shape.

[0019] In the method and apparatus for encoding / decoding video signals according to the present invention, if the merging estimation region has a non-square shape, the number of candidate shapes that the merging estimation region can have is limited to a predefined number.

[0020] The features briefly outlined above are merely illustrative aspects of the invention as described in the detailed description below, and do not limit the scope of the invention.

[0021] Beneficial effects

[0022] According to the present invention, efficient inter-frame prediction can be performed on the encoded / decoded target block.

[0023] According to the present invention, merging candidates can be derived based on blocks having a predetermined shape or a predetermined size.

[0024] According to the present invention, merging can be performed in parallel in units of predetermined shape or predetermined size.

[0025] According to the present invention, intra-frame prediction can be performed on the encoded / decoded target block by selecting at least one of a plurality of reference lines.

[0026] According to the present invention, reference lines can be derived based on blocks having a predetermined shape or having a size equal to or greater than a predetermined size.

[0027] According to the present invention, an intra-frame filter can be applied to at least one of a plurality of reference lines.

[0028] According to the present invention, the number of intra-prediction modes or intra-prediction modes can be adaptively determined based on the reference line used for intra-prediction of the current block.

[0029] The effects achievable by this invention are not limited to those mentioned above, and those skilled in the art can clearly understand from the following description other effects not mentioned. Attached Figure Description

[0030] Figure 1 This is a block diagram illustrating an apparatus for encoding video according to an embodiment of the present invention.

[0031] Figure 2 This is a block diagram illustrating an apparatus for decoding video according to an embodiment of the present invention.

[0032] Figure 3 This is a diagram illustrating an example of hierarchical partitioning of coding blocks based on a tree structure according to an embodiment of the present invention.

[0033] Figure 4 This is a diagram illustrating partitioning types that allow binary tree-based partitioning according to an embodiment of the present invention.

[0034] Figure 5 This is a diagram illustrating an example of binary tree-based partitioning that allows only predetermined types according to an embodiment of the present invention.

[0035] Figure 6 This is a diagram illustrating an example of how information related to the allowed number of binary tree partitions is encoded / decoded according to an embodiment of the present invention.

[0036] Figure 7 This is a diagram illustrating a partitioning pattern applicable to coded blocks according to an embodiment of the present invention.

[0037] Figure 8 This is a flowchart illustrating an inter-frame prediction method according to an embodiment of the present invention.

[0038] Figure 9 This is a diagram illustrating the process of deriving motion information for the current block when a merge mode is applied to the current block.

[0039] Figure 10 This demonstrates the process of deriving motion information for the current block when the AMVP mode is applied to the current block.

[0040] Figure 11 This is a graph showing the spatial merging candidates for the current block.

[0041] Figure 12 This is a diagram showing the co-located blocks of the current block.

[0042] Figure 13 This is a diagram illustrating an example of obtaining temporally merged candidate motion vectors by scaling the motion vectors of the co-position blocks.

[0043] Figure 14 This is a diagram illustrating an example of deriving merge candidates for non-square blocks based on square blocks.

[0044] Figure 15 This is a diagram used to illustrate an example of merging candidate binary tree partitions derived from the upper node block.

[0045] Figure 16 This is a diagram illustrating an example of determining the availability of spatial merging candidates based on the merging estimation region.

[0046] Figure 17 This is a flowchart illustrating the process of obtaining residual samples according to an embodiment of the present invention. Detailed Implementation

[0047] Various modifications can be made to this invention, and various embodiments of the invention exist. Examples of embodiments will now be provided with reference to the accompanying drawings, and described in detail. However, the invention is not limited thereto, and exemplary embodiments can be interpreted as including all modifications, equivalents, or alternatives within the technical concept and scope of the invention. In describing the drawings, similar reference numerals refer to similar elements.

[0048] The terms "first," "second," etc., used in this specification may be used to describe various components, but these components are not to be construed as limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the invention, a "first" component may be referred to as a "second" component, and a "second" component may similarly be referred to as a "first" component. The term "and / or (and / or)" includes a combination of multiple items or any one of multiple terms.

[0049] It should be understood that in this specification, when an element is simply referred to as "connected to" or "coupled to" another element rather than "directly connected to" or "directly coupled to" another element, the element may be "directly connected to" or "directly coupled to" another element, or the element may be connected to or coupled to another element with other elements in between. Conversely, it should be understood that when an element is referred to as "directly coupled to" or "directly connected to" another element, there are no intermediate elements.

[0050] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Expressions used in the singular include expressions in the plural unless the expression clearly has a different meaning in the context. It should be understood in this specification that terms such as “comprising,” “having,” etc., are intended to indicate the presence of features, numbers, steps, actions, elements, portions, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, portions, or combinations thereof may be present or added.

[0051] Preferred embodiments of the invention will be described in detail below with reference to the accompanying drawings. In the drawings, the same constituent elements are indicated by the same reference numerals, and repeated descriptions of the same elements will be omitted.

[0052] Figure 1 This is a block diagram illustrating an apparatus for encoding video according to an embodiment of the present invention.

[0053] Reference Figure 1 The apparatus (100) for encoding video may include: an image segmentation module (110), a prediction module (120, 125), a transformation module (130), a quantization module (135), a rearrangement module (160), an entropy coding module (165), an inverse quantization module (140), an inverse transformation module (145), a filter module (150), and a memory (155).

[0054] Figure 1The components shown are presented independently to represent different functional features within the apparatus for encoding video. Therefore, it is not implied that each component is composed of separate hardware or software units. In other words, for convenience, each component includes every one of the listed components. Thus, at least two components of each component can be combined to form a single component, or a single component can be divided into multiple components to perform each function. Embodiments of combining each component and embodiments of dividing a component are also included within the scope of this invention without departing from its spirit.

[0055] Furthermore, some components may not be essential for performing the basic functions of the invention, but rather optional components used only to improve the performance of the invention. The invention can be implemented by excluding components used to improve performance and including only the essential components for achieving the essence of the invention. Structures that exclude optional components used only to improve performance and include only the essential components are also included within the scope of the invention.

[0056] The image partitioning module (110) can partition an input image into one or more processing units. Here, the processing unit can be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The image partitioning module (110) can partition an image into a combination of multiple coding units, prediction units, and transform units, and can encode the image by selecting a combination of coding units, prediction units, and transform units using a predetermined criterion (e.g., a cost function).

[0057] For example, an image can be divided into multiple coding units. A recursive tree structure, such as a quadtree, can be used to divide the image into coding units. Coding units that are further divided into other coding units, with the image or the largest coding unit as the root, can be divided in such a way that the number of child nodes corresponds to the number of coding units. Coding units that are no longer divided according to predetermined constraints are used as leaf nodes. That is, when it is assumed that only a square partition is feasible for a coding unit, a coding unit can be divided into at most four other coding units.

[0058] In the following, in embodiments of the present invention, the encoding unit may refer to a unit that performs encoding or a unit that performs decoding.

[0059] A prediction unit can be one of the partitions in a single coding unit that are divided into square or rectangular shapes of the same size, or a prediction unit can be one of the partitions in a single coding unit that are divided into partitions with different shapes / sizes.

[0060] When a prediction unit to be subjected to intra-frame prediction is generated based on a coding unit and the coding unit is not the smallest coding unit, intra-frame prediction can be performed without dividing the coding unit into multiple prediction units N×N.

[0061] The prediction modules (120, 125) may include an inter-frame prediction module (120) that performs inter-frame prediction and an intra-frame prediction module (125) that performs intra-frame prediction. It can be determined whether inter-frame or intra-frame prediction is performed for a prediction unit, and details based on each prediction method (e.g., intra-frame prediction mode, motion vectors, reference image, etc.) can be determined. Here, the processing unit to be predicted may be different from the processing unit for which the prediction method and details have been determined. For example, the prediction method, prediction mode, etc., can be determined by the prediction unit, and the prediction can be performed by the transformation unit. The residual value (residual block) between the generated prediction block and the original block can be input to the transformation module (130). Furthermore, prediction mode information, motion vector information, etc., used for prediction can be encoded together with the residual value by the entropy coding module (165) and can be transmitted to the device for decoding the video. When using a specific coding mode, the original block can be encoded as is and transmitted to the device for decoding the video without generating a prediction block through the prediction modules (120, 125).

[0062] The inter-frame prediction module (120) can predict the prediction unit based on information from at least one of the previous or subsequent images of the current image, or in some cases, it can predict the prediction unit based on information from some coded regions in the current image. The inter-frame prediction module (120) may include a reference picture interpolation module, a motion prediction module, and a motion compensation module.

[0063] The reference image interpolation module can receive reference image information from the memory (155) and generate pixel information in whole pixels or smaller than whole pixels based on the reference image. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information in units of 1 / 4 pixels or smaller than whole pixels. In the case of chrominance signals, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information in units of 1 / 8 pixels or smaller than whole pixels.

[0064] The motion prediction module can perform motion prediction based on a reference image interpolated by the reference image interpolation module. Various methods can be used to calculate motion vectors, such as Full Search-Based Block Matching (FBMA), Three-Step Search (TSS), and New Three-Step Search (NTS). Based on the interpolated pixels, the motion vector can have motion vector values ​​in units of 1 / 2 pixel or 1 / 4 pixel. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods can be used as motion prediction methods, such as skipping methods, merging methods, AMVP (Advanced Motion Vector Prediction) methods, and intra-block copying methods.

[0065] The intra-frame prediction module (125) can generate prediction units based on reference pixel information adjacent to the current block, which serves as pixel information in the current image. When the neighboring block of the current prediction unit is a block undergoing inter-frame prediction and therefore the reference pixel is a pixel undergoing inter-frame prediction, the reference pixel information of the neighboring block undergoing intra-frame prediction can be used to replace the reference pixel included in the block undergoing inter-frame prediction. That is, when the reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information.

[0066] Intra-frame prediction modes can include directional prediction modes that use reference pixel information based on the prediction direction and non-directional prediction modes that do not use directional information when performing prediction. The mode used to predict luminance information can be different from the mode used to predict chrominance information, and to predict chrominance information, the intra-frame prediction mode information used to predict luminance information or the predicted luminance signal information can be used.

[0067] When performing intra-prediction, if the size of the prediction unit is the same as the size of the transform unit, intra-prediction can be performed on the prediction unit based on the pixels located to its left, upper left, and top. However, when performing intra-prediction, if the size of the prediction unit is different from the size of the transform unit, intra-prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-prediction using an N×N partition can be used only for the smallest coding unit.

[0068] In intra-frame prediction methods, prediction blocks can be generated after applying an AIS (Adaptive Intra-Frame Smoothing) filter to a reference pixel based on the prediction mode. The type of AIS filter applied to the reference pixel can vary. To perform intra-frame prediction, the intra-frame prediction mode of the current prediction unit can be predicted based on the intra-frame prediction modes of prediction units adjacent to the current prediction unit. When predicting the prediction mode of the current prediction unit using mode information predicted by neighboring prediction units, if the intra-frame prediction mode of the current prediction unit is the same as that of the neighboring prediction units, predetermined tagging information can be used to transmit information indicating that the prediction modes of the current prediction unit and those of the neighboring prediction units are the same. If the prediction mode of the current prediction unit is different from that of the neighboring prediction units, entropy coding can be performed to encode the prediction mode information of the current block.

[0069] Furthermore, residual blocks containing information related to residual values, which are the differences between the predicted unit and the original block of the predicted unit, can be generated based on the predicted unit generated by the prediction modules (120, 125). The generated residual blocks can be input to the transformation module (130).

[0070] The transform module (130) can transform the residual block, which includes information about the residual values ​​between the original block and the prediction units generated by the prediction modules (120, 125), using transform methods such as Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and KLT. The choice between applying DCT, DST, or KLT to transform the residual block can be determined based on the intra-frame prediction mode information of the prediction units used to generate the residual block.

[0071] The quantization module (135) can quantize the values ​​transformed to the frequency domain by the transformation module (130). The quantization coefficients can vary depending on the importance or block of the image. The values ​​calculated by the quantization unit (135) can be provided to the inverse quantization module (140) and the rearrangement module (160).

[0072] The rearrangement module (160) can rearrange the coefficients of the quantized residuals.

[0073] The rearrangement module (160) can transform coefficients in two-dimensional block form into coefficients in one-dimensional vector form using a coefficient scanning method. For example, the rearrangement module (160) can use a zigzag scanning method to scan from DC coefficients to coefficients in the high-frequency domain in order to transform the coefficients into one-dimensional vector form. Depending on the size of the transform unit and the intra-frame prediction mode, a vertical scan along the column direction or a horizontal scan along the row direction can be used instead of a zigzag scan. That is, the choice between a zigzag scan, a vertical scan, and a horizontal scan can be determined based on the size of the transform unit and the intra-frame prediction mode.

[0074] The entropy coding module (165) can perform entropy coding based on the value calculated by the rearrangement module (160). Entropy coding can use various coding methods, such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC).

[0075] The entropy coding module (165) can encode various information from the rearrangement module (160) and the prediction module (120, 125), such as residual coefficient information and block type information of coding units, prediction mode information, partitioning unit information, prediction unit information, transform unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc.

[0076] The entropy coding module (165) can entropy code the coefficients of the coding units input from the rearrangement module (160).

[0077] The inverse quantization module (140) can inverse quantize the value quantized by the quantization module (135), and the inverse transform module (145) can inverse transform the value transformed by the transform module (130). The residual value generated by the inverse quantization module (140) and the inverse transform module (145) can be combined with the prediction units predicted by the motion estimation module, motion compensation module and intra-frame prediction module of the prediction modules (120, 125) to generate a reconstruction block.

[0078] The filter module (150) may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).

[0079] Deblocking filters remove block distortion caused by boundaries between blocks in a reconstructed image. To determine whether to perform deblocking, the pixels included in several rows or columns of a block can be the basis for deciding whether to apply a deblocking filter to the current block. When applying a deblocking filter to a block, a strong or weak filter can be applied depending on the desired deblocking intensity. Furthermore, horizontal and vertical filtering can be processed in parallel when applying a deblocking filter.

[0080] The offset correction module can correct the offset between the deblocked image and the original image on a pixel-by-pixel basis. To perform offset correction on a specific image, a method that considers the edge information of each pixel can be used to apply the offset, or the following method can be used: divide the image's pixels into a predetermined number of regions, determine the regions to be offset, and apply the offset to the determined regions.

[0081] Adaptive Loop Filtering (ALF) can be performed based on values ​​obtained by comparing the filtered reconstructed image with the original image. Pixels included in the image can be divided into predetermined groups, the filter to be applied to each group can be determined, and filtering can be performed individually for each group. Information regarding whether ALF is applied and the luminance signal can be transmitted via the coding unit (CU). The shape and filter coefficients of the filter used for ALF can vary depending on each block. Furthermore, a filter of the same shape (fixed shape) for ALF can be applied regardless of the characteristics of the target block.

[0082] The memory (155) can store the reconstructed blocks or reconstructed images calculated by the filter module (150). The stored reconstructed blocks or reconstructed images can be provided to the prediction modules (120, 125) when performing inter-frame prediction.

[0083] Figure 2 This is a block diagram illustrating an apparatus for decoding video according to an embodiment of the present invention.

[0084] Reference Figure 2 The device (200) for decoding video may include: an entropy decoding module (210), a rearrangement module (215), an inverse quantization module (220), an inverse transform module (225), a prediction module (230, 235), a filter module (240), and a memory (245).

[0085] When a video bitstream is input from a device used for encoding video, the input bitstream can be decoded by reversing the process of the device used for encoding video.

[0086] The entropy decoding module (210) can perform entropy decoding according to the inverse process of entropy encoding performed by the entropy encoding module of the device for encoding video. For example, various methods can be applied corresponding to the method performed by the device for encoding video, such as exponential Golomb coding, context adaptive variable-length coding (CAVLC), and context adaptive binary arithmetic coding (CABAC).

[0087] The entropy decoding module (210) can decode information about intra-frame prediction and inter-frame prediction performed by the means for encoding the video.

[0088] The rearrangement module (215) can perform rearrangement on the bitstream entropy decoded by the entropy decoding module (210) based on the rearrangement method used in the apparatus for encoding video. The rearrangement module can reconstruct and rearrange coefficients in one-dimensional vector form into coefficients in two-dimensional block form. The rearrangement module (215) can receive information related to the coefficient scan performed in the apparatus for encoding video, and can perform rearrangement via a method of inverse scanning of coefficients based on the scan order performed in the apparatus for encoding video.

[0089] The inverse quantization module (220) can perform inverse quantization based on the quantization parameters received from the means for encoding the video and the coefficients of the rearranged blocks.

[0090] The inverse transform module (225) can perform inverse transforms on the quantization results of the means for encoding video, which are the inverse processes of the transforms (i.e., DCT, DST, and KLT) performed by the transform module, namely, inverse DCT, inverse DST, and inverse KLT. The inverse transform can be performed based on the transmission unit determined by the means for encoding video. The inverse transform module (225) of the means for decoding video can selectively perform transform schemes (e.g., DCT, DST, KLT) based on multiple pieces of information such as prediction method, current block size, prediction direction, etc.

[0091] The prediction modules (230, 235) can generate prediction blocks based on information about the prediction blocks received from the entropy decoding module (210) and information about previously decoded blocks or images received from the memory (245).

[0092] As described above, similar to the operation of a device for encoding video, when performing intra-frame prediction, if the size of the prediction unit is the same as the size of the transform unit, intra-frame prediction can be performed based on the pixels located to the left, upper left, and top of the prediction unit. When performing intra-frame prediction, if the size of the prediction unit is different from the size of the transform unit, intra-frame prediction can be performed using reference pixels based on the transform unit. Furthermore, intra-frame prediction using an N×N partition can be used only for the smallest coding unit.

[0093] The prediction modules (230, 235) may include a prediction unit determination module, an inter-frame prediction module, and an intra-frame prediction module. The prediction unit determination module can receive various information from the entropy decoding module (210), such as prediction unit information, prediction mode information of the intra-frame prediction method, motion prediction information about the inter-frame prediction method, etc., and can divide the current coding unit into prediction units and determine whether to perform inter-frame prediction or intra-frame prediction on the prediction unit. By using the information required for inter-frame prediction of the current prediction unit received from the means for encoding video, the inter-frame prediction module (230) can perform inter-frame prediction on the current prediction unit based on information from at least one of the previous or subsequent images of the current image including the current prediction unit. Alternatively, inter-frame prediction can be performed based on information from some pre-reconstructed regions in the current image including the current prediction unit.

[0094] To perform inter-frame prediction, it can be determined for the coding unit which of the skip mode, merge mode, AMVP mode, and inter-block copy mode will be used as the motion prediction method for the prediction unit included in the coding unit.

[0095] The intra-prediction module (235) can generate prediction blocks based on pixel information in the current image. When the prediction unit is a prediction unit undergoing intra-prediction, intra-prediction can be performed based on the intra-prediction mode information of the prediction unit received from the means for encoding the video. The intra-prediction module (235) may include an adaptive intra-smoothing (AIS) filter, a reference pixel interpolation module, and a DC filter. The AIS filter performs filtering on the reference pixels of the current block, and can determine whether to apply the filter based on the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode of the prediction unit and the AIS filter information received from the means for encoding the video. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.

[0096] When the prediction mode of the prediction unit is a prediction mode that performs intra-frame prediction based on pixel values ​​obtained by interpolating reference pixels, the reference pixel interpolation module can interpolate the reference pixels to generate reference pixels that are integers or smaller than integers. When the prediction mode of the current prediction unit is a prediction mode that generates prediction blocks without interpolating reference pixels, interpolation of reference pixels is not required. When the prediction mode of the current block is DC mode, the DC filter can generate prediction blocks through filtering.

[0097] A reconstructed block or reconstructed image can be provided to the filter module (240). The filter module (240) may include a deblocking filter, an offset correction module, and an ALF.

[0098] Information regarding whether a deblocking filter should be applied to a corresponding block or image can be received from the means for encoding video, as well as information about which filter, strong or weak, should be applied when the deblocking filter is applied. The deblocking filter of the means for decoding video can receive information about the deblocking filter from the means for encoding video and can perform deblocking filtering on the corresponding block.

[0099] The offset correction module can perform offset correction on the reconstructed image based on the type and offset value information of the offset correction applied to the image during encoding.

[0100] ALF can be applied to the coding unit based on information received from the device used to encode the video, such as whether ALF is applied and ALF coefficient information. ALF information can be provided as included in a specific parameter set.

[0101] The memory (245) can store reconstructed images or reconstructed blocks for use as reference images or reference blocks, and can provide reconstructed images to the output module.

[0102] As described above, in embodiments of the present invention, for ease of explanation, the term "encoding unit" is used to refer to a unit used for encoding; however, the term "encoding unit" can also be used to refer to a unit that performs both decoding and encoding.

[0103] Furthermore, the current block can represent the target block to be encoded / decoded. And, depending on the encoding / decoding steps, the current block can represent a coding tree block (or coding tree unit), a coding block (or coding unit), a transform block (or transform unit), a prediction block (or prediction unit), etc.

[0104] Images can be encoded / decoded by dividing them into basic blocks of square or non-square shapes. These basic blocks are called coding tree units (CRUs). A CRU can be defined as the largest allowed coding unit within a sequence or slice. Information about the shape of the CRU, whether it is square or non-square, or its size, can be signaled via sequence parameter sets, image parameter sets, or slice headers. A CRU can be further divided into smaller partitions. For example, if the depth of a partition generated by dividing a CRU is 1, then the depth of a partition generated by dividing a partition with depth 1 can be defined as 2. That is, a partition generated by dividing a CRU into partitions with depth k can be defined as having depth k+1.

[0105] A partition of arbitrary size generated by segmenting coding tree units can be defined as a coding unit. A coding unit can be recursively segmented or divided into basic units for performing prediction, quantization, transform, or loop filtering. For example, a partition of arbitrary size generated by segmenting coding units can be defined as a coding unit, or it can be defined as a transform unit or prediction unit, which are basic units for performing prediction, quantization, transform, or loop filtering.

[0106] The partitioning of a coding tree unit or coding unit can be performed based on at least one of vertical and horizontal lines. Furthermore, the number of vertical or horizontal lines used to partition the coding tree unit or coding unit can be at least one or more. For example, a coding tree unit or coding unit can be divided into two partitions using one vertical line or one horizontal line, or into three partitions using two vertical lines or two horizontal lines. Alternatively, a coding tree unit or coding unit can be divided into four partitions with a length and width of 1 / 2 using one vertical line and one horizontal line.

[0107] When a coding tree unit or coding unit is divided into multiple partitions using at least one vertical line or at least one horizontal line, the partitions can have a uniform size or different sizes. Alternatively, any partition can have a different size than the other partitions.

[0108] In the embodiments described below, it is assumed that the coding tree unit or coding unit is divided into a quadtree structure or a binary tree structure. However, it is also possible to use a greater number of vertical lines or a greater number of horizontal lines to divide the coding tree unit or coding unit.

[0109] Figure 3 This is a diagram illustrating an example of hierarchical partitioning of coding blocks based on a tree structure according to an embodiment of the present invention.

[0110] The input video signal is decoded in predetermined block units. The default unit for decoding the input video signal is the coding block. A coding block can be a block that performs intra / inter-frame prediction, transform, and quantization. Furthermore, a prediction mode (e.g., intra-frame prediction mode or inter-frame prediction mode) is determined on a block-by-block basis, and prediction blocks included in a coding block can share the determined prediction mode. A coding block can be a square or non-square block with any size ranging from 8×8 to 64×64, or it can be a square or non-square block with a size of 128×128, 256×256, or larger.

[0111] Specifically, the coded blocks can be hierarchically partitioned based on at least one of quadtrees and binary trees. Here, quadtree-based partitioning means dividing a 2N×2N coded block into four N×N coded blocks, and binary tree-based partitioning means dividing one coded block into two coded blocks. Even when performing binary tree-based partitioning, square-shaped coded blocks may still exist at lower depths.

[0112] Binary tree-based partitioning can be performed symmetrically or asymmetrically. The coded blocks resulting from binary tree-based partitioning can be square or non-square, for example, rectangular in shape. For example, partitioning types that allow binary tree-based partitioning may include at least one of the following: symmetric types 2N×N (horizontal non-square coded units) or N×2N (vertical non-square coded units); asymmetric types nL×2N; nR×2N; 2N×nU or 2N×nD.

[0113] Binary tree-based partitioning can be restricted to either symmetric or asymmetric partitioning. In this case, constructing a coding tree unit with square blocks corresponds to a quadtree CU partition, and constructing a coding tree unit with symmetric non-square blocks corresponds to a binary tree partition. Constructing coding tree units with both square blocks and symmetric non-square blocks corresponds to quadtree CU partitioning and binary tree CU partitioning, respectively.

[0114] Binary tree-based partitioning can be performed on coded blocks that no longer require quadtree-based partitioning. Quadtree-based partitioning can be omitted from coded blocks that already require binary tree-based partitioning.

[0115] Furthermore, the partitioning at lower depths can be determined based on the partitioning type at higher depths. For example, if binary tree-based partitioning is allowed at two or more depths, then only the same type of binary tree partitioning as the higher depth partitioning can be allowed at lower depths. For instance, if a binary tree-based partitioning at a higher depth is performed using a 2N×N type, then a binary tree-based partitioning at a lower depth is also performed using a 2N×N type. Alternatively, if a binary tree-based partitioning at a higher depth is performed using an N×2N type, then a binary tree-based partitioning at a lower depth is also performed using an N×2N type.

[0116] In contrast, it is also possible to allow only different types of binary tree partitioning at lower depths compared to those at higher depths.

[0117] Binary tree-based partitions of a specific type can be restricted to sequences, slices, coded tree units, or coded units. For example, for coded tree units, only 2N×N or N×2N type binary tree-based partitions can be allowed. Available partition types can be predefined in the encoder or decoder. Alternatively, information related to available partition types or unavailable partition types can be encoded and then transmitted as a signal via a bitstream.

[0118] Figure 5 This is a diagram illustrating an example of binary tree-based partitioning that only allows specific types of partitioning. Figure 5 (a) shows an example that only allows binary tree-based partitioning of type N×2N, and Figure 5 (b) shows an example of binary tree-based partitioning that only allows 2N×N types. To achieve adaptive partitioning based on quadtrees or binary trees, the following information can be used: information indicating quadtree-based partitioning, information related to the size / depth of the coding blocks that allow quadtree-based partitioning, information indicating binary tree-based partitioning, information related to the size / depth of the coding blocks that allow binary tree-based partitioning, information related to the size / depth of the coding blocks that do not allow binary tree-based partitioning, information related to whether binary tree-based partitioning is performed in the vertical or horizontal direction, etc.

[0119] Furthermore, information relating to the number of allowed binary tree partitions, the depth of allowed binary tree partitions, or the number of allowed binary tree partition depths can be obtained for a specific coding tree unit or coding unit. This information can be encoded at the coding tree unit or coding unit level and transmitted to the decoder via a bitstream.

[0120] For example, the syntax "max_binary_depth_idx_minus1" indicating the maximum allowed depth of a binary tree partition can be encoded / decoded via a bitstream. In this case, max_binary_depth_idx_minus1+1 can indicate the maximum allowed depth of a binary tree partition.

[0121] Reference Figure 6 The example shown is in Figure 6 In the bitstream, binary tree partitioning has already been performed for coding units of depth 2 and depth 3. Therefore, at least one of the following information can be encoded / decoded via the bitstream: information indicating the number of times binary tree partitioning has been performed in the coding unit (i.e., 2 times), information indicating the maximum depth of binary tree partitioning allowed in the coding unit (i.e., depth 3), or information indicating the number of depths of binary tree partitioning performed in the coding unit (i.e., 2 (depth 2 and depth 3)).

[0122] As another example, at least one of the following can be obtained for each sequence or slice: information relating to the number of allowed binary tree partitions, information relating to the depth of allowed binary tree partitions, or information relating to the number of depths of allowed binary tree partitions. For example, the information can be encoded in units of sequences, images, or slices and transmitted via a bitstream. Therefore, at least one of the following—the number of binary tree partitions in the first slice, the maximum depth of allowed binary tree partitions in the first slice, or the number of depths at which binary tree partitions are performed in the first slice—can differ from that in the second slice. For example, in the first slice, binary tree partitions can be allowed for only one depth, while in the second slice, binary tree partitions can be allowed for two depths.

[0123] As another example, the allowed number of binary tree partitions, the allowed depth of binary tree partitions, or the allowed depth of binary tree partitions can be set differently based on the temporal ID of the slice or image. Here, the temporal ID is used to identify each of multiple video layers that have scalability in at least one of view, space, time, or quality.

[0124] like Figure 3 As shown, a first coding block 300 with a partitioning depth (splitting depth) of k can be divided into multiple second coding blocks based on a quadtree. For example, second coding blocks 310 to 340 can be square blocks with a width and height that are half the width and height of the first coding block, and the partitioning depth of the second coding blocks can be increased to k+1.

[0125] A second coding block 310 with a partition depth of k+1 can be partitioned into multiple third coding blocks with a partition depth of k+2. The partitioning of the second coding block 310 can be performed by selectively using either a quadtree or a binary tree, depending on the partitioning method. Here, the partitioning method can be determined based on at least one of information indicating quadtree-based partitioning and information indicating binary tree-based partitioning.

[0126] When the second coding block 310 is partitioned based on a quadtree, it can be divided into four third coding blocks 310a, each with a width and height half that of the second coding block, and the partitioning depth of the third coding blocks 310a can be increased to k+2. In contrast, when the second coding block 310 is partitioned based on a binary tree, it can be divided into two third coding blocks. Here, each of the two third coding blocks can be a non-square block with one of half the width and half the height of the second coding block, and the partitioning depth can be increased to k+2. The second coding block can be determined as a horizontal or vertical non-square block based on the partitioning direction, and the partitioning direction can be determined based on information about whether the binary tree-based partitioning is performed along the vertical or horizontal direction.

[0127] Meanwhile, the second coding block 310 can be determined as a leaf coding block that is no longer based on a quadtree or binary tree partition. In this case, the leaf coding block can be used as a prediction block or a transform block.

[0128] Similar to the division of the second coding block 310, the third coding block 310a can be determined as a leaf coding block, or it can be further divided based on a quadtree or a binary tree.

[0129] Simultaneously, the third coding block 310b based on binary tree partitioning can be further partitioned into vertical coding blocks 310b-2 or horizontal coding blocks 310b-3 based on the binary tree, and the partitioning depth of the relevant coding blocks can be increased to k+3. Alternatively, the third coding block 310b can be determined as a leaf coding block 310b-1 that is no longer based on binary tree partitioning. In this case, coding block 310b-1 can be used as a prediction block or a transform block. However, the above partitioning process can be performed restrictively based on at least one of the following information: information related to the size / depth of coding blocks that allow quadtree-based partitioning, information related to the size / depth of coding blocks that allow binary tree-based partitioning, and information related to the size / depth of coding blocks that do not allow binary tree-based partitioning.

[0130] The number of candidates representing the size of a coded block can be limited to a predetermined number, or the size of the coded block in a predetermined unit can have a fixed value. For example, the size of a coded block in a sequence or image can be limited to 256×256, 128×128, or 32×32. Information indicating the size of the coded blocks in a sequence or image can be sent via signals through the sequence header or image header.

[0131] As a result of partitioning based on quadtrees and binary trees, the coding unit can be represented as a square or rectangle of arbitrary size.

[0132] The coded block is encoded using at least one of skip mode, intra-frame prediction, inter-frame prediction, or skip method. Once the coded block is determined, the prediction block can be determined by the prediction partitioning of the coded block. The prediction partitioning of the coded block can be performed by a partition mode (Part_mode) that indicates the partitioning type of the coded block. The size or shape of the prediction block can be determined based on the partition mode of the coded block. For example, the size of the prediction block determined based on the partition mode can be equal to or smaller than the size of the coded block.

[0133] Figure 7 This is a diagram showing the partitioning modes that can be applied to a coding block when encoding the coding block via inter-frame prediction.

[0134] When encoding a coded block using inter-frame prediction, one of eight partitioning patterns can be applied to the coded block, such as... Figure 4 As shown in the example.

[0135] When encoding a coded block using intra-frame prediction, the partitioning pattern PART_2N×2N or the partitioning pattern PART_N×N can be applied to the coded block.

[0136] When the coded block has a minimum size, PART_N×N can be applied. Here, the minimum size of the coded block can be predefined in the encoder and decoder. Alternatively, information related to the minimum size of the coded block can be sent via a signal through the bitstream. For example, the minimum size of the coded block can be sent via a signal in the chip header, allowing the minimum size of the coded block to be defined on a chip-by-chip basis.

[0137] Typically, prediction blocks can have sizes ranging from 64×64 to 4×4. However, when encoding blocks via inter-frame prediction, restrictions can be placed so that prediction blocks do not have a 4×4 size to reduce memory bandwidth when performing motion compensation.

[0138] Figure 8 This is a flowchart illustrating an inter-frame prediction method according to an embodiment of the present invention.

[0139] Reference Figure 8 The motion information of the current block is determined (S810). The motion information of the current block may include at least one of the motion vectors related to the current block, the reference image index of the current block, or the inter-frame prediction direction of the current block.

[0140] The motion information of the current block can be obtained based on at least one of the information transmitted by signal via bit stream or the motion information of neighboring blocks adjacent to the current block.

[0141] Figure 9 This is a diagram illustrating the process of deriving motion information for the current block when a merge mode is applied to the current block.

[0142] If a merge mode is applied to the current block, spatial merge candidates can be derived from the spatially adjacent blocks of the current block (S910). Spatially adjacent blocks can refer to at least one of the blocks that are adjacent to the top, left, or corner (e.g., at least one of the top left, top right, or bottom left corners) of the current block.

[0143] The motion information of spatial merging candidates can be set to be the same as the motion information of spatially adjacent blocks.

[0144] Temporal merging candidates can be derived from temporally adjacent blocks of the current block (S920). A temporally adjacent block can refer to a co-located block included in a collocated picture. The collocated picture has a different picture order count (POC) than the current picture that includes the current block. The collocated picture can be determined as a picture with a predefined index from a list of reference pictures, or it can be determined by an index sent from the bitstream via a signal. A temporally adjacent block can be determined as a block in the collocated picture that has the same position and size as the current block, or a block adjacent to a block that has the same position and size as the current block. For example, at least one of the blocks in the collocated picture that include the center coordinates of a block with the same position and size as the current block, or blocks adjacent to the lower right boundary of that block, can be determined as a temporally adjacent block.

[0145] Motion information for temporal merging candidates can be determined based on motion information from temporally adjacent blocks. For example, the motion vector of a temporal merging candidate can be determined based on the motion vector of a temporally adjacent block. Furthermore, the inter-frame prediction direction of the temporal merging candidate can be set to the same as the inter-frame prediction direction of the temporally adjacent blocks. However, the reference image index of the temporal merging candidate can have a fixed value. For example, the reference image index of the temporal merging candidate can be set to "0".

[0146] Reference Figures 11 to 16 An example of deriving merge candidates will be described in more detail below.

[0147] Subsequently, a merge candidate list including spatial merge candidates and temporal merge candidates can be generated (S930). If the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, a merge candidate that combines two or more merge candidates or a merge candidate with zero motion vector (0,0) can be included in the merge candidate list.

[0148] When generating the merge candidate list, at least one of the merge candidates to be included in the merge candidate list can be specified based on the merge candidate index (S940).

[0149] The motion information of the current block can be set to be the same as the motion information of the merge candidate specified by the merge candidate index (S950). For example, when a spatial merge candidate is selected by the merge candidate index, the motion information of the current block can be set to be the same as the motion information of the spatially adjacent block. Alternatively, when a temporal merge candidate is selected by the merge candidate index, the motion information of the current block can be set to be the same as the motion information of the temporally adjacent block.

[0150] Figure 10 This demonstrates the process of deriving motion information for the current block when AMVP mode is applied to the current block.

[0151] When the AMVP mode is applied to the current block, at least one of the inter-frame prediction direction or the reference image index of the current block can be decoded from the bitstream (S1010). That is, when the AMVP mode is applied, at least one of the inter-frame prediction direction or the reference image index of the current block can be determined based on the information encoded by the bitstream.

[0152] Spatial motion vector candidates can be determined based on the motion vectors of the spatially adjacent blocks of the current block (S1020). The spatial motion vector candidates can include at least one of a first spatial motion vector candidate derived from the top adjacent block of the current block and a second spatial motion vector candidate derived from the left adjacent block of the current block. Here, the top adjacent block can include at least one of the blocks adjacent to the top or upper right corner of the current block, and the left adjacent block of the current block can include at least one of the blocks adjacent to the left or lower left corner of the current block. The block adjacent to the upper left corner of the current block can be considered either a top adjacent block or a left adjacent block.

[0153] When the reference images between the current block and its spatial neighbors are different, the spatial motion vector can be obtained by scaling the motion vectors of the spatial neighbors.

[0154] Candidate temporal motion vectors can be determined based on the motion vectors of the temporally adjacent blocks of the current block (S1030). If the reference images between the current block and the temporally adjacent blocks are different, the temporal motion vectors can be obtained by scaling the motion vectors of the temporally adjacent blocks.

[0155] A list of motion vector candidates, including spatial motion vector candidates and temporal motion vector candidates, can be generated (S1040).

[0156] When generating a motion vector candidate list, at least one of the motion vector candidates included in the motion vector candidate list can be specified based on information from at least one of the specified motion vector candidate lists (S1050).

[0157] The motion vector candidate specified by this information is set as the motion vector prediction value for the current block. Then, the motion vector for the current block is obtained by adding the motion vector difference to the motion vector prediction value (S1060). At this point, the motion vector difference can be parsed from the bitstream.

[0158] When the motion information of the current block is obtained, motion compensation can be performed on the current block based on the obtained motion information (S820). More specifically, motion compensation can be performed on the current block based on the inter-frame prediction direction, reference image index, and motion vector of the current block.

[0159] The maximum number of merge candidates that can be included in the merge candidate list can be sent via a bitstream signal. For example, information indicating the maximum number of merge candidates can be sent via a sequence parameter or a picture parameter.

[0160] The number of spatial and temporal merge candidates that can be included in the merge candidate list can be determined based on the maximum number of merge candidates. Specifically, the number of spatial and temporal merge candidates can be adjusted so that the total number of spatial and temporal merge candidates does not exceed the maximum number of merge candidates N. For example, when the maximum number of merge candidates is 5, 4 spatial merge candidates selected from the 5 spatial merge candidates of the current block can be added to the merge candidate list, and 1 temporal merge candidate selected from the 2 temporal merge candidates of the current block can be added to the merge candidate list. The number of temporal merge candidates can be adjusted based on the number of spatial merge candidates added to the merge candidate list, or the number of spatial merge candidates can be adjusted based on the number of temporal merge candidates added to the merge candidate list. If the number of merge candidates added to the merge candidate list is less than 5, a combined merge candidate that combines at least two merge candidates can be added to the merge candidate list, or a merge candidate with a motion vector of (0, 0) can be added to the merge candidate list.

[0161] Merge candidates can be added to the merge candidate list in a predefined order. For example, the merge candidate list can be generated in the order of spatial merge candidates, temporal merge candidates, combined merge candidates, and merge candidates with zero motion vectors. It is also possible to define an order in which merge candidates are added that differs from the enumeration order.

[0162] Figure 11 This is a diagram showing the spatial merge candidates for the current block. Spatial merge candidates for the current block can be derived from spatially adjacent blocks. For example, spatial merge candidates may include merge candidate A1 derived from the block to the left of the current block, merge candidate B1 derived from the block at the top of the current block, merge candidate A0 derived from the block to the lower left of the current block, merge candidate B0 derived from the block to the upper right of the current block, or merge candidate B2 derived from the block to the upper left of the current block. Spatial merge candidates can be searched in a predetermined order. For example, the search order for spatial merge candidates may be A1, B1, B0, A0, and B2. In this case, B2 can be included in the merge candidate list only if there is no block corresponding to A1, B1, B0, or A0, or if the block corresponding to A1, B1, B0, or A0 is unavailable. For example, if the block corresponding to A1, B1, B0, or A0 is encoded in intra-frame prediction, it can be determined that the block is unavailable. Alternatively, if the number of spatial and temporal merge candidates included in the merge candidate list is less than or equal to the maximum number of merge candidates, B2 can be added to the merge candidate list as the next in order of temporal merge candidates.

[0163] To derive time merge candidates, a collated picture (col_pic) can be selected from the list of reference pictures. The collated picture can be the picture in the list that has the smallest Picture Order Count (POC) difference from the current picture, or a picture specified by the reference picture index. Time merge candidates can be derived based on the co-occurring blocks of the current block within the collated block. In this case, the reference picture list information used to specify the co-occurring blocks can be encoded at the block, header, or picture level and can be transmitted via a bitstream.

[0164] Figure 12 This is a diagram showing the corresponding blocks of the current block. A corresponding block indicates a block in the juxtaposed image that corresponds to the position of the current block. For example, a corresponding block can be identified as block H, which is the lower right corner of a block in the juxtaposed image that has the same coordinates and size as the current block, or block C3, which includes the center position of that block. In this case, block C3 can be identified as a corresponding block when the position of block H is unavailable, when block H is encoded via intra-frame prediction, or when block H is located outside the LCU that includes the current block.

[0165] Alternatively, a corner block of a neighboring block in the juxtaposed image that has the same coordinates and size as the current block can be identified as a corresponding block, or a block containing coordinates within that block can be identified as a corresponding block. For example, in Figure 12 In the example shown, blocks TL, BL, or C0 can be identified as co-occurring blocks.

[0166] It is also possible to derive multiple time merge candidates for the current block from multiple co-located blocks.

[0167] Motion vectors for temporal merging can be obtained by scaling the motion vectors of co-positioned blocks in juxtaposed images. Figure 13 This is a diagram illustrating an example of obtaining motion vectors for temporal merging candidates by scaling the motion vectors of co-location blocks. Motion vectors for temporal merging candidates can be obtained by scaling the motion vectors of co-location blocks using at least one of the temporal distance tb between the current image and the reference image of the current block, and the temporal distance td between the juxtaposed image and the reference image of the co-location block.

[0168] Merging candidates can be derived based on blocks having a predetermined shape or blocks having a size equal to or greater than a predetermined size. Therefore, if the current block does not have a predetermined shape or if the current block's size is smaller than a predetermined size, merging candidates for the current block can be derived based on blocks that include the current block's predetermined shape or blocks that include the current block's predetermined size or larger. For example, merging candidates for non-square-shaped coding units can be derived based on square-shaped coding units that include non-square-shaped coding units.

[0169] Figure 14 This is a diagram illustrating an example of deriving merge candidates for non-square blocks based on square blocks.

[0170] Merge candidates for non-square blocks can be derived based on square blocks that include non-square blocks. For example, in Figure 14 In the example shown, merge candidates for non-square coded block 0 and non-square coded block 1 can be derived based on square blocks. Therefore, coded block 0 and coded block 1 can use at least one of the spatial merge candidates A0, A1, A2, A3 and A4 derived based on square blocks.

[0171] Although not shown in the figure, time merging candidates for non-square blocks can also be derived based on square blocks. For example, coded block 0 and coded block 1 can use time merging candidates derived from time-adjacent blocks determined based on square blocks.

[0172] Alternatively, at least one of the spatial merge candidate and the temporal merge candidate can be derived based on square blocks, and the other can be derived based on non-square blocks. For example, coded blocks 0 and coded blocks 1 can use the same spatial merge candidate derived based on square blocks, while coded blocks 0 and coded blocks 1 can use different temporal merge candidates derived by their respective positions.

[0173] The example above illustrates deriving merge candidates based on square blocks, but merge candidates can also be derived based on non-square blocks of a predetermined shape. For example, if the current block is a non-square block of shape 2N×n (where n is 1 / 2N), then merge candidates for the current block can be derived based on non-square blocks of shape 2N×N, and if the current block is a non-square block of shape nx2N, then merge candidates for the current block can be derived based on non-square blocks of shape Nx2N.

[0174] Information indicating the shape or size of the block that will form the basis for deriving merge candidates can be transmitted via a bitstream signal. For example, block shape information indicating whether it is square or non-square can be transmitted via a bitstream signal. Alternatively, the encoder / decoder can derive merge candidates according to predefined rules, such as blocks having a predefined shape or blocks having a size equal to or greater than a predefined size.

[0175] In another example, merge candidates can be derived based on quadtree partitioning units. Here, a quadtree partitioning unit can represent a block unit partitioned by a quadtree. For example, if the current block is partitioned by a binary tree, merge candidates for the current block can be derived based on the parent block partitioned by the quadtree. If there is no parent node partitioned by the quadtree for the current block, merge candidates for the current block can be derived based on the LCU that includes the current block or a block of a specific size.

[0176] Figure 15 This is a diagram used to illustrate an example of merging candidate binary tree partitions derived from the upper node block.

[0177] Non-square binary tree partition 0 and non-square binary tree partition 1 can use at least one of the spatial merge candidates A0, A1, A2, A3 and A4 derived from the upper block based on quadtree units. Therefore, block 0 and block 1 can use the same spatial merge candidate.

[0178] Furthermore, non-square binary tree partitioning blocks 2, 3, and 4 can use at least one of B0, B1, B2, B3, and B4 derived from the upper blocks based on quadtree units. Therefore, blocks 2, 3, and 4 can use the same spatial merging candidates.

[0179] Although not shown in the figure, time merging candidates for binary tree partitioning blocks can be derived from the upper blocks based on the quadtree. Therefore, blocks 0 and 1 can use the same time merging candidates derived from the time-adjacent blocks determined by the quadtree block units. Blocks 2, 3, and 4 can also use the same time merging candidates derived from the time-adjacent blocks determined by the quadtree block units.

[0180] Alternatively, at least one of spatial merge candidates and temporal merge candidates can be derived based on binary tree block units, and the other can be derived based on quadtree block units. For example, block 0 and block 1 can use the same spatial merge candidate derived based on quadtree block units, but can use different temporal merge candidates derived based on their respective positions.

[0181] Information indicating whether the merging candidate is derived based on a quadtree partitioning unit or a binary tree partitioning unit can be transmitted via a bitstream signal. Based on this information, it can be determined whether to derive the merging candidate for the binary tree partitioning block based on the upper node block of the quadtree partition. Alternatively, the encoder / decoder can derive the merging candidate based on either a quadtree partitioning unit or a binary tree partitioning unit according to predefined rules.

[0182] As described above, merge candidates for the current block can be derived on a block or predefined unit basis (e.g., on a coding block or prediction block basis). In this case, if any spatial merge candidate for the current block exists in a predetermined region, it can be determined that it is unavailable and can then be excluded from the spatial merge candidate pool. For example, if a parallel processing region is defined for parallel processing between blocks, merge candidates included in the parallel processing region from the spatial merge candidates of the current block can be determined as unavailable. The parallel processing region can be called the merge estimation region (MER). Blocks within a parallel processing region have the advantage of being able to be merged in parallel.

[0183] The merged estimation region can be square or non-square. Non-square merged estimation regions can be limited to a predetermined shape. For example, a non-square merged estimation region can be 2N×N or N×2N in shape.

[0184] Information indicating the shape or size of the merged estimation region can be transmitted via a bitstream signal. For example, information related to the shape or size of the merged estimation region can be transmitted via a sequence header, picture parameters, or sequence parameters.

[0185] Information indicating the shape of the merged estimated region can be a 1-bit flag. For example, the syntax "isrectagular_mer_flag" indicating whether the merged estimated region has a square shape or a non-square shape can be sent as a signal via a bitstream. If the value of isrectagular_mer_flag is 1, it indicates that the merged estimated region has a non-square shape, and if the value of isrectagular_mer_flag is 0, it indicates that the merged estimated region has a square shape.

[0186] If the merged estimation region has a non-square shape, at least one of the information related to width, height, or the ratio between width and height can be transmitted via a bitstream signal. Based on this information, the size and / or shape of the non-square merged estimation region can be derived.

[0187] Figure 16 This is a diagram illustrating an example of determining the availability of spatial merging candidates based on the merging estimation region.

[0188] If the merge estimation region has an N×2N shape and a predetermined size, then spatial merge candidates B0 and B3, which are included in the same merge estimation region as block 1, cannot be used as spatial merge candidates for block 1. Therefore, the spatial merge candidates for block 1 can consist of at least one of B1, B2, and B4, other than merge candidates B0 and B3.

[0189] Similarly, spatial merge candidate C0, which is included in the same merge estimation region as block 3, cannot be used as a spatial merge candidate for block 3. Therefore, the spatial merge candidates for block 3 can consist of at least one of C1, C2, C3, and C4, other than merge candidate C0.

[0190] The above implementation has been mainly described in terms of decoding processing; however, encoding processing can be performed in the same order as described or in the reverse order.

[0191] Figure 17 This is a flowchart illustrating the process of obtaining residual samples according to an embodiment of the present invention.

[0192] First, the residual coefficients of the current block can be obtained (S1710). The decoder can obtain the residual coefficients through a coefficient scan method. For example, the decoder can perform a coefficient scan using a jig-zag scan, an up-write scan, a vertical scan, or a horizontal scan, and the decoder can obtain the residual coefficients in the form of a two-dimensional block.

[0193] Inverse quantization can be performed on the residual coefficients of the current block (S1720).

[0194] It can be determined whether to skip the inverse transform of the dequantized residual coefficients of the current block (S1730). Specifically, the decoder can determine whether to skip the inverse transform in at least one direction, either horizontal or vertical, of the current block. When it is determined that an inverse transform is to be applied in at least one direction, either horizontal or vertical, the residual samples of the current block can be obtained by performing an inverse transform on the dequantized residual coefficients of the current block. Here, at least one of DCT, DST, and KLT can be used to perform the inverse transform.

[0195] When the inverse transform is skipped in both the horizontal and vertical directions of the current block, the inverse transform is not performed in the horizontal and vertical directions of the current block. In this case, the residual sample of the current block can be obtained by scaling the inverse quantization residual coefficients with a predetermined value.

[0196] Skipping the inverse transform in the horizontal direction means performing the inverse transform in the vertical direction instead of the horizontal one. In this case, scaling can be performed in the horizontal direction.

[0197] Skipping the inverse vertical transformation means performing the inverse transformation in the horizontal direction instead of the vertical one. In this case, scaling can be performed in the vertical direction.

[0198] Whether an inverse transform skipping technique can be used for the current block can be determined based on the type of partitioning of the current block. For example, if the current block is generated through a binary tree-based partition, the inverse transform skipping scheme can be restricted for the current block. Therefore, when the current block is generated through a binary tree-based partition, the residual sample of the current block can be obtained by performing an inverse transform on the current block. Furthermore, when the current block is generated through a binary tree-based partition, the encoding / decoding of information indicating whether to skip the inverse transform (e.g., transform_skip_flag) can be omitted.

[0199] Alternatively, when generating the current block via binary tree-based partitioning, the inverse transform skipping scheme can be restricted to at least one of the horizontal or vertical directions. Here, the direction in which the inverse transform skipping scheme is restricted can be determined based on information decoded from the bitstream, or adaptively based on at least one of the current block size, the current block shape, or the current block's intra-frame prediction mode.

[0200] For example, when the current block is a non-square block with a width greater than its height, inverse transformation skipping schemes can be allowed only in the vertical direction and restricted in the horizontal direction. That is, when the current block is 2N×N, the inverse transformation is performed in the horizontal direction of the current block, and the inverse transformation can be selectively performed in the vertical direction.

[0201] On the other hand, when the current block is a non-square block with a height greater than its width, inverse transformation skipping schemes can be allowed only in the horizontal direction and restricted in the vertical direction. That is, when the current block is N×2N, the inverse transformation is performed in the vertical direction of the current block, and the inverse transformation can be selectively performed in the horizontal direction.

[0202] In contrast to the example above, when the current block is a non-square block with a width greater than its height, the inverse transformation skipping scheme can be allowed only in the horizontal direction, and when the current block is a non-square block with a height greater than its width, the inverse transformation skipping scheme can be allowed only in the vertical direction.

[0203] Information indicating whether to skip the inverse transform in the horizontal direction or the inverse transform in the vertical direction can be transmitted via a bitstream signal. For example, the information indicating whether to skip the inverse transform in the horizontal direction is a 1-bit flag “hor_transform_skip_flag”, and the information indicating whether to skip the inverse transform in the vertical direction is a 1-bit flag “ver_transform_skip_flag”. The encoder can encode at least one of “hor_transform_skip_flag” or “ver_transform_skip_flag” based on the shape of the current block. Furthermore, the decoder can determine whether to skip the inverse transform in the horizontal or vertical direction by using at least one of “hor_transform_skip_flag” or “ver_transform_skip_flag”.

[0204] It can be configured to skip the inverse transformation in either direction for the current block based on the partitioning type of the current block. For example, if the current block is generated by a binary tree-based partition, the inverse transformation in either the horizontal or vertical direction can be skipped. That is, if the current block is generated by a binary tree-based partition, it can be determined whether to skip the inverse transformation in at least one of the horizontal or vertical directions without encoding / decoding information indicating whether to skip the inverse transformation of the current block (e.g., transform_skip_flag, hor_transform_skip_flag, ver_transform_skip_flag).

[0205] According to embodiments of this disclosure, the following can also be described:

[0206] Solution 1. A method for decoding video, the method comprising:

[0207] Export the space merge candidates for the current block;

[0208] A list of merging candidates for the current block is generated based on the spatial merging candidates;

[0209] The motion information of the current block is obtained based on the merge candidate list; and

[0210] The motion information is used to perform motion compensation on the current block.

[0211] If the current block does not have a predefined shape or does not have a size equal to or greater than the predefined size, then space merging candidates for the current block are derived based on blocks having the predefined shape or having a size equal to or greater than the predefined size, and the blocks include the current block.

[0212] Option 2. According to the method described in Option 1, wherein the predefined shape is a square shape.

[0213] Option 3. According to the method of Option 2, wherein the current block has the same spatial merging candidate as the adjacent blocks included in the block having the square shape of the current block.

[0214] Option 4. According to the method of Option 1, wherein if the current block and the spatial merge candidate are included in the same merge estimation region, the spatial merge candidate is determined to be unavailable.

[0215] Option 5. The method according to Option 4, wherein the merged estimation region has a square shape or a non-square shape.

[0216] Scheme 6. According to the method of Scheme 5, if the merge estimation region has a non-square shape, the number of candidate shapes that the merge estimation region can have for merging is limited to a predefined number.

[0217] Option 7. A method for encoding video, the method comprising:

[0218] Export the space merge candidates for the current block;

[0219] A list of merging candidates for the current block is generated based on the spatial merging candidates;

[0220] The motion information of the current block is obtained based on the merge candidate list; and

[0221] The motion information is used to perform motion compensation on the current block.

[0222] If the current block does not have a predefined shape or does not have a size equal to or greater than the predefined size, then space merging candidates for the current block are derived based on blocks having the predefined shape or having a size equal to or greater than the predefined size, and the blocks include the current block.

[0223] Option 8. The method according to Option 7, wherein the predefined shape is a square shape.

[0224] Option 9. The method according to Option 8, wherein the current block has the same spatial merging candidate as the adjacent blocks included in the block having the square shape of the current block.

[0225] Option 10. The method according to Option 7, wherein if the current block and the spatial merge candidate are included in the same merge estimation region, the spatial merge candidate is determined to be unavailable.

[0226] Option 11. The method according to Option 10, wherein the merged estimation region has a square shape or a non-square shape.

[0227] Scheme 12. The method according to Scheme 11, wherein if the merge estimation region has a non-square shape, the number of candidate shapes that the merge estimation region can have for merging is limited to a predefined number.

[0228] Option 13. An apparatus for decoding video, the apparatus comprising:

[0229] The prediction unit is configured to: derive spatial merging candidates for the current block; generate a merging candidate list for the current block based on the spatial merging candidates; obtain motion information of the current block based on the merging candidate list; and perform motion compensation on the current block using the motion information.

[0230] If the current block does not have a predefined shape or does not have a size equal to or greater than the predefined size, then space merging candidates for the current block are derived based on blocks having the predefined shape or having a size equal to or greater than the predefined size, and the blocks include the current block.

[0231] Although the above embodiments have been described based on a series of steps or flowcharts, they do not limit the timing of the invention and can be executed simultaneously or in different orders as needed. Furthermore, each of the components (e.g., units, modules, etc.) constituting the block diagrams in the above embodiments can be implemented by hardware devices or software and multiple components. Alternatively, multiple components can be combined and implemented by a single hardware device or software. The above embodiments can be implemented in the form of program instructions that can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium can include one or a combination of program commands, data files, data structures, etc. Examples of computer-readable media include: magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical recording media, such as CD-ROMs and DVDs; magneto-optical media, such as optical-magnetic floppy disks; media; and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, flash memory, etc. The hardware device can be configured to operate as one or more software modules for performing the processing according to the invention, and vice versa.

[0232] Industrial application

[0233] This invention can be applied to electronic devices capable of encoding / decoding video.

Claims

1. A method for decoding video, the method comprising: The current block is determined based on tree-based block partitioning; Generate a merge candidate list for the current block, the merge candidate list including spatial merge candidates and temporal merge candidates; The motion information of the current block is obtained based on the merge candidate list; as well as Based on the motion information, a prediction sample for the current block is obtained. Specifically, the current block is asymmetrically divided into two partitions based on either a vertical line or a horizontal line. The two partitions share the merge candidate list used for inter-frame prediction of the current block. Even if the current block has a non-square shape in which one of its width and height is greater than the other, the merge candidate list is allowed to be shared between the two partitions if the size of the current block is greater than or equal to a predefined size.

2. The method according to claim 1, wherein, The current block is not square, and In response to the case where the current block is divided into two partitions based on the vertical line, the first partition of the two partitions has a smaller width than the second partition of the two partitions.

3. The method according to claim 2, wherein, In response to the first partition being the left partition of the two partitions, the merge candidate list for the first partition includes multiple spatial merge candidates, and The plurality of spatial merge candidates includes a first spatial merge candidate derived from the top adjacent block that is adjacent to the second partition but not adjacent to the first partition.

4. The method according to claim 3, wherein, The plurality of spatial merge candidates also includes a second spatial merge candidate derived from the left adjacent block that is adjacent to the first partition but not adjacent to the second partition.

5. The method according to claim 1, wherein, The current block is not square, and In response to the case where the current block is divided into two partitions based on the horizontal line, the first partition of the two partitions has a smaller height than the second partition of the two partitions.

6. The method according to claim 5, wherein, In response to the first partition being the top partition of the two partitions, the merge candidate list for the first partition includes multiple space merge candidates, and Among them, the plurality of spatial merge candidates includes a first spatial merge candidate derived from the lower left adjacent block that is adjacent to the second partition but not adjacent to the first partition.

7. The method according to claim 6, wherein, The plurality of spatial merge candidates also includes a second spatial merge candidate derived from the upper right adjacent block that is adjacent to the first partition but not adjacent to the second partition.

8. The method according to claim 1, wherein, The current block is determined based on tree-based block partitioning, including: The first depth coding block is divided into two second depth coding blocks; Determine whether to divide the second-depth coding block into two third-depth coding blocks; and If it is determined that the second depth-coded block needs to be divided, the second depth-coded block is divided into two third depth-coded blocks, and the current block is one of the two third depth-coded blocks. The first depth coding block is divided in either the horizontal or vertical direction, and In response to the first depth coding block having a size of 128x128, the second depth coding block is not allowed to be divided in the same direction as the first depth coding block.

9. A method for encoding video, the method comprising: The current block is determined based on tree-based block partitioning; Generate a merge candidate list for the current block, the merge candidate list including spatial merge candidates and temporal merge candidates; The motion information of the current block is obtained based on the merge candidate list; as well as Based on the motion information, a prediction sample for the current block is obtained. Specifically, the current block is asymmetrically divided into two partitions based on either a vertical line or a horizontal line. The two partitions share the merge candidate list used for inter-frame prediction of the current block. Even if the current block has a non-square shape in which one of its width and height is greater than the other, the merge candidate list is allowed to be shared between the two partitions if the size of the current block is greater than or equal to a predefined size.

10. A method for transmitting video data, comprising: Obtain the bitstream of the video data, wherein the bitstream is generated by performing the following operations: The current block is determined based on tree-based block partitioning. Generate a merge candidate list for the current block, the merge candidate list including spatial merge candidates and temporal merge candidates. The motion information of the current block is obtained based on the merge candidate list. Based on the motion information, a prediction sample for the current block is obtained, and The current block is encoded based on the predicted samples; and Transmit the bit stream, Specifically, the current block is asymmetrically divided into two partitions based on either a vertical line or a horizontal line. The two partitions share the merge candidate list used for inter-frame prediction of the current block. Even if the current block has a non-square shape in which one of its width and height is greater than the other, the merge candidate list is allowed to be shared between the two partitions if the size of the current block is greater than or equal to a predefined size.

Citation Information

Patent Citations

  • Inter prediction method and apparatus therefor

    WO2013036071A2