Method for decoding and encoding video, and apparatus for transmitting compressed video data
By using multiple merge candidate lists to perform motion compensation during the video signal encoding/decoding process, the cost increase caused by the large amount of data in high-resolution video signal processing is solved, and more efficient inter-frame prediction and data compression are achieved.
Patent Information
- Application Number
- CN202510424833.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-23
- Filing Date
- 2019-05-23
- Publication Date
- 2025-05-27
AI Technical Summary
When processing high-resolution and high-quality video signals, the prior art faces the problem of increasing transmission and storage costs caused by large data volume.
Motion compensation is performed by using multiple merge candidate lists during the video signal encoding/decoding process, and the merge index is effectively encoded/decoded.
The efficiency of inter-frame prediction is improved, the amount of data is reduced, and the cost of transmission and storage is reduced.
Smart Images

Figure CN120050433A_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application No. 201980034764.3, whose international application date is May 23, 2019, international application number is PCT / KR2019 / 006216, and which entered the Chinese national phase on November 23, 2020, and whose invention name is “Method and device for processing video signals”. Technical Field
[0002] The invention relates to a method and an apparatus for processing a video signal. Background Art
[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, has increased in various application fields. However, compared with conventional image data, higher resolution and quality image data has an increased data volume. Therefore, when the image data is transmitted by using a medium such as a conventional wired and wireless broadband network, or when the image data is stored by using a conventional storage medium, the cost of transmission and storage increases. In order to solve these problems arising as the resolution and quality of image data increase, efficient image encoding / decoding technology can be used.
[0004] The image compression technology includes various technologies, including: an inter-frame prediction technology that predicts a pixel value included in a current picture based on a previous picture or a subsequent picture of the current picture; an intra-frame prediction technology that predicts a pixel value included in the current picture by using pixel information in the current picture; an entropy coding technology that assigns a short code to a value with a high frequency of occurrence and a long code to a value with a low frequency of occurrence, etc. Image data can be effectively compressed by using such an image compression technology and can be transmitted or stored.
[0005] Meanwhile, along with the demand for high-resolution images, the demand for stereoscopic image content as a new image service is also increasing. Video compression technology for efficiently providing stereoscopic image content with high resolution and ultra-high resolution is being studied. Summary of the invention
[0006] Technical issues
[0007] The present invention provides a method and apparatus for efficiently performing inter-frame prediction on an encoding / decoding target block when encoding / decoding a video signal.
[0008] The present invention provides a method and apparatus for performing motion compensation by using a plurality of merge candidate lists when encoding / decoding a video signal.
[0009] The present invention provides a method and apparatus for efficiently encoding / decoding a merge index when encoding / decoding a video signal.
[0010] Technical problems that can be obtained from the present invention are not limited to the above-mentioned technical tasks, and a person of ordinary skill in the technical field to which the present invention belongs can clearly understand other unmentioned technical tasks according to the following description.
[0011] Technical Solution
[0012] The video signal decoding method and apparatus according to the present invention can derive a merge candidate based on a neighboring block adjacent to a current block, generate a first merge candidate list including the merge candidate, decode information for specifying one of the merge candidates included in the first merge candidate list, and derive motion information of the current block based on the merge candidate assigned an index determined by the information. In this case, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, the merge candidate included in the second merge candidate list can be added to the first merge candidate list.
[0013] The video signal encoding method and apparatus according to the present invention can derive a merge candidate based on a neighboring block adjacent to a current block, generate a first merge candidate list including the merge candidate, encode information for specifying one of the merge candidates included in the first merge candidate list, and derive motion information of the current block based on the merge candidate assigned an index determined by the information. In this case, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, the merge candidate included in the second merge candidate list can be added to the first merge candidate list.
[0014] For the video signal encoding / decoding method and apparatus according to the present invention, the information may include an index prefix and an index suffix.
[0015] For the video signal encoding / decoding method and apparatus according to the present invention, when the value of the index prefix is less than a threshold value, the index may be set to be the same as the index prefix.
[0016] For the video signal encoding / decoding method and apparatus according to the present invention, when the value of the index prefix is greater than a threshold, the index may be derived by adding the value of the index suffix to a value derived based on the index prefix.
[0017] For the video signal encoding / decoding method and apparatus according to the present invention, the threshold value may be determined based on the number of merging candidates included in the first merging candidate list.
[0018] For the video signal encoding / decoding method and apparatus according to the present invention, the second merge candidate list may include merge candidates derived from blocks that are not adjacent to the current block.
[0019] For the video signal encoding / decoding method and apparatus according to the present invention, the non-adjacent blocks may be on the same line as the blocks adjacent to the current block.
[0020] The method for decoding a video according to the present invention includes: deriving at least one spatial merge candidate of a current coding block, the spatial merge candidate being derived from a neighboring block adjacent to the current coding block; deriving a temporal merge candidate of the current coding block; generating a first candidate list including at least one spatial merge candidate and a temporal merge candidate of the current coding block, wherein when the number of merge candidates included in the first candidate list is less than a predetermined value, adding a motion information candidate included in a second candidate list to the first candidate list; selecting two merge candidates from a plurality of merge candidates included in the first candidate list for two triangular-shaped partitions in the current coding block based on two index information; obtaining a prediction block of the current coding block based on two motion information derived from the two merge candidates; and reconstructing the current coding block based on the prediction block. The motion information candidate included in the second candidate list is derived from a spatial block, which is decoded before decoding the current coding block and is included in a current picture including the current coding block. The first index information in the two index information specifies a first merge candidate among a plurality of merge candidates. The second index information in the two index information specifies a second merge candidate among the remaining merge candidates except the first merge candidate. Both the first index information and the second index information consist of only a prefix portion.
[0021] The method for encoding a video according to the present invention includes: deriving at least one spatial merge candidate of a current coding block, the spatial merge candidate being derived from a neighboring block adjacent to the current coding block; deriving a temporal merge candidate of the current coding block; generating a first candidate list including at least one spatial merge candidate and a temporal merge candidate of the current coding block, wherein when the number of merge candidates included in the first candidate list is less than a predetermined value, adding a motion information candidate included in a second candidate list to the first candidate list; selecting two merge candidates from a plurality of merge candidates included in the first candidate list for two triangular-shaped partitions in the current coding block; obtaining a prediction block of the current coding block based on two motion information derived from the two merge candidates; obtaining a residual block of the current coding block based on the prediction block; and encoding two index information. The motion information candidate included in the second candidate list is derived from a spatial block, which is encoded before encoding the current coding block and is included in a current picture including the current coding block. The first index information in the two index information specifies a first merge candidate among a plurality of merge candidates. The second index information in the two index information specifies a second merge candidate among the remaining merge candidates except the first merge candidate. Both the first index information and the second index information are encoded by only the prefix portion.
[0022] The device for transmitting compressed video data according to the present invention includes one or more hardware units, which are configured to: obtain compressed video data; and transmit compressed video data. Obtaining compressed video data includes: deriving at least one spatial merge candidate of the current coding block, the spatial merge candidate is derived from the neighboring blocks adjacent to the current coding block; deriving the temporal merge candidate of the current coding block; generating a first candidate list including at least one spatial merge candidate and a temporal merge candidate of the current coding block, wherein when the number of merge candidates included in the first candidate list is less than a predetermined value, the motion information candidate included in the second candidate list is added to the first candidate list; selecting two merge candidates from the multiple merge candidates included in the first candidate list for the division of two triangular shapes in the current coding block; obtaining a prediction block of the current coding block based on two motion information derived from the two merge candidates; obtaining a residual block of the current coding block based on the prediction block; and encoding two index information. The motion information candidate included in the second candidate list is derived from a spatial block, which is encoded before encoding the current coding block and included in the current picture including the current coding block. The first index information of the two index information specifies the first merge candidate among the multiple merge candidates. The second index information of the two index information specifies the second merge candidate of the remaining merge candidates except the first merge candidate. Both the first index information and the second index information are encoded only by the prefix part.
[0023] It should be understood that the above-summarized features are exemplary aspects of the following detailed description of the invention and do not limit the scope of the invention.
[0024] Beneficial Effects
[0025] According to the present invention, the efficiency of inter prediction can be improved by performing motion compensation by using a plurality of merge candidate lists.
[0026] According to the present invention, the efficiency of inter-frame prediction can be improved by obtaining motion information based on a plurality of merge candidates.
[0027] According to the present invention, an efficient encoding / decoding method of a merged index can be provided.
[0028] Effects that can be obtained from the present invention may not be limited to the above-mentioned effects, and other unmentioned effects may be clearly understood by a person of ordinary skill in the technical field to which the present invention pertains from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a block diagram showing an apparatus for encoding a video according to an embodiment of the present invention.
[0030] Figure 2 is a block diagram showing an apparatus for decoding a video according to an embodiment of the present invention.
[0031] Figure 3 is a diagram showing partition mode candidates that can be applied to a coding block when the coding block is encoded by inter prediction.
[0032] Figure 4 An example of hierarchical division of coding blocks based on a tree structure is shown as an embodiment to which the present invention is applied.
[0033] Figure 5 1 is a diagram showing a partition shape that allows binary tree-based partitioning as an embodiment to which the present invention is applied.
[0034] Figure 6 A ternary tree partition shape is shown.
[0035] Figure 7 is a diagram showing an example of binary tree-based partitioning that allows only specific shapes.
[0036] Figure 8 is a diagram for describing an example in which information related to the number of times binary tree division is allowed is encoded / decoded according to an embodiment to which the present invention is applied.
[0037] Fig. 9 : is a flowchart showing an inter-frame prediction method as an embodiment to which the present invention is applied.
[0038] Fig.10 is a diagram illustrating a process of deriving motion information of a current block when a merge mode is applied to the current block.
[0039] Fig.11 is a diagram showing an example of spatially neighboring blocks.
[0040] Fig.12 is a diagram showing an example of deriving a motion vector of a temporal merging candidate.
[0041] Fig.13 is a diagram showing the positions of candidate blocks that can be used as co-located blocks.
[0042] Fig.14 is a diagram illustrating a process of deriving motion information of a current block when the AMVP mode is applied to the current block.
[0043] Fig.15 is a diagram showing an example of deriving a merge candidate from a second merge candidate block when a first merge candidate block is unavailable.
[0044] Fig.16is a diagram showing an example of deriving a merge candidate from a second merge candidate block located on the same line as a first merge candidate block.
[0045] Figures 17 to 20 is a diagram showing the order of searching for merge candidate blocks.
[0046] Fig.21 is a diagram showing an example of deriving merging candidates of non-square blocks based on square blocks.
[0047] Fig. 22 is a diagram showing an example of deriving merge candidates based on high-level node blocks.
[0048] Fig.23 is a diagram showing an example of determining the availability of spatially neighboring blocks based on a merged estimated region.
[0049] Fig.24 is a diagram showing an example of deriving a merge candidate based on a merge estimation region. DETAILED DESCRIPTION
[0050] Various modifications can be made to the present invention, and there are various embodiments of the present invention, examples of which will now be provided with reference to the accompanying drawings, and examples of which will be described in detail. However, the present invention is not limited thereto, and the exemplary embodiments may be interpreted as including all modifications, equivalents or alternatives within the technical concept and technical scope of the present invention. In the accompanying drawings described, similar reference numerals refer to similar elements.
[0051] The terms 'first', 'second', etc. used in the specification may be used to describe various components, but these components should not be interpreted as being limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the present invention, a 'first' component may be referred to as a 'second' component, and a 'second' component may also be similarly referred to as a 'first' component. The term 'and / or' includes a combination of a plurality of items or any one of a plurality of items.
[0052] In the present disclosure, when an element is referred to as being "connected" or "coupled" to another element, it should be understood that it includes not only that the element is directly connected or coupled to the other element, but also that another element may exist between them. When an element is referred to as being "directly connected" or "directly coupled" to another element, it should be understood that there are no other elements between them.
[0053] The terms used in this specification are only used to describe specific embodiments and are not intended to limit the present invention. Expressions used in the singular include expressions in the plural form unless the expression has a significantly different meaning in the context. In this specification, it should be understood that terms such as "including", "having" etc. are intended to indicate the existence of features, numbers, steps, actions, elements, parts or combinations thereof disclosed in this specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, parts or combinations thereof may exist or may be added.
[0054] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. Hereinafter, the same constituent elements in the drawings are denoted by the same reference numerals, and repeated description of the same elements will be omitted.
[0055] Figure 1 is a block diagram showing an apparatus for encoding a video according to an embodiment of the present invention.
[0056] Reference Figure 1 , the device 100 for encoding a video may include: a picture division module 110, prediction modules 120 and 125, a transformation module 130, a quantization module 135, a rearrangement module 160, an entropy coding module 165, an inverse quantization module 140, an inverse transformation module 145, a filter module 150 and a memory 155.
[0057] Figure 1 The components shown in are shown independently to represent the different characteristic functions in the device for encoding the video. Therefore, this does not mean that each component is composed of a separate hardware or software component unit. In other words, for convenience, each component includes each of the listed components. Therefore, at least two components in each component can be combined to form a component, or a component can be divided into multiple components to perform each function. Without departing from the essence of the present invention, the embodiment of combining each component and the embodiment of dividing a component are also included in the scope of the present invention.
[0058] In addition, some of the components may not be indispensable components for performing the basic functions of the present invention, but selective components that only improve the performance of the present invention. The present invention can be implemented by excluding components for improving performance and only including indispensable components for implementing the essence of the present invention. Structures that exclude selective components that only improve performance and only include indispensable components are also included in the scope of the present invention.
[0059] The picture partitioning module 110 may partition an input picture into one or more processing units. Here, a processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture partitioning module 110 may partition a picture into a combination of a plurality of coding units, prediction units, and transform units, and may encode a picture by selecting a combination of coding units, prediction units, and transform units using a predetermined criterion (e.g., a cost function).
[0060] For example, a picture may be divided into a plurality of coding units. A recursive tree structure such as a quadtree structure may be used to divide a picture into coding units. A coding unit divided into other coding units with a picture or a maximum coding unit as a root may be divided into child nodes corresponding to the number of divided coding units. Coding units that cannot be further divided according to predetermined restrictions are used as leaf nodes. That is, when it is assumed that only square division is feasible for a coding unit, a coding unit may be divided into up to four other coding units.
[0061] Hereinafter, in an embodiment of the present invention, a coding unit may mean a unit that performs encoding or a unit that performs decoding.
[0062] The prediction unit may be one of partitions split into a square shape or a rectangular shape having the same size in a single coding unit, or the prediction unit may be one of partitions split so as to have different shapes / sizes in a single coding unit.
[0063] When a prediction unit subject to intra prediction is generated based on a coding unit and the coding unit is not a minimum coding unit, intra prediction may be performed without splitting the coding unit into a plurality of prediction units NxN.
[0064] The prediction modules 120 and 125 may include an inter prediction module 120 that performs inter prediction and an intra prediction module 125 that performs intra prediction. It may be determined whether inter prediction or intra prediction is performed for a prediction unit, and detailed information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. Here, the processing unit subjected to prediction may be different from the processing unit for which the prediction method and detailed content are determined. For example, the prediction method, prediction mode, etc. may be determined by the prediction unit, and the prediction may be performed by the transform unit. The residual value (residual block) between the generated prediction block and the original block may be input to the transform module 130. In addition, the prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value by the entropy encoding module 165, and may be transmitted to a device for decoding a video. When a specific encoding mode is used, it may be transmitted to a device for decoding a video by encoding the original block as it is without generating a prediction block through the prediction modules 120 and 125.
[0065] The inter-frame prediction module 120 may predict a prediction unit based on information of at least one of a previous picture or a subsequent picture of the current picture, or in some cases, may predict a prediction unit based on information of some coding regions in the current picture. The inter-frame prediction module 120 may include a reference picture interpolation module, a motion prediction module, and a motion compensation module.
[0066] The reference picture interpolation module may receive reference picture information from the memory 155, and may generate pixel information of integer pixels or less than integer pixels according to the reference picture. In the case of luma pixels, an 8-tap interpolation filter based on DCT with different filter coefficients may be used to generate pixel information of integer pixels or less than integer pixels in units of 1 / 4 pixels. In the case of chrominance signals, a 4-tap interpolation filter based on DCT with different filter coefficients may be used to generate pixel information of integer pixels or less than integer pixels in units of 1 / 8 pixels.
[0067] The motion prediction module can perform motion prediction based on the reference picture interpolated by the reference picture interpolation module. Various methods such as full search based block matching algorithm (FBMA), three-step search (TSS), new three-step search algorithm (NTS), etc. can be used as methods for calculating motion vectors. Based on the interpolated pixels, the motion vector can have a motion vector value in units of 1 / 2 pixel or 1 / 4 pixel. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods such as skip method, merge method, AMVP (Advanced Motion Vector Prediction) method, intra-frame block copy method, etc. can be used as motion prediction methods.
[0068] The intra prediction module 125 may generate a prediction unit based on reference pixel information adjacent to the current block, which is pixel information in the current picture. When the adjacent block of the current prediction unit is a block subject to inter-frame prediction and thus the reference pixel is a pixel subject to inter-frame prediction, the reference pixel information of the adjacent block subject to intra-frame prediction may be used to replace the reference pixel included in the block subject to inter-frame prediction. That is, when the reference pixel is unavailable, at least one reference pixel among the available reference pixels may be used to replace the unavailable reference pixel information.
[0069] The prediction mode of the intra prediction may include a directional prediction mode using reference pixel information according to a prediction direction and a non-directional prediction mode that does not use directional information when performing prediction. The mode for predicting luminance information may be different from the mode for predicting chrominance information, and in order to predict chrominance information, the intra prediction mode information for predicting luminance information or the predicted luminance signal information may be used.
[0070] When performing intra prediction, when the size of the prediction unit is the same as the size of the transform unit, intra prediction can be performed on the prediction unit based on pixels located on the left, upper left, and upper sides of the prediction unit. However, when performing intra prediction, when the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed using reference pixels based on the transform unit. In addition, intra prediction using NxN partitioning can be used only for the minimum coding unit.
[0071] In the intra prediction method, according to the prediction mode, a prediction block can be generated after applying an AIS (Adaptive Intra Smoothing) filter to a reference pixel. The type of the AIS filter applied to the reference pixel may be different. In order to perform the intra prediction method, the intra prediction mode of the current prediction unit may be predicted according to the intra prediction mode of a prediction unit adjacent to the current prediction unit. When the prediction mode of the current prediction unit is predicted by using the mode information predicted according to the adjacent prediction unit, when the intra prediction mode of the current prediction unit is the same as the intra prediction mode of the adjacent prediction unit, predetermined flag information may be used to transmit information indicating that the prediction mode of the current prediction unit and the prediction mode of the adjacent prediction unit are the same as each other. When the prediction mode of the current prediction unit is different from the prediction mode of the adjacent prediction unit, entropy coding may be performed to encode the prediction mode information of the current block.
[0072] In addition, a residual block including information on a residual value, which is a difference between the prediction unit subjected to prediction and the original block of the prediction unit, may be generated based on the prediction unit generated by the prediction modules 120 and 125. The generated residual block may be input to the transform module 130.
[0073] The transform module 130 may transform the residual block including information on the residual value between the original block and the prediction unit generated by the prediction modules 120 and 125 by using a transform method such as discrete cosine transform (DCT), discrete sine transform (DST), and KLT. Whether to apply DCT, DST, or KLT in order to transform the residual block may be determined based on intra prediction mode information of the prediction unit used to generate the residual block.
[0074] The quantization module 135 may quantize the value transformed to the frequency domain by the transform module 130. The quantization coefficient may be different according to the block or importance of the picture. The value calculated by the quantization module 135 may be provided to the inverse quantization module 140 and the rearrangement module 160.
[0075] The rearrangement module 160 may rearrange coefficients of the quantized residual value.
[0076] The rearrangement module 160 can change the coefficient of the two-dimensional block form into the coefficient of the one-dimensional vector form by the coefficient scanning method. For example, the rearrangement module 160 can use a zigzag scanning method to scan from the DC coefficient to the coefficient of the high frequency domain so that the coefficient is changed into a one-dimensional vector form. According to the size of the transform unit and the intra-frame prediction mode, the zigzag scanning can be replaced by the vertical scanning of the coefficient of the two-dimensional block form in the column direction or the horizontal scanning of the coefficient of the two-dimensional block form in the row direction. That is to say, it can be determined according to the size of the transform unit and the intra-frame prediction mode which scanning method to use among the zigzag scanning, the vertical scanning and the horizontal scanning.
[0077] The entropy encoding module 165 may perform entropy encoding based on the value calculated by the rearrangement module 160. The entropy encoding may use various encoding methods such as exponential Golomb coding, context adaptive variable length coding (CAVLC), and context adaptive binary arithmetic coding (CABAC).
[0078] The entropy coding module 165 can encode various information from the rearrangement module 160 and the prediction modules 120 and 125, such as residual value coefficient information and block type information of the coding unit, prediction mode information, partition unit information, prediction unit information, transformation unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc.
[0079] The entropy encoding module 165 may entropy encode coefficients of the coding unit input from the rearrangement module 160 .
[0080] The inverse quantization module 140 may inversely quantize the value quantized by the quantization module 135, and the inverse transform module 145 may inversely transform the value transformed by the transform module 130. The residual values generated by the inverse quantization module 140 and the inverse transform module 145 may be combined with the prediction units predicted by the motion estimation module, the motion compensation module, and the intra prediction module of the prediction modules 120 and 125, so that a reconstructed block may be generated.
[0081] The filter module 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).
[0082] The deblocking filter can remove block distortion that occurs due to the boundaries between blocks in the reconstructed picture. In order to determine whether to perform deblocking, the pixels in several rows or columns included in the block can be the basis for determining whether to apply the deblocking filter to the current block. When the deblocking filter is applied to the block, a strong filter or a weak filter can be applied according to the required deblocking filter strength. In addition, when applying the deblocking filter, horizontal filtering and vertical filtering can be processed in parallel.
[0083] The offset correction module can correct the offset from the original picture in units of pixels in the picture subjected to deblocking. In order to perform offset correction on a specific picture, a method of applying an offset in consideration of edge information of each pixel or a method of dividing the pixels of the picture into a predetermined number of regions, determining the region to be subjected to performing the offset, and applying the offset to the determined region can be used.
[0084] Adaptive loop filtering (ALF) can be performed based on a value obtained by comparing a filtered reconstructed picture with an original picture. The pixels included in the picture can be divided into predetermined groups, a filter to be applied to each group in the group can be determined, and filtering can be performed separately for each group. Information on whether ALF is applied and a luminance signal can be transmitted by a coding unit (CU). The shape and filter coefficients of the filter used for ALF can be different for each block. In addition, regardless of the characteristics of the application target block, a filter for ALF of the same shape (fixed shape) can be applied.
[0085] The memory 155 may store the reconstructed block or the reconstructed picture calculated by the filter module 150. The stored reconstructed block or the reconstructed picture may be provided to the prediction modules 120 and 125 when performing inter prediction.
[0086] Figure 2 is a block diagram showing an apparatus for decoding a video according to an embodiment of the present invention.
[0087] Reference Figure 2The apparatus 200 for decoding a video may include an entropy decoding module 210 , a rearrangement module 215 , an inverse quantization module 220 , an inverse transform module 225 , prediction modules 230 and 235 , a filter module 240 , and a memory 245 .
[0088] When a video bitstream is input from the apparatus for encoding a video, the input bitstream may be decoded according to an inverse process of the apparatus for encoding a video.
[0089] The entropy decoding module 210 may perform entropy decoding according to an inverse process of entropy encoding performed by an entropy encoding module of a device for encoding a video. For example, corresponding to a method performed by a device for encoding a video, various methods such as exponential Golomb coding, context adaptive variable length coding (CAVLC), and context adaptive binary arithmetic coding (CABAC) may be applied.
[0090] The entropy decoding module 210 may decode information about intra prediction and inter prediction performed by the apparatus for encoding a video.
[0091] The rearrangement module 215 may perform rearrangement on the bitstream entropy-decoded by the entropy decoding module 210 based on the rearrangement method used in the apparatus for encoding a video. The rearrangement module may reconstruct and rearrange coefficients in the form of a one-dimensional vector into coefficients in the form of a two-dimensional block. The rearrangement module 215 may receive information related to coefficient scanning performed in the apparatus for encoding a video, and may perform rearrangement via a method of inversely scanning the coefficients based on the scanning order performed in the apparatus for encoding a video.
[0092] The inverse quantization module 220 may perform inverse quantization based on a quantization parameter received from the apparatus for encoding a video and coefficients of the rearranged block.
[0093] The inverse transform module 225 may perform an inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the quantization result of the apparatus for encoding the video, which is an inverse process of the transform, i.e., DCT, DST, and KLT, performed by the transform module. The inverse transform may be performed based on a transmission unit determined by the apparatus for encoding the video. The inverse transform module 225 of the apparatus for decoding the video may selectively perform a transform scheme (e.g., DCT, DST, and KLT) according to multiple pieces of information such as a prediction method, a size of a current block, a prediction direction, and the like.
[0094] The prediction modules 230 and 235 may generate a prediction block based on the information on prediction block generation received from the entropy decoding module 210 and previously decoded block or picture information received from the memory 245 .
[0095] As described above, similar to the operation of the apparatus for encoding a video, when performing intra prediction, when the size of the prediction unit is the same as the size of the transform unit, intra prediction can be performed on the prediction unit based on pixels located on the left, upper left, and upper sides of the prediction unit. When performing intra prediction, when the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed using reference pixels based on the transform unit. In addition, intra prediction using NxN partitioning can be used only for the minimum coding unit.
[0096] The prediction modules 230 and 235 may include a prediction unit determination module, an inter-frame prediction module, and an intra-frame prediction module. The prediction unit determination module may receive various information such as prediction unit information, prediction mode information of an intra-frame prediction method, information about motion prediction of an inter-frame prediction method, etc. from the entropy decoding module 210, may divide the current coding unit into prediction units, and may determine whether to perform inter-frame prediction or intra-frame prediction on the prediction unit. By using the information required in the inter-frame prediction of the current prediction unit received from the device for encoding a video, the inter-frame prediction module 230 may perform inter-frame prediction on the current prediction unit based on information of at least one of a previous picture or a subsequent picture of the current picture including the current prediction unit. Alternatively, inter-frame prediction may be performed based on information of some pre-reconstructed areas in the current picture including the current prediction unit.
[0097] In order to perform inter prediction, which mode of the skip mode, the merge mode, the AMVP mode, and the inter block copy mode to use as a motion prediction method of a prediction unit included in the coding unit may be determined for the coding unit.
[0098] The intra prediction module 235 can generate a prediction block based on pixel information in the current picture. When the prediction unit is a prediction unit subjected to intra prediction, intra prediction can be performed based on intra prediction mode information of the prediction unit received from the device for encoding the video. The intra prediction module 235 may include an adaptive intra smoothing (AIS) filter, a reference pixel interpolation module, and a DC filter. The AIS filter performs filtering on the reference pixels of the current block, and it can be determined whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block by using the AIS filter information received from the device for encoding the video and the prediction mode of the prediction unit. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.
[0099] When the prediction mode of the prediction unit is a prediction mode in which intra prediction is performed based on pixel values obtained by interpolating reference pixels, the reference pixel interpolation module may interpolate the reference pixels to generate reference pixels that are integer pixels or smaller than integer pixels. When the prediction mode of the current prediction unit is a prediction mode in which a prediction block is generated without interpolating reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is a DC mode, a DC filter may generate a prediction block by filtering.
[0100] The reconstructed block or picture may be provided to the filter module 240. The filter module 240 may include a deblocking filter, an offset correction module, and an ALF.
[0101] Information about whether to apply a deblocking filter to a corresponding block or picture and information about which filter of a strong filter and a weak filter to apply when applying the deblocking filter can be received from a device for encoding a video. The deblocking filter of a device for decoding a video can receive information about the deblocking filter from a device for encoding a video, and can perform deblocking filtering on the corresponding block.
[0102] The offset correction module may perform offset correction on the reconstructed picture based on the type of offset correction applied to the picture when encoding is performed and offset value information.
[0103] The ALF may be applied to the coding unit based on information on whether the ALF is applied, ALF coefficient information, etc. received from the apparatus for encoding a video. The ALF information may be provided to be included in a specific parameter set.
[0104] The memory 245 may store the reconstructed picture or block to be used as a reference picture or block, and may provide the reconstructed picture to the output module.
[0105] As described above, in the embodiments of the present invention, for convenience of explanation, a coding unit is used as a term indicating a unit for encoding, but the coding unit may be used as a unit that performs decoding as well as encoding.
[0106] In addition, the current block may represent a target block to be encoded / decoded. And, according to the encoding / decoding step, the current block may represent a coding tree block (or coding tree unit), a coding block (or coding unit), a transform block (or transform unit), a prediction block (or prediction unit), etc. In the present specification, a 'unit' may represent a basic unit for performing a specific encoding / decoding process, and a 'block' may represent a sample array of a predetermined size. Unless otherwise specified, 'block' and 'unit' may be used as the same meaning. For example, in the examples mentioned later, it can be understood that the coding block and the coding unit have the same meaning as each other.
[0107] A picture can be encoded / decoded by being divided into basic blocks having a square shape or a non-square shape. At this time, the basic block can be referred to as a coding tree unit. The coding tree unit can be defined as a coding unit of the maximum size allowed within a sequence or slice. Information indicating whether the coding tree unit has a square shape or a non-square shape or information related to the size of the coding tree unit can be signaled through a sequence parameter set, a picture parameter set, or a slice header. The coding tree unit can be divided into smaller-sized partitions. At this time, if it is assumed that the depth of the partition generated by splitting the coding tree unit is 1, the depth of the partition generated by splitting the partition with a depth of 1 can be defined as 2. That is, the partition generated by splitting the partition with a depth of k in the coding tree unit can be defined as a depth of k+1.
[0108] A partition of any size generated by splitting a coding tree unit may be defined as a coding unit. A coding unit may be recursively split or divided into a basic unit for performing prediction, quantization, transformation, or loop filtering, etc. For example, a partition of any size generated by splitting a coding unit may be defined as a coding unit, or may be defined as a transformation unit or a prediction unit, which is a basic unit for performing prediction, quantization, transformation, or loop filtering, etc.
[0109] Alternatively, a prediction block having the same size as or smaller than the coding block can be determined by predictive partitioning of the coding block. For the predictive partitioning of the coding block, any one of the partitioning mode (Part_mode) candidates representing the partitioning shape of the coding block can be specified. Information for determining a partitioning index indicating any one of the partitioning mode candidates can be sent by a bitstream signal. Alternatively, the partitioning index of the coding block can be determined based on at least one of the size, shape, or coding mode of the coding block. The size or shape of the prediction block can be determined based on the partitioning mode specified by the partitioning index. The partitioning mode candidates may include asymmetric partitioning shapes (e.g., nLx2N, nRx2N, 2NxnU, 2NxnD). The number or type of asymmetric partitioning mode candidates available for the coding block can be determined based on at least one of the size, shape, or coding mode of the coding block.
[0110] Figure 3 is a diagram showing partition mode candidates that can be applied to a coding block when the coding block is encoded by inter prediction.
[0111] When encoding a coding block by inter-frame prediction, Figure 3 Any one of the 8 partition mode candidates shown in is applied to the coding block.
[0112] On the other hand, when the coding block is encoded by intra prediction, only the square partitioning mode may be applied to the coding block. In other words, when the coding block is encoded by intra prediction, the partitioning mode, PART_2Nx2N or PART_NxN, may be applied to the coding block.
[0113] When the coding block has a minimum size, PART_NxN can be applied. In this article, the minimum size of the coding block can be predefined in the encoder and decoder. Alternatively, information related to the minimum size of the coding block can be signaled through the bitstream. In an example, the minimum size of the coding block can be signaled through the slice header. Therefore, the minimum size of the coding block can be determined differently for each slice.
[0114] In another example, the partition mode candidates available for the coding block may be determined differently according to at least one of the size or shape of the coding block. In an example, the number or type of partition mode candidates available for the coding block may be determined differently according to at least one of the size or shape of the coding block.
[0115] Alternatively, the type or number of asymmetric partitioning pattern candidates available for the coding block may be determined based on the size or shape of the coding block. The number or type of asymmetric partitioning pattern candidates available for the coding block may be determined differently according to at least one of the size or shape of the coding block. In the example, when the coding block has a non-square shape with a width greater than the height, at least one of PART_2NxN, PART_2NxnU, or PART_2NxnD may not be used as a partitioning pattern candidate for the coding block. When the coding block has a non-square shape with a height greater than the width, at least one of PART_Nx2N, PART_nLx2N, PART_nRx2N may not be used as a partitioning pattern candidate for the coding block.
[0116] Generally, the prediction block may have a size of 4x4 to 64x64. However, when the coding block is encoded by inter-frame prediction, the prediction block may be restricted to not have a size of 4x4 to reduce the memory bandwidth when performing motion compensation.
[0117] Based on the partition mode, the coding block may be recursively partitioned. In other words, based on the partition mode determined by the partition index, the coding block may be partitioned and each partition generated by partitioning the coding block may be defined as a coding block.
[0118] Hereinafter, the method of dividing the coding unit will be described in more detail. In the examples mentioned later, the coding unit may refer to a coding tree unit or a coding unit included in a coding tree unit. In addition, "division" generated by dividing a coding block may refer to a "coding block". The division method mentioned later may be applied when dividing a coding block into a plurality of prediction blocks or transform blocks.
[0119] The coding unit may be divided by at least one line. In this case, the angle of the line dividing the coding unit may be a value in the range of 0 to 360 degrees. For example, the angle of the horizontal line may be 0 degrees, the angle of the vertical line may be 90 degrees, the angle of the diagonal line in the upper right direction may be 45 degrees, and the angle of the upper left diagonal line may be 135 degrees.
[0120] When the coding unit is divided by multiple lines, all the lines in the multiple lines may have the same angle. Alternatively, at least one of the multiple lines may have an angle different from the other lines. Alternatively, the multiple lines that divide the coding tree unit or the coding unit may have a predefined angle difference (e.g., 90 degrees).
[0121] Information about the lines dividing the coding unit may be determined by a division pattern. Alternatively, information about at least one of the number, direction, angle, or position of the lines in the block may be encoded.
[0122] For convenience of description, in an example mentioned later, it is assumed that a coding unit is divided into a plurality of coding units by using at least one of a vertical line or a horizontal line.
[0123] The number of vertical lines or horizontal lines that divide the coding unit may be at least one or more. In the example, the coding unit may be divided into 2 partitions by using one vertical line or one horizontal line. Alternatively, the coding unit may be divided into 3 partitions by using two vertical lines or two horizontal lines. Alternatively, the coding unit may be divided into 4 partitions by using one vertical line or one horizontal line, and the width and height of the 4 partitions are both halved relative to the coding unit.
[0124] When a coding unit is divided into multiple partitions by using at least one vertical line or at least one horizontal line, the partitions may have a uniform size. Alternatively, one partition may have a different size from the other partitions, or each partition may have a different size. In the example, when the coding unit is divided by two horizontal lines or two vertical lines, the coding unit may be divided into 3 partitions. In this case, the width ratio or height ratio of the 3 partitions may be n:2n:n, 2n:n:n, or n:n:2n.
[0125] In the examples mentioned later, dividing the coding block into 4 partitions is called quadtree-based partitioning. And, dividing the coding block into 2 partitions is called binary tree-based partitioning. In addition, dividing the coding block into 3 partitions is called ternary tree-based partitioning.
[0126] In the figures mentioned later, it will be shown that the coding unit is divided using one vertical line and / or one horizontal line, but it will be described that dividing the coding unit into more divisions than shown by using more vertical lines and / or more horizontal lines than shown or dividing the coding unit into fewer divisions than shown is also included in the scope of the present invention.
[0127] Figure 4 An example of hierarchical division of coding blocks based on a tree structure is shown as an embodiment to which the present invention is applied.
[0128] The input video signal is decoded in predetermined block units, and the basic unit for decoding the input video signal is called a coding block. A coding block may be a unit that performs intra / inter prediction, transformation, and quantization. In addition, a prediction mode (e.g., an intra prediction mode or an inter prediction mode) may be determined in units of coding blocks, and the prediction blocks included in the coding block may share the determined prediction mode. A coding block may be a square block or a non-square block of any size in the range of 8x8 to 64x64, or may be a square block or a non-square block having a size of 128x128, 256x256, or larger.
[0129] Specifically, the coding block may be hierarchically divided based on at least one of a quadtree partitioning method, a binary tree partitioning method, or a ternary tree partitioning method. Quadtree-based partitioning may mean a method of dividing a 2Nx2N coding block into four NxN coding blocks. Binary tree-based partitioning may mean a method of dividing one coding block into two coding blocks. Ternary tree-based partitioning may mean a method of dividing one coding block into three coding blocks. Even when ternary tree-based or binary tree-based partitioning is performed, square coding blocks may exist at a lower depth.
[0130] The partitions generated by the binary tree-based partitioning may be symmetric or asymmetric. In addition, the coding blocks based on the binary tree partitioning may be square blocks or non-square blocks (eg, rectangular).
[0131] Figure 52 is a diagram showing the division shape of a coding block based on binary tree division. The division shape of the coding block based on binary tree division may include a symmetric type such as 2NxN (non-square coding unit in the horizontal direction) or Nx2N (non-square coding unit in the vertical direction), or an asymmetric type such as nLx2N, nRx2N, 2NxnU, or 2NxnD. Only one of the symmetric type or the asymmetric type may be allowed as the division shape of the coding block.
[0132] The ternary tree partition shape may include at least one of a shape of dividing the coding block into 2 vertical lines or a shape of dividing the coding block into 2 horizontal lines. Three non-square partitions may be generated by the ternary tree partition.
[0133] Figure 6 A ternary tree partition shape is shown.
[0134] The ternary tree partition shape may include a shape of dividing the coding block into 2 horizontal lines or a shape of dividing the coding block into 2 vertical lines. The width ratio or height ratio of the partition generated by dividing the coding block may be n:2n:n, 2n:n:n, or n:n:2n.
[0135] The position of the partition having the largest width or height among the 3 partitions may be predefined in the encoder and the decoder. Alternatively, information showing the partition having the largest width or height among the 3 partitions may be signaled in the bitstream.
[0136] For coding units, only square-shaped or non-square-shaped symmetrical partitions may be allowed. In this case, the partitioning of the coding unit into square partitions may correspond to a quadtree CU partition, and the partitioning of the coding unit into a symmetrically shaped non-square partition may correspond to a binary tree partition. The partitioning of the coding tree unit into a square partition and a symmetrically shaped non-square partition may correspond to a quadtree and binary tree CU partition (QTBT).
[0137] Binary tree or ternary tree based partitioning may be performed on a coding block for which quadtree based partitioning is no longer performed. A coding block generated by binary tree or ternary tree based partitioning may be divided into smaller coding blocks. In this case, at least one of quadtree partitioning, ternary tree partitioning, or binary tree partitioning may be set not to be applied to the coding block. Alternatively, for a coding block, binary tree partitioning in a predetermined direction or ternary tree partitioning in a predetermined direction may not be allowed. In an example, for a coding block generated by binary tree or ternary tree based partitioning, quadtree partitioning and ternary tree partitioning may be set to be not allowed. For the coding block, only binary tree partitioning may be allowed.
[0138] Alternatively, only the largest coding block among the three coding blocks generated by the ternary tree-based partitioning may be partitioned into smaller coding blocks. Alternatively, only the largest coding block among the three coding blocks generated by the ternary tree-based partitioning may be allowed to be partitioned based on a binary tree or partitioned based on a ternary tree.
[0139] The division shape of the lower depth division can be determined accordingly based on the division shape of the higher depth division. In the example, when the higher division and the lower division are divided based on the binary tree, for the lower depth division, only the binary tree-based division of the same shape as the binary tree division shape of the higher depth division can be allowed. For example, when the binary tree division shape of the higher depth division is 2NxN, the binary tree division shape of the lower depth division can also be set to 2NxN. Alternatively, when the binary tree division shape of the higher depth division is Nx2N, the division shape of the lower depth division can also be set to Nx2N.
[0140] Alternatively, for the maximum partition among the partitions generated by ternary-tree based partitioning, binary tree partitioning in the same partition direction as a partition at a higher depth or ternary tree partitioning in the same partition direction as a partition at a higher depth may be set to be disallowed.
[0141] Alternatively, the partition shape of the partition of the lower depth can be determined by considering the partition shape of the partition of the higher depth and the partition shape of the adjacent partition of the lower depth. Specifically, if the partition of the higher depth is partitioned based on a binary tree, the partition shape of the partition of the lower depth can be determined so that the same result as the result of partitioning the partition of the higher depth based on a quadtree will not occur. In the example, when the partition shape of the partition of the higher depth is 2NxN and the partition shape of the adjacent partition of the lower depth is Nx2N, the partition shape of the current partition of the lower depth may not be set to Nx2N. This is because when the partition shape of the current partition of the lower depth is Nx2N, it leads to the same result as the result of partitioning the partition of the higher depth based on a quadtree of an NxN shape. When the partition shape of the partition of the higher depth is Nx2N and the partition shape of the adjacent partition of the lower depth is 2NxN, the partition shape of the current partition of the lower depth may not be set to 2NxN. In other words, when the binary tree partition shape of a higher depth partition is different from the binary tree partition shape of an adjacent lower depth partition, the binary tree partition shape of the current lower depth partition may be set to be the same as that of the higher depth partition.
[0142] Alternatively, the binary tree division shape of the division at a lower depth may be set to be different from the binary tree division shape of the division at a higher depth.
[0143] The allowed binary tree partition shape can be determined in units of sequence, slice or coding unit. In an example, the binary tree partition shape allowed by the coding tree unit can be limited to 2NxN or Nx2N shape. The allowed partition shape can be predefined in the encoder or decoder. Alternatively, information about the allowed partition shape or the disallowed partition shape can be encoded and sent by signaling through the bitstream.
[0144] Figure 7 is a diagram showing an example of binary tree-based partitioning that allows only specific shapes.
[0145] Figure 7 (a) represents an example in which only partitioning based on a binary tree of Nx2N shape is allowed, and Figure 7 (b) shows an example in which only partitioning based on a binary tree in a 2NxN shape is allowed.
[0146] In order to represent various partition shapes, information about quadtree partition, information about binary tree partition, or information about ternary tree partition may be used. The information about quadtree partition may include at least one of information indicating whether quadtree-based partition is performed or information about the size / depth of the coding block that allows quadtree-based partition. The information about binary tree partition may include at least one of the following: information indicating whether binary tree-based partition is performed, information about whether binary tree-based partition is vertical or horizontal, information about the size / depth of the coding block that allows binary tree-based partition, or information about the size / depth of the coding block that does not allow binary tree-based partition. The information about ternary tree partition may include at least one of the following: information indicating whether ternary tree-based partition is performed, information about whether ternary tree-based partition is vertical or horizontal, information about the size / depth of the coding block that allows ternary tree-based partition, or information about the size / depth of the coding block that does not allow ternary tree-based partition. The information about the size of the coding block may represent at least one minimum or maximum value of the width, height, product of width and height, or ratio of width to height of the coding block.
[0147] In an example, when the width or height of a coding block is smaller than the minimum size of allowed binary tree partitioning, or when the partition depth of the coding block is greater than the maximum depth of allowed binary tree partitioning, binary tree-based partitioning may not be allowed for the coding block.
[0148] In an example, when the width or height of a coding block is smaller than the minimum size allowed for ternary tree partitioning, or when the partition depth of the coding block is greater than the maximum depth allowed for ternary tree partitioning, ternary tree-based partitioning may not be allowed for the coding block.
[0149] Information about the conditions for allowing binary tree-based or ternary tree-based partitioning may be signaled through a bitstream. The information may be encoded in units of a sequence, a picture, or a partial image. A partial image may mean at least one of a slice, a tile group, a tile, a brick, a coding block, a prediction block, or a transform block.
[0150] In an example, a syntax 'max_mtt_depth_idx_minus1' indicating a maximum depth allowed for binary / ternary tree partitioning may be encoded / decoded by a bitstream. In this case, max_mtt_depth_idx_minus1+1 may indicate a maximum depth allowed for binary / ternary tree partitioning.
[0151] In an example, at least one of the number of allowed binary / ternary tree partitions, the maximum depth of allowed binary / ternary tree partitions, or the number of depths of allowed binary / ternary tree partitions may be signaled at a sequence or slice level. Thus, at least one of the number of allowed binary / ternary tree partitions, the maximum depth of allowed binary / ternary tree partitions, or the number of depths of allowed binary / ternary tree partitions may be different for a first slice and a second slice. In an example, while for a first slice, binary / ternary tree partitions may be allowed at only one depth, for a second slice, binary / ternary tree partitions may be allowed at two depths.
[0152] exist Figure 8 In the example shown in Figure 8 It is shown that binary tree partitioning is performed on a coding unit having a depth of 2 and a coding unit having a depth of 3. Therefore, at least one of the following can be encoded / decoded through a bitstream: information indicating the number of times binary tree partitioning is performed in a coding tree unit (2 times), information indicating the maximum depth (depth 3) of partitions generated by binary tree partitioning in a coding tree unit, or information indicating the number of partition depths (2 depths, depth 2 and depth 3) to which binary tree partitioning is applied in a coding tree unit.
[0153] Alternatively, the number of times binary / ternary tree divisions are allowed, the depth of binary / ternary tree divisions allowed, or the number of depths of binary / ternary tree divisions allowed may be predefined in the encoder and decoder. Alternatively, the number of times binary / ternary tree divisions are allowed, the depth of binary / ternary tree divisions allowed, or the number of depths of binary / ternary tree divisions allowed may be determined based on at least one of the index of the sequence or slice or the size / shape of the coding unit. In the example, for the first slice, binary / ternary tree division may be allowed at one depth, and for the second slice, binary / ternary tree division may be allowed at two depths.
[0154] In another example, at least one of the number of times binary tree partitioning is allowed, the depth of binary tree partitioning is allowed, or the number of depths of binary tree partitioning is allowed can be set differently according to a temporal level identifier (TemporalID) of a slice or picture. In this article, the temporal level identifier (TemporalID) is used to identify each of a plurality of layers of scalability having at least one or more perspective, space, time, or quality in an image.
[0155] like Figure 4 As shown in , the first coding block 300 with a division depth (split depth) of k can be divided into multiple second coding blocks based on a quadtree. For example, the second coding blocks 310 to 340 can be square blocks with a width and height halved relative to the first coding block, and the division depth of the second coding block can be increased to k+1.
[0156] The second coding block 310 with a division depth of k+1 may be divided into a plurality of third coding blocks with a division depth of k+2. The division of the second coding block 310 may be performed by selectively using one of a quadtree or a binary tree according to a division method. In this case, the division method may be determined based on at least one of information indicating a quadtree-based division or information indicating a binary tree-based division.
[0157] When the second coding block 310 is divided based on a quadtree, the second coding block 310 may be divided into four third coding blocks 310a whose width and height are halved relative to the second coding block, and the division depth of the third coding block 310a may be increased to k+2. In other words, when the second coding block 310 is divided based on a binary tree, the second coding block 310 may be divided into two third coding blocks. In this case, each of the two third coding blocks may be a non-square block whose width and height are halved relative to the second coding block, and the division depth may be increased to k+2. The second coding block may be determined as a non-square block in the horizontal direction or in the vertical direction according to the division direction, and the division direction may be determined based on information about whether the binary tree-based division is performed in the vertical direction or in the horizontal direction.
[0158] Meanwhile, the second coding block 310 may be determined as a leaf coding block that is no longer divided based on the quadtree or the binary tree, and in this case, the corresponding coding block may be used as a prediction block or a transform block.
[0159] Similar to the division of the second coding block 310, the third coding block 310a may be determined as a leaf coding block, or may be further divided based on a quadtree or a binary tree.
[0160] On the other hand, the third coding block 310b divided based on the binary tree can be further divided into a coding block 310b-2 in the vertical direction or a coding block 310b-3 in the horizontal direction based on the binary tree, and the division depth of the corresponding coding block can be increased to k+3. Alternatively, the third coding block 310b can be determined as a leaf coding block 310b-1 that is no longer divided based on the binary tree, and in this case, the corresponding coding block 310b-1 can be used as a prediction block or a transform block. However, the above-mentioned division process can be restrictively performed based on at least one of the following: information about the size / depth of the coding block that allows quadtree-based division, information about the size / depth of the coding block that allows binary tree-based division, or information about the size / depth of the coding block that does not allow binary tree-based division.
[0161] The number of candidates representing the size of the coding block may be limited to a predetermined number, or the size of the coding block in a predetermined unit may have a fixed value. In an example, the size of the coding block in a sequence or in a picture may be limited to any one of 256x256, 128x128, or 32x32. Information representing the size of the coding block in the sequence or in the picture may be signaled in a sequence header or a picture header.
[0162] As a result of quadtree and binary tree based partitioning, the coding unit can be represented as a square or rectangular shape of any size.
[0163] like Figure 4 As shown in , the first coding block 300 with a division depth (split depth) of k can be divided into multiple second coding blocks based on a quadtree. For example, the second coding blocks 310 to 340 can be square blocks with a width and height halved relative to the first coding block, and the division depth of the second coding block can be increased to k+1.
[0164] The second coding block 310 with a division depth of k+1 may be divided into a plurality of third coding blocks with a division depth of k+2. The division of the second coding block 310 may be performed by selectively using one of a quadtree or a binary tree according to a division method. In this case, the division method may be determined based on at least one of information indicating a quadtree-based division or information indicating a binary tree-based division.
[0165] When the second coding block 310 is divided based on a quadtree, the second coding block 310 may be divided into four third coding blocks 310a whose width and height are halved relative to the second coding block, and the division depth of the third coding block 310a may be increased to k+2. In other words, when the second coding block 310 is divided based on a binary tree, the second coding block 310 may be divided into two third coding blocks. In this case, each of the two third coding blocks may be a non-square block whose width and height are halved relative to the second coding block, and the division depth may be increased to k+2. The second coding block may be determined as a non-square block in the horizontal direction or in the vertical direction according to the division direction, and the division direction may be determined based on information about whether the binary tree-based division is performed in the vertical direction or in the horizontal direction.
[0166] Meanwhile, the second coding block 310 may be determined as a leaf coding block that is no longer divided based on the quadtree or the binary tree, and in this case, the corresponding coding block may be used as a prediction block or a transform block.
[0167] Similar to the division of the second coding block 310, the third coding block 310a may be determined as a leaf coding block, or may be further divided based on a quadtree or a binary tree.
[0168] On the other hand, the third coding block 310b divided based on the binary tree can be further divided into a coding block 310b-2 in the vertical direction or a coding block 310b-3 in the horizontal direction based on the binary tree, and the division depth of the corresponding coding block can be increased to k+3. Alternatively, the third coding block 310b can be determined as a leaf coding block 310b-1 that is no longer divided based on the binary tree, and in this case, the corresponding coding block 310b-1 can be used as a prediction block or a transform block. However, the above-mentioned division process can be restrictively performed based on at least one of the following: information about the size / depth of the coding block that allows quadtree-based division, information about the size / depth of the coding block that allows binary tree-based division, or information about the size / depth of the coding block that does not allow binary tree-based division.
[0169] The number of candidates representing the size of the coding block may be limited to a predetermined number, or the size of the coding block in a predetermined unit may have a fixed value. In an example, the size of the coding block in a sequence or in a picture may be limited to any one of 256x256, 128x128, or 32x32. Information representing the size of the coding block in the sequence or in the picture may be signaled in a sequence header or a picture header.
[0170] As a result of quadtree and binary tree based partitioning, the coding unit can be represented as a square or rectangular shape of any size.
[0171] The transform skip may be set not to be used for a coding unit generated by binary tree-based partitioning or ternary tree-based partitioning. Alternatively, the transform skip may be set to be applied to at least one of a vertical direction or a horizontal direction in a non-square coding unit. In the example, when the transform skip is applied to the horizontal direction, it means that scaling is performed only in the horizontal direction without performing transform / inverse transform, and transform / inverse transform using DCT or DST is performed in the vertical direction. When the transform skip is applied to the vertical direction, it means that scaling is performed only in the vertical direction without performing transform / inverse transform, and transform / inverse transform using DCT or DST is performed in the horizontal direction.
[0172] Information on whether to skip inverse transform in the horizontal direction or information on whether to skip inverse transform in the vertical direction may be signaled through a bitstream. In an example, information on whether to skip inverse transform in the horizontal direction may be a 1-bit flag 'hor_transform_skip_flag', and information on whether to skip inverse transform in the vertical direction may be a 1-bit flag 'ver_transform_skip_flag'.
[0173] The encoder may determine whether to encode 'hor_transform_skip_flag' or 'ver_transform_skip_flag' according to the size and / or shape of the current block. In an example, when the current block has an Nx2N shape, hor_transform_skip_flag may be encoded and the encoding of ver_transform_skip_flag may be omitted. When the current block has a 2NxN shape, ver_transform_skip_flag may be encoded and hor_transform_skip_flag may be omitted.
[0174] Alternatively, based on the size and / or shape of the current block, it may be determined whether to perform transform skipping in the horizontal direction or whether to perform transform skipping in the vertical direction. In an example, when the current block has an Nx2N shape, transform skipping may be applied to the horizontal direction, and transform / inverse transform may be performed in the vertical direction. When the current block has a 2NxN shape, transform skipping may be applied to the vertical direction, and transform / inverse transform may be performed in the horizontal direction. Transform / inverse transform may be performed based on at least one of DCT or DST.
[0175] As a result of partitioning based on a quadtree, a binary tree, or a ternary tree, a coding block that is no longer divided can be used as a prediction block or a transform block. In other words, a coding block generated by quadtree partitioning or binary tree partitioning can be used as a prediction block or a transform block. In the example, a predicted image can be generated in units of coding blocks, and the difference between the residual signal, the original image, and the predicted image can be transformed in units of coding blocks. In order to generate a predicted image in units of coding blocks, motion information can be determined based on the coding blocks, or an intra-frame prediction mode can be determined based on the coding blocks. Therefore, the coding block can be encoded by using at least one of a skip mode, an intra-frame prediction, or an inter-frame prediction.
[0176] Alternatively, multiple coding blocks generated by dividing the coding block may be set to share at least one of motion information, merge candidates, reference samples, reference sample rows, or intra-frame prediction modes. In an example, when the coding block is divided by a ternary tree, the divisions generated by dividing the coding block may share at least one of motion information, merge candidates, reference samples, reference sample rows, or intra-frame prediction modes according to the size or shape of the coding block. Alternatively, only a portion of the multiple coding blocks may be set to share information, and the remaining coding blocks may be set not to share information.
[0177] In another example, a prediction block or a transform block smaller than the coding block may be used by partitioning the coding block.
[0178] Hereinafter, a method of performing inter prediction on a coding block or a prediction block generated by splitting the coding block will be described in detail.
[0179] Fig. 9 : is a flowchart showing an inter-frame prediction method as an embodiment to which the present invention is applied.
[0180] Reference Fig. 9 , the motion information of the current block may be determined S910. The motion information of the current block may include at least one of a motion vector of the current block, a reference picture index of the current block, an inter-prediction direction of the current block, or a weight of weighted prediction. The weight of weighted prediction may represent a weight applied to the L0 reference block and a weight applied to the L1 reference block.
[0181] The motion vector of the current block may be determined based on information signaled by the bitstream. The precision of the motion vector represents the basic unit used to express the motion vector of the current block. For example, the precision of the motion vector of the current block may be determined as one of integer pixels, 1 / 2 pixels, 1 / 4 pixels, or 1 / 8 pixels. The precision of the motion vector may be determined on a per-picture basis, on a per-slice basis, on a per-tile group basis, on a per-tile basis, or on a per-block basis. A block may represent a coding tree unit, a coding unit, a prediction unit, or a transform unit.
[0182] The motion information of the current block may be obtained based on at least one of information signaled through a bitstream or motion information of a neighboring block adjacent to the current block.
[0183] Fig.10 is a diagram illustrating a process of deriving motion information of a current block when a merge mode is applied to the current block.
[0184] The merge mode indicates a method of deriving motion information of a current block from neighboring blocks.
[0185] When the merge mode is applied to the current block, a spatial merge candidate may be derived according to the spatial neighboring blocks of the current block S1010. The spatial neighboring blocks may include at least one of the blocks adjacent to the upper boundary, left boundary, or corner (e.g., at least one of the upper left corner, upper right corner, or lower left corner) of the current block.
[0186] Fig.11 is a diagram showing an example of spatially neighboring blocks.
[0187] As Fig.11 In the example shown in , the spatial neighboring blocks may include at least one of the following: a neighboring block A adjacent to the left side of the current block 1 , the neighboring block B1 adjacent to the upper side of the current block, the neighboring block A adjacent to the lower left corner of the current block 0 , the adjacent block B adjacent to the upper right corner of the current block 0 and the neighboring block B adjacent to the upper left corner of the current block 2 For example, assume that the position of the top left corner sample of the current block is (0, 0), the width of the current block is W, and the height of the current block is H. Block A 1 The sample at position (-1, H-1) may be included. Block B 1 The sample at position (W-1, -1) may be included. Block A 0 The sample at position (-1, H) may be included. Block B 0 The sample at position (W, -1) may be included. Block B 2 A sample at position (-1, -1) may be included.
[0188] Further expansion Fig.11In an example, a spatial merge candidate may be derived based on a block adjacent to the upper left sample of the current block or a block adjacent to the upper center sample of the current block. For example, a block adjacent to the upper left sample of the current block may include at least one of a block including a sample at position (0, -1) or a block including a sample at position (-1, 0). Alternatively, a spatial merge candidate may be derived based on at least one of a block adjacent to the upper center sample of the current block or a block adjacent to the left center sample of the current block. For example, a block adjacent to the upper center sample of the current block may include a sample at position (W / 2, -1). A block adjacent to the left center sample of the current block may include a sample at position (-1, H / 2).
[0189] Based on the size and / or shape of the current block, the position of the upper adjacent block and / or the left adjacent block for deriving the spatial merge candidate can be determined. In the example, when the size of the current block is greater than a threshold, the spatial merge candidate can be derived based on the block adjacent to the upper center sample of the current block and the block adjacent to the left center sample of the current block. On the other hand, when the size of the current block is less than the threshold, the spatial merge candidate can be derived based on the block adjacent to the upper right sample of the current block and the block adjacent to the lower left sample of the current block. In this article, the size of the current block can be expressed based on at least one of the width, height, the sum of the width and the height, the product of the width and the height, or the ratio of the width to the height. The threshold can be an integer, such as 2, 4, 8, 16, 32, or 128.
[0190] According to the shape of the current block, the availability of the extended spatial neighboring blocks can be determined. In the example, when the current block is a non-square block with a width greater than a height, it can be determined that the block adjacent to the upper left sample of the current block, the block adjacent to the left center sample, or the block adjacent to the lower left sample of the current block is unavailable. At the same time, when the current block is a block with a height greater than a width, it can be determined that the block adjacent to the upper left sample of the current block, the block adjacent to the upper center sample, or the block adjacent to the upper right sample of the current block is unavailable.
[0191] The motion information of the spatial merging candidate may be set to be the same as the motion information of the spatial neighboring blocks.
[0192] Spatial merging candidates may be determined by searching adjacent blocks in a predetermined order. Fig.11 In the example shown in FIG, block A 1 , B 1 , B 0 , A 0 and B 2 The search is performed in the order of A to determine the spatial merging candidate. In this paper, when there are no other blocks (i.e., A 1 , B 1 , B 0 and A 0) or at least one of them is encoded by an intra-frame prediction mode, block B can be used 2 .
[0193] The order of searching for spatial merging candidates may be predefined in the encoder / decoder. Alternatively, the order of searching for spatial merging candidates may be adaptively determined according to the size or shape of the current block. Alternatively, the order of searching for spatial merging candidates may be determined based on information signaled through the bitstream.
[0194] A temporal merge candidate may be derived from the temporal neighboring blocks of the current block S1020. The temporal neighboring blocks may refer to the co-located blocks included in the co-located picture. The co-located picture has a different POC from the current picture including the current block. The co-located picture may be determined as a picture with a predefined index in a reference picture list or a picture with the smallest POC difference from the current picture. Alternatively, the co-located picture may be determined by information signaled by a bitstream. The information signaled by the bitstream may include at least one of: information indicating a reference picture list (e.g., an L0 reference picture list or an L1 reference picture list) including the co-located picture and an index indicating the co-located picture in the reference picture list. Information for determining the co-located picture may be signaled in at least one of a picture parameter set, a slice header, and a block level.
[0195] The motion information about the temporal merge candidate may be determined based on the motion information of the co-located block. In an example, the motion vector of the temporal merge candidate may be determined based on the motion vector of the co-located block. For example, the motion vector of the temporal merge candidate may be set to be the same as the motion vector of the co-located block. Alternatively, the motion vector of the temporal merge candidate may be derived by scaling the motion vector of the co-located block based on at least one of the POC difference between the current picture of the current block and the reference picture and the POC difference between the co-located co-located picture and the reference picture.
[0196] Fig.12 is a diagram showing an example of deriving a motion vector of a temporal merging candidate.
[0197] exist Fig.12 In the example shown in , tb represents the POC difference between the current picture curr_pic and the reference picture curr_ref of the current picture, and td represents the POC difference between the co-located picture col_pic and the reference picture col_ref of the co-located block. The motion vector of the temporal merge candidate can be derived by scaling the motion vector of the co-located block col_PU based on tb and / or td.
[0198] Alternatively, considering whether the co-located block is available, the motion vector of the co-located block and the motion vector obtained by scaling the motion vector of the co-located block can be used as the motion vector of the temporal merge candidate. In an example, the motion vector of the co-located block is set as the motion vector of the first temporal merge candidate, and the value obtained by scaling the motion vector of the co-located block can be set as the motion vector of the second temporal merge candidate.
[0199] The inter-frame prediction direction of the temporal merge candidate may be set to be the same as the inter-frame prediction direction of the temporal neighboring block. However, the reference picture index of the temporal merge candidate may have a fixed value. In the example, the reference picture index of the temporal merge candidate may be set to "0". Alternatively, the reference picture index of the temporal merge candidate may be adaptively determined based on at least one of the reference picture index of the spatial merge candidate and the reference picture index of the current picture.
[0200] A specific block having the same position and size as the current block within the co-located picture, or a block adjacent to a block adjacent to a block having the same position and size as the current block may be determined as the co-located block.
[0201] Fig.13 is a diagram showing the positions of candidate blocks that can be used as co-located blocks.
[0202] The candidate blocks may include at least one of: a block adjacent to the upper left corner of the current block in the same picture, a block adjacent to the center sample of the current block in the same picture, and a block adjacent to the lower left corner of the current block in the same picture.
[0203] In an example, the candidate blocks may include at least one of: a block TL within the co-located picture including the position of the upper left sample of the current block, a block BR within the co-located picture including the position of the lower right sample of the current block, a block H within the co-located picture adjacent to the lower right corner of the current block, a block C3 within the co-located picture including the position of the center sample of the current block, and a block C0 within the co-located picture adjacent to the center sample of the current block (e.g., a block including the position of a sample separated from the center sample of the current block by (-1, -1)).
[0204] Apart from Fig.13 In addition to the example shown in , a block within the co-located picture including a position of a neighboring block adjacent to a predetermined boundary of the current block may be selected as the co-located block.
[0205] The number of temporal merging candidates may be 1 or more. In an example, at least one temporal merging candidate may be derived based on at least one co-located block.
[0206] Information about the maximum number of temporal merge candidates may be encoded by an encoder and sent by a signal. Alternatively, the maximum number of temporal merge candidates may be derived based on the maximum number of merge candidates that may be included in the merge candidate list and / or the maximum number of spatial merge candidates. Alternatively, the maximum number of temporal merge candidates may be determined based on the number of available co-located blocks.
[0207] Whether the candidate block is available can be determined according to a predetermined priority, and at least one co-located block can be determined based on the above determination and the maximum number of candidates for temporal merging. In the example, when block C3 including the position of the center sample of the current block and block H adjacent to the lower right corner of the current block are candidate blocks, any one of block C3 and block H can be determined as a co-located block. When block H is available, block H can be determined as a co-located block. However, when block H is not available (for example, when block H is encoded by intra-frame prediction, when block H is not available, or when block H is located outside the maximum coding unit (LCU), etc.), block C3 can be determined as a co-located block.
[0208] In another example, when at least one of the plurality of blocks adjacent to the lower right corner position of the current block in the co-located picture is unavailable (e.g., block H and / or block BR), the unavailable block may be replaced with another available block. The other available block replacing the unavailable block may include at least one block adjacent to the center sample position of the current block in the co-located picture (e.g., C0 and / or C3) and a block adjacent to the lower left corner of the current block in the co-located picture (e.g., TL).
[0209] When at least one of the multiple blocks adjacent to the center sample position of the current block in the same picture is unavailable, or when at least one of the multiple blocks adjacent to the upper left corner position of the current block in the same picture is unavailable, the unavailable block can be replaced by another available block.
[0210] Subsequently, a merge candidate list including a spatial merge candidate and a temporal merge candidate may be generated (S1030). When configuring the merge candidate list, a merge candidate having the same motion information as an existing merge candidate may be removed from the merge candidate list.
[0211] Information about the maximum number of merge candidates may be signaled via the bitstream. In an example, information indicating the maximum number of merge candidates may be signaled via a sequence parameter or a picture parameter. In an example, when the maximum number of merge candidates is six, a total of six may be selected from the spatial merge candidates and the temporal merge candidates. For example, five spatial merge candidates may be selected from five merge candidates, and one temporal merge candidate may be selected from two temporal merge candidates.
[0212] Alternatively, the maximum number of merge candidates may be predefined in the encoder and the decoder. For example, the maximum number of merge candidates may be 2, 3, 4, 5, or 6. Alternatively, the maximum number of merge candidates may be determined based on at least one of whether merge with MVD (MMVD) is performed, whether combined prediction is performed, or whether triangulation is performed.
[0213] If the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, the merge candidates included in the second merge candidate list may be added to the merge candidate list.
[0214] The second merge candidate list may include merge candidates derived based on motion information of a block encoded / decoded by inter-frame prediction before the current block. In an example, if motion compensation of a block whose coding mode is inter-frame prediction is performed, the merge candidate derived based on the motion information of the block may be added to the second merge candidate list. If encoding / decoding of the current block is completed, the motion information of the current block may be added to the second merge candidate list for inter-frame prediction of subsequent blocks.
[0215] The second merge candidate list may be initialized in units of CTU, tile, or slice. The maximum number of merge candidates that may be included in the second merge candidate list may be predefined in the encoder and decoder. Alternatively, information indicating the maximum number of merge candidates that may be included in the second merge candidate list may be signaled via a bitstream.
[0216] The indexes of the merge candidates included in the second merge candidate list may be determined based on the order of addition to the second merge candidate list. In an example, the index assigned to the Nth merge candidate added to the second merge candidate list may have a smaller value than the index assigned to the N+1th merge candidate added to the second merge candidate list. For example, the index of the N+1th merge candidate may be set to a value increased by 1 relative to the index of the Nth merge candidate. Alternatively, the index of the Nth merge candidate may be set to the index of the N+1th merge candidate, and the value of the index of the Nth merge candidate may be subtracted by 1.
[0217] Alternatively, the index assigned to the Nth merge candidate added to the second merge candidate list may have a greater value than the index assigned to the N+1th merge candidate added to the second merge candidate list. For example, the index of the Nth merge candidate may be set to the index of the N+1th merge candidate, and the value of the index of the Nth merge candidate may be increased by 1.
[0218] Based on whether the motion information of the block on which motion compensation is performed is the same as the motion information of the merge candidate included in the second merge candidate list, it can be determined whether to add the merge candidate derived from the block to the second merge candidate list. In the example, when a merge candidate having the same motion information as the block is included in the second merge candidate list, the merge candidate derived based on the motion information of the block may not be added to the second merge candidate list. Alternatively, when a merge candidate having the same motion information as the block is included in the second merge candidate list, the merge candidate may be deleted from the second merge candidate list, and the merge candidate derived based on the motion information of the block may be added to the second merge candidate list.
[0219] When the number of merge candidates included in the second merge candidate list is the same as the maximum number of merge candidates, the merge candidate with the lowest index or the merge candidate with the highest index may be deleted from the second merge candidate list, and the merge candidate derived based on the motion information of the block may be added to the second merge candidate list. In other words, after deleting the oldest merge candidate among the merge candidates included in the second merge candidate list, the merge candidate derived based on the motion information of the block may be added to the second merge candidate list.
[0220] When the number of merge candidates included in the merge candidate list has not reached the maximum number of merge candidates, a combined merge candidate obtained by combining two or more merge candidates or a merge candidate having a (0, 0) motion vector (zero motion vector) may be included in the merge candidate list.
[0221] Alternatively, an average merge candidate in which the average value of the motion vectors of two or more merge candidates is found may be added to the merge candidate list. The average merge candidate may be derived by finding the average value of the motion vectors of two or more merge candidates included in the merge candidate list. In the example, when the first merge candidate and the second merge candidate are added to the merge candidate list, the average value of the motion vector of the first merge candidate and the motion vector of the second merge candidate may be calculated so as to obtain the average merge candidate. In detail, the L0 motion vector of the average merge candidate may be derived by calculating the average value of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate, and the L1 motion vector of the average merge candidate may be derived by calculating the average value of the L1 motion vector of the first merge candidate and the L1 motion vector of the second merge candidate. When bidirectional prediction is applied to any one of the first merge candidate and the second merge candidate, and unidirectional prediction is performed on the other merge candidate, the motion vector of the bidirectional merge candidate may be set as is to the L0 motion vector or the L1 motion vector of the average merge candidate. In the example, when the L0 direction prediction and the L1 direction prediction are performed on the first merge candidate, but the L0 direction prediction is performed on the second merge candidate, the L0 motion vector of the average merge candidate can be derived by calculating the average of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate. At the same time, the L1 motion vector of the average merge candidate can be derived as the L1 motion vector of the first merge candidate.
[0222] When the reference picture of the first merge candidate is different from the second merge candidate, the motion vector of the first merge candidate or the second merge candidate may be scaled according to the distance between the reference picture of each merge candidate and the current picture (i.e., POC difference). For example, after scaling the motion vector of the second merge candidate, the average merge candidate may be derived by calculating the average of the motion vector of the first merge candidate and the scaled motion vector of the second merge candidate. In this article, priority may be set based on the value of the reference picture index of each merge candidate, the distance between the reference picture of each merge candidate and the current block, or whether bidirectional prediction is applied, and scaling may be applied to the motion vector of a merge candidate with a high (or low) priority.
[0223] The reference picture index of the average merge candidate may be set to indicate a reference picture at a specific position within the reference picture list. In an example, the reference picture index of the average merge candidate may indicate the first reference picture or the last reference picture within the reference picture list. Alternatively, the reference picture index of the average merge candidate may be set to be the same as the reference picture index of the first merge candidate or the second merge candidate. In an example, when the reference picture index of the first merge candidate is the same as the second merge candidate, the reference picture index of the average merge candidate may be set to be the same as the reference picture index of the first merge candidate and the second merge candidate. When the reference picture index of the first merge candidate is different from the second merge candidate, the priority may be set based on the value of the reference picture index of each merge candidate, the distance between the reference picture of each merge candidate and the current block, or whether bidirectional prediction is applied, and the reference picture index of the merge candidate with high (or low) priority may be set to the reference picture index of the average merge candidate. In an example, when bidirectional prediction is applied to the first merge candidate and unidirectional prediction is applied to the second merge candidate, the reference picture index of the first merge candidate to which bidirectional prediction is applied may be determined as the reference picture index of the average merge candidate.
[0224] Based on the priority between the combinations of merge candidates, a sequence of combinations for generating average merge candidates can be determined. The priorities can be predefined in the encoder and the decoder. Alternatively, the sequence of combinations can be determined based on whether bidirectional prediction of the merge candidates is performed. For example, a combination of merge candidates encoded using bidirectional prediction can be set to have a higher priority than a combination of merge candidates encoded using unidirectional prediction. Alternatively, the sequence of combinations can be determined based on the reference pictures of the merge candidates. For example, a combination of merge candidates with the same reference picture can have a higher priority than a combination of merge candidates with different reference pictures.
[0225] The merge candidates may be included in the merge candidate list according to a predefined priority. Merge candidates with high priority may be assigned to have small index values. In the example, the spatial merge candidate may be added to the merge candidate list before the temporal merge candidate. In addition, the spatial merge candidates may be added to the merge candidate list in the order of spatial merge candidates for the left neighboring block, spatial merge candidates for the upper neighboring block, spatial merge candidates for the block adjacent to the upper right corner, spatial merge candidates for the block adjacent to the lower left corner, and spatial merge candidates for the block adjacent to the upper left corner. Alternatively, it may be set so that the spatial merge candidates are added to the merge candidate list according to the order of spatial merge candidates for the neighboring block adjacent to the upper left corner of the current block ( Fig.11 The spatial merge candidate derived from B2) is added to the merge candidate list later than the temporal merge candidate.
[0226] In another example, the priority between the merge candidates may be determined according to the size or shape of the current block. In the example, when the current block has a rectangular shape with a width greater than a height, the spatial merge candidate of the left neighboring block may be added to the merge candidate list before the spatial merge candidate of the upper neighboring block. In other words, when the current block has a rectangular shape with a height greater than a width, the spatial merge candidate of the upper neighboring block may be added to the merge candidate list before the spatial merge candidate of the left neighboring block.
[0227] In another example, the priority between the merge candidates can be determined according to the motion information of each merge candidate. In the example, the merge candidate with bidirectional motion information can have a higher priority than the merge candidate with unidirectional motion information. Therefore, the merge candidate with bidirectional motion information can be added to the merge candidate list before the merge candidate with unidirectional motion information.
[0228] In another example, a merge candidate list may be generated according to a predefined priority, and then the merge candidates may be rearranged. The rearrangement may be performed based on the motion information of the merge candidate. In the example, the rearrangement may be performed based on whether the merge candidate has bidirectional motion information, the size of the motion vector, the precision of the motion vector, or the POC difference between the reference picture of the merge candidate and the current picture. In detail, a merge candidate with bidirectional motion information may be rearranged to have a higher priority than a merge candidate with unidirectional motion information. Alternatively, a merge candidate with a motion vector having a precision value of fractional pixels may be rearranged to have a higher priority than a merge candidate with a motion vector having a precision of integer pixels.
[0229] When generating a merge candidate list, at least one of the merge candidates included in the merge candidate list may be specified based on the merge candidate index ( S1040 ).
[0230] The motion information of the current block may be set to be the same as the motion information of the merge candidate specified by the merge candidate index (S1050). In the example, when a spatial merge candidate is selected by the merge candidate index, the motion information of the current block may be set to be the same as the motion information of the spatial neighboring block. Alternatively, when a temporal merge candidate is selected by the merge candidate index, the motion information of the current block may be set to be the same as the motion information of the temporal neighboring block.
[0231] Fig.14 is a diagram illustrating a process of deriving motion information of a current block when the AMVP mode is applied to the current block.
[0232] When the AMVP mode is applied to the current block, at least one of the inter prediction direction and the reference picture index of the current block may be decoded from the bitstream S1410. In other words, when the AMVP mode is applied, at least one of the inter prediction direction and the reference picture index of the current block may be determined based on information encoded by the bitstream.
[0233] A spatial motion vector candidate may be determined based on the motion vector of the spatial neighboring block of the current block S1420. The spatial motion vector candidate may include at least one of the following: a first spatial motion vector candidate derived from an upper neighboring block of the current block and a second spatial motion vector candidate derived from a left neighboring block of the current block. In this article, the upper neighboring block may include at least one of the blocks adjacent to the upper side and upper right corner of the current block, and the left neighboring block of the current block includes at least one of the blocks adjacent to the left side and lower left corner of the current block. The block adjacent to the upper left corner of the current block may be used as an upper neighboring block or may be used as a left neighboring block.
[0234] Alternatively, a spatial motion vector candidate may be derived from a spatial non-adjacent block that is not adjacent to the current block. In an example, a spatial motion vector candidate for the current block may be derived by using at least one of: a block located on the same vertical line as a block adjacent to the upper side, upper right corner, or upper left corner of the current block; a block located on the same horizontal line as a block adjacent to the left side, lower left corner, or upper left corner of the current block; and a block located on the same diagonal line as a block adjacent to a corner of the current block. When spatially adjacent blocks are not available, spatial motion vector candidates may be derived by using spatially non-adjacent blocks.
[0235] In another example, at least two spatial motion vector candidates may be derived by using spatial adjacent blocks and spatial non-adjacent blocks. In an example, a first spatial motion vector candidate and a second spatial motion vector candidate may be derived by using adjacent blocks adjacent to the current block. Meanwhile, a third spatial motion vector candidate and / or a fourth spatial motion vector candidate may be derived based on blocks that are not adjacent to the current block but are adjacent to the above adjacent blocks.
[0236] When the current block is different from the spatial neighboring blocks on the reference picture, the spatial motion vector may be obtained by scaling the motion vectors of the spatial neighboring blocks. A temporal motion vector candidate may be determined based on the motion vectors of the temporal neighboring blocks of the current block S1430. When the current block is different from the temporal neighboring blocks on the reference picture, the temporal motion vector may be obtained by scaling the motion vectors of the temporal neighboring blocks. In this article, when the number of spatial motion vector candidates is equal to or less than a predetermined number, the temporal motion vector candidate may be derived.
[0237] A motion vector candidate list including spatial motion vector candidates and temporal motion vector candidates may be generated S1440.
[0238] When generating the motion vector candidate list, at least one of the motion vector candidates included in the motion vector candidate list may be specified based on information specifying at least one of the motion vector candidate list S1450.
[0239] The motion vector candidate specified by the information may be set as a prediction value of the motion vector of the current block, and the motion vector of the current block may be obtained by adding the residual value of the motion vector and the prediction value of the motion vector S1460. Herein, the residual value of the motion vector may be parsed through the bitstream.
[0240] When the motion information of the current block is obtained, motion compensation of the current block may be performed based on the obtained motion information S920. In detail, motion compensation of the current block may be performed based on the inter prediction direction, the reference picture index, and the motion vector of the current block. The inter prediction direction indicates whether L0 prediction, L1 prediction, or bi-prediction is performed. When the current block is encoded by bi-prediction, a prediction block of the current block may be obtained based on a weighted sum operation or an average operation of the L0 reference block and the L1 reference block.
[0241] When the prediction sample is obtained by performing motion compensation, the current block can be reconstructed based on the generated prediction sample. In detail, the reconstructed sample can be obtained by adding the prediction sample of the current block to the residual sample.
[0242] As in the above example, based on the motion information of the block encoded / decoded using inter-frame prediction before the current block, the merge candidate of the current block can be derived. For example, based on the motion information of the neighboring block at the predefined position adjacent to the current block, the merge candidate of the current block can be derived. Examples of the neighboring blocks may include at least one of the following: a block adjacent to the left side of the current block, a block adjacent to the upper side of the current block, a block adjacent to the upper left corner of the current block, a block adjacent to the upper right corner of the current block, and a block adjacent to the lower left corner of the current block.
[0243] The merging candidate of the current block can be derived based on the motion information of the blocks other than the neighboring blocks. For ease of description, the neighboring block at a predefined position adjacent to the current block is referred to as a first merging candidate block, and the block at a position different from the first merging candidate block is referred to as a second merging candidate block.
[0244] The second merge candidate block may include at least one of a block encoded / decoded using inter prediction before the current block, a block adjacent to the first merge candidate block, or a block located on the same line as the first merge candidate block. Fig.15 A second merge candidate block adjacent to the first merge candidate block is shown, and Fig.16 A second merging candidate block located on the same line as the first merging candidate block is shown.
[0245] When the first merge candidate block is not available, the merge candidate derived based on the motion information of the second merge candidate block is added to the merge candidate list. Alternatively, even if at least one of the spatial merge candidate and the temporal merge candidate is added to the merge candidate list, when the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, the merge candidate derived based on the motion information of the second merge candidate block is added to the merge candidate list.
[0246] Fig.15 is a diagram showing an example of deriving a merge candidate from a second merge candidate block when a first merge candidate block is unavailable.
[0247] When the first merge candidate block AN (herein, N ranges from 0 to 4) is unavailable, a merge candidate of the current block is derived based on motion information of a second merge candidate block BM (herein, M ranges from 0 to 6). That is, the merge candidate of the current block can be derived by replacing the unavailable first merge candidate block with the second merge candidate block.
[0248] Among the blocks adjacent to the first merge candidate block, a block placed in a predefined direction from the first merge candidate block may be set as a second merge candidate block. The predefined direction may be a left direction, a right direction, an upward direction, a downward direction, or a diagonal direction. The predefined direction may be set for each first merge candidate block. For example, the predefined direction of the first merge candidate block adjacent to the left side of the current block may be a left direction. The predefined direction of the first merge candidate block adjacent to the upper side of the current block may be an upward direction. The predefined direction of the first merge candidate block adjacent to the corner of the current block may include at least one of a left direction, an upward direction, or a diagonal direction.
[0249] For example, when A0 adjacent to the left side of the current block is not available, a merge candidate for the current block is derived based on B0 adjacent to A1. When A1 adjacent to the upper side of the current block is not available, a merge candidate for the current block is derived based on B1 adjacent to A1. When A2 adjacent to the upper right corner of the current block is not available, a merge candidate for the current block is derived based on B2 adjacent to A2. When A3 adjacent to the lower left corner of the current block is not available, a merge candidate for the current block is derived based on B3 adjacent to A3. When A4 adjacent to the upper left corner of the current block is not available, a merge candidate for the current block is derived based on at least one of B4 to B6 adjacent to A4.
[0250] Fig.15 The example shown in is only used to describe the embodiments of the present invention, but does not limit the present invention. The position of the second merge candidate block may be set to Fig.15. For example, a second merge candidate block adjacent to a first merge candidate block may be positioned in an upward direction or a downward direction of the first merge candidate block, the first merge candidate block being adjacent to the left side of the current block. Alternatively, a second merge candidate block adjacent to a first merge candidate block may be positioned in a left direction or a right direction of the first merge candidate block, the first merge candidate block being adjacent to the upper side of the current block.
[0251] Fig.16 is a diagram showing an example of deriving a merge candidate from a second merge candidate block located on the same line as a first merge candidate block.
[0252] The blocks located on the same line as the first merge candidate block may include at least one of the following: a block located on the same horizontal line as the first merge candidate block, a block located on the same vertical line as the first merge candidate block, or a block located on the same diagonal line as the first merge candidate block. The y coordinate positions of the blocks located on the same horizontal line are the same. The x coordinate positions of the blocks located on the same vertical line are the same. The difference between the x coordinate positions of the blocks located on the same diagonal line is the same as the difference between the y coordinate positions.
[0253] Assume that the top left sample of the current block is located at (0, 0), and the width and height of the current block are W and H respectively. Fig.18 , it is shown that the position of the second merge candidate block (e.g., B4, C6) located on the same vertical line as the first merge candidate block is determined based on the rightmost block on the upper side of the coding block (e.g., block A1 including coordinates (W-1, -1)). In addition, Fig.18 , it is shown that the position of the second merge candidate block (e.g., B1, C1) located on the same horizontal line as the first merge candidate block is determined based on the lowest block on the left side of the coding block (e.g., block A0 including coordinates (-1, H-1)).
[0254] In another example, the position of the second merge candidate block may be determined based on the leftmost block on the upper side of the coding block (e.g., a block including coordinates (0, -1)) or a block located at the upper center of the coding block (e.g., a block including coordinates (W / 2, -1)). In addition, the position of the second merge candidate block may be determined based on the uppermost block on the left side of the coding block (e.g., a block including coordinates (-1, 0)) or a block located at the left center of the coding block (e.g., a block including coordinates (-1, H / 2)).
[0255] In another example, when there are multiple upper adjacent blocks adjacent to the upper side of the current block, the second merge candidate block can be determined by using all or part of the multiple upper adjacent blocks. In the example, the second merge candidate block can be determined by using a block at a specific position among the multiple upper adjacent blocks (for example, at least one of the upper adjacent block located at the leftmost side, the upper adjacent block located at the rightmost side, or the upper adjacent block located at the center). The number of upper adjacent blocks used to determine the second merge candidate block among the multiple upper adjacent blocks can be 1, 2, 3 or more. In addition, when there are multiple left adjacent blocks adjacent to the left side of the current block, the second merge candidate block can be determined by using all or part of the multiple left adjacent blocks. In the example, the second merge candidate block can be determined by using a block at a specific position among the multiple left adjacent blocks (for example, at least one of the left adjacent block located at the bottom, the left adjacent block located at the top, or the left adjacent block located at the center). The number of left adjacent blocks used to determine the second merge candidate block among the multiple left adjacent blocks can be 1, 2, 3 or more.
[0256] Depending on the size and / or shape of the current block, the position and / or number of the upper adjacent blocks and / or the left adjacent blocks used to determine the second merge candidate block can be determined differently. In the example, when the size of the current block is greater than the threshold, the second merge candidate block can be determined based on the upper center block and / or the left center block. On the other hand, when the size of the current block is less than the threshold, the second merge candidate block can be determined based on the upper right block and / or the lower left block. The threshold can be an integer, such as 8, 16, 32, 64, or 128.
[0257] A first merge candidate list and a second merge candidate list may be constructed, and motion compensation of the current block may be performed based on at least one of the first merge candidate list or the second merge candidate list.
[0258] The first merge candidate list may include at least one of: a spatial merge candidate derived based on motion information of a neighboring block at a predefined position adjacent to the current block, or a temporal merge candidate derived based on motion information of a co-located block.
[0259] The second merge candidate list may include merge candidates derived based on the motion information of the second merge candidate block.
[0260] As an embodiment of the present invention, the first merge candidate list may be constructed to include merge candidates derived from the first merge candidate block, and the second merge candidate list may be constructed to include merge candidates derived from the second merge candidate block. Fig.15In the example shown in , the merge candidates derived from blocks A0 to A4 can be added to the first merge candidate list, and the merge candidates derived from blocks B0 to B6 can be added to the second merge candidate list. Fig.16 In the example shown in , merge candidates derived from blocks A0 to A4 may be added to the first merge candidate list, and merge candidates derived from blocks B0 to B5, C0 to C7 may be added to the second merge candidate list.
[0261] Alternatively, the second merge candidate list may include merge candidates derived based on motion information of a block encoded / decoded using inter-frame prediction before the current block. For example, when motion compensation of a block whose coding mode is inter-frame prediction is performed, a merge candidate derived based on the motion information of the block is added to the second merge candidate list. When encoding / decoding of the current block is completed, the motion information of the current block is added to the second merge candidate list for inter-frame prediction of subsequent blocks.
[0262] The indexes of the merge candidates included in the second merge candidate list may be determined based on the order in which the merge candidates are added to the second merge candidate list. For example, the index assigned to the Nth merge candidate added to the second merge candidate list may have a lower value than the index assigned to the N+1th merge candidate added to the second merge candidate list. For example, the index of the N+1th merge candidate may be set to have a value that is 1 higher than the index of the Nth merge candidate. Alternatively, the index of the Nth merge candidate may be set to the index of the N+1th merge candidate, and the value of the index of the Nth merge candidate is reduced by 1.
[0263] Alternatively, the index assigned to the Nth merge candidate added to the second merge candidate list may have a higher value than the index assigned to the N+1th merge candidate added to the second merge candidate list. For example, the index of the Nth merge candidate may be set to the index of the N+1th merge candidate, and the value of the index of the Nth merge candidate is increased by 1.
[0264] Based on whether the motion information of the block subjected to motion compensation is the same as the motion information of the merge candidate included in the second merge candidate list, it can be determined whether to add the merge candidate derived from the block to the second merge candidate list. For example, when a merge candidate having the same motion information as the block is included in the second merge candidate list, the merge candidate derived based on the motion information of the block is not added to the second merge candidate list. Alternatively, when a merge candidate having the same motion information as the block is included in the second merge candidate list, the merge candidate is deleted from the second merge candidate list, and the merge candidate derived based on the motion information of the block is added to the second merge candidate list.
[0265] When the number of merge candidates included in the second merge candidate list is the same as the maximum number of merge candidates, a merge candidate with a lowest index or a merge candidate with a highest index is detected from the second merge candidate list, and a merge candidate derived based on the motion information of the block is added to the second merge candidate list. That is, after deleting the oldest merge candidate among the merge candidates included in the second merge candidate list, a merge candidate derived based on the motion information of the block may be added to the second merge candidate list.
[0266] The second merge candidate list may be initialized in units of CTU, tile, or slice. In other words, a block different from the current block included in the CTU, tile, or slice may be set to be unavailable as a second merge candidate block. The maximum number of merge candidates that may be included in the second merge candidate list may be predefined in the encoder and the decoder. Alternatively, information indicating the maximum number of merge candidates that may be included in the second merge candidate list may be signaled via a bitstream.
[0267] The first merge candidate list or the second merge candidate list may be selected, and the selected merge candidate list may be used to perform inter prediction of the current block. Specifically, based on the index information, any one of the merge candidates included in the merge candidate list may be selected, and motion information of the current block may be obtained from the merge candidate.
[0268] Information specifying the first merge candidate list or the second merge candidate list may be signaled through the bitstream, and the decoder may select the first merge candidate list or the second merge candidate list based on the information.
[0269] Alternatively, among the first merge candidate list and the second merge candidate list, a merge candidate list including a larger number of available merge candidates may be selected.
[0270] Alternatively, the first merge candidate list or the second merge candidate list may be selected based on at least one of the size, shape, and split depth of the current block.
[0271] Alternatively, the merge candidate list is configured by adding (or appending) the other one of the first merge candidate list and the second merge candidate list to any one of the first merge candidate list and the second merge candidate list.
[0272] For example, inter prediction may be performed based on a merge candidate list having at least one merge candidate included in a first merge candidate list and at least one merge candidate included in a second merge candidate list.
[0273] For example, the merge candidates included in the second merge candidate list may be added to the first merge candidate list. Alternatively, the merge candidates included in the first merge candidate list may be added to the second merge candidate list.
[0274] When the number of merge candidates included in the first merge candidate list is less than the maximum number, or when the first merge candidate block is unavailable, merge candidates included in the second merge candidate list are added to the first merge candidate list.
[0275] Alternatively, when the first merge candidate block is not available, a merge candidate derived from a block adjacent to the first merge candidate block among the merge candidates included in the second merge candidate list is added to the first merge candidate list. Fig.15 , when A0 is not available, add the merge candidate derived based on the motion information of B0 among the merge candidates included in the second merge candidate list to the first merge candidate list. When A1 is not available, add the merge candidate derived based on the motion information of B1 among the merge candidates included in the second merge candidate list to the first merge candidate list. When A2 is not available, add the merge candidate derived based on the motion information of B2 among the merge candidates included in the second merge candidate list to the first merge candidate list. When A3 is not available, add the merge candidate derived based on the motion information of B3 among the merge candidates included in the second merge candidate list to the first merge candidate list. When A4 is not available, add the merge candidate derived based on the motion information of B4, B5 or B6 among the merge candidates included in the second merge candidate list to the first merge candidate list.
[0276] Alternatively, the merge candidate to be added to the first merge candidate list may be determined according to the priority of the merge candidates included in the second merge candidate list. The priority may be determined based on the index value assigned to each merge candidate. For example, when the number of merge candidates included in the first merge candidate list is less than the maximum number, or when the first merge candidate block is unavailable, the merge candidate with the smallest index value or the merge candidate with the largest index value among the merge candidates included in the second merge candidate list is added to the first merge candidate list.
[0277] When a merge candidate having the same motion information as a merge candidate with the highest priority among the merge candidates included in the second merge candidate list is included in the first merge candidate list, the merge candidate with the highest priority may not be added to the first merge candidate list. In addition, it may be determined whether a merge candidate with the next priority (for example, a merge candidate to which an index value greater than the index value assigned to the merge candidate with the highest priority is assigned by 1 or a merge candidate to which an index value less than the index value assigned to the merge candidate with the highest priority is assigned by 1) can be added to the first merge candidate list.
[0278] Alternatively, a merge candidate list including merge candidates derived based on motion information of the first merge candidate block and merge candidates derived based on motion information of the second merge candidate block may be generated. The merge candidate list may be a combination of the first merge candidate list and the second merge candidate list.
[0279] For example, according to a predetermined search order, a merge candidate list may be generated by searching the first merge candidate block and the second merge candidate block.
[0280] Figures 17 to 20 is a diagram showing the order of searching for merge candidate blocks.
[0281] Figures 17 to 20 The order of searching for merge candidates is shown as follows.
[0282] A0→A1→A2→A3→A4→B0→B1→B2→B3→B4→(B5)→(B6).
[0283] Only when block B4 is unavailable or when the number of merge candidates included in the merge candidate list is equal to or less than a preset number, the search for blocks B5 and B6 is performed.
[0284] Can be set with Figures 17 to 20 The examples shown in the different search orders.
[0285] A merge candidate list may be generated that includes a combination of at least one merge candidate included in a first merge candidate list and at least one merge candidate included in a second merge candidate list. For example, the combined merge candidate list may include N merge candidates among the merge candidates included in the first merge candidate list and M merge candidates among the merge candidates included in the second merge candidate list. The letters N and M may represent the same number or different numbers. Alternatively, at least one of N and M may be determined based on at least one of the number of merge candidates included in the first merge candidate list and the number of merge candidates included in the second merge candidate list. Alternatively, information for determining at least one of N and M may be signaled via a bitstream. Any one of N and M may be derived by subtracting the other of N and M from the maximum number of merge candidates in the combined merge candidate list.
[0286] The merge candidates to be added to the combined merge candidate list may be determined according to a predefined priority. The predefined priority may be determined based on an index assigned to the merge candidate.
[0287] Alternatively, the merge candidates to be added to the combined merge candidate list may be determined based on the association between the merge candidates. For example, when A0 included in the first merge candidate list is added to the combined merge candidate list, a merge candidate (e.g., B0) at a position adjacent to A0 is not added to the combined merge list.
[0288] When the number of merge candidates included in the first merge candidate list is less than N, more than M merge candidates among the merge candidates included in the second merge candidate list are added to the combined merge candidate list. For example, when N is 4 and M is 2, 4 merge candidates among the merge candidates included in the first merge candidate list are added to the combined merge candidate list, and 2 merge candidates among the merge candidates included in the second merge candidate list are added to the combined merge candidate list. When the number of merge candidates included in the first merge candidate list is less than 4, two or more merge candidates among the merge candidates included in the second merge candidate list are added to the combined merge candidate list. When the number of merge candidates included in the second merge candidate list is less than 2, four or more merge candidates among the merge candidates included in the first merge candidate list are added to the combined merge candidate list.
[0289] That is, the value of N or M may be adjusted according to the number of merge candidates included in each merge candidate list. By adjusting the value of N or M, the total number of merge candidates included in the combined merge candidate list may be fixed. When the total number of merge candidates included in the combined merge candidate list is less than the maximum number of merge candidates, a combined merge candidate, an average merge candidate, or a zero motion vector candidate is added.
[0290] Motion compensation of the current block may be performed by using at least one of the merge candidates included in the first merge candidate list and the second merge candidate list. The encoder may encode index information for specifying any one of the plurality of merge candidates. In an example, 'merge_idx' may specify any one of the plurality of merge candidates. In an example, Table 1 represents the index information for specifying any one of the plurality of merge candidates. Fig.16 The merge index of each merge candidate derived from the first merge candidate block and the second merge candidate block shown in .
[0291]
Table 1
[0292]
[0293]
[0294] However, as the number of merge candidates included in the merge candidate list increases, the codeword used to encode the merge index becomes longer. Therefore, there is a problem of reducing encoding / decoding efficiency. In order to reduce the length of the codeword, the merge index can be determined by using a prefix and a suffix. In the example, the merge index can be determined by using merge_idx_prefix representing the prefix of the merge index and merge_idx_suffix representing the suffix of the merge index.
[0295] Table 2 shows the merged index prefix value and the merged index suffix value of each merged index, and Table 3 shows the process of determining the merged index based on the merged index prefix value and the merged index suffix value.
[0296]
Table 2
[0297]
[0298]
[0299]
Table 3
[0300]
[0301] As shown in Tables 2 and 3, when the merge index prefix value is less than the threshold value, the merge index may be set to be the same as the merge index prefix value. On the other hand, when the merge index prefix value is greater than the threshold value, the merge index may be determined by subtracting a reference value from the merge index prefix and adding a merge index suffix to a value that shifts the result. The reference value may be a threshold value, or may be a value obtained by subtracting 1 from the threshold value.
[0302] Tables 2 and 3 show that the threshold is 4. The threshold may be determined based on at least one of the number of merge candidates included in the merge candidate list, the number of second merge candidate blocks, or the number of lines including the second merge candidate blocks. Alternatively, the threshold may be predefined in the encoder and the decoder.
[0303] Whether to use the prefix and suffix to determine the merge index may be determined according to the number of merge candidates included in the merge candidate list or the maximum number of merge candidates that can be included in the merge candidate list. In an example, when the maximum number of merge candidates that can be included in the merge candidate list is greater than a threshold, a merge index prefix and a merge index suffix used to determine the merge index may be sent with a signal. On the other hand, when the maximum number of merge candidates is less than the threshold, the merge index may be sent with a signal.
[0304] The rectangular block may be divided into a plurality of triangular blocks. Merging candidates of the triangular blocks may be derived based on the rectangular block including the triangular blocks. The triangular blocks may share the same merging candidate.
[0305] A merge index may be signaled for each triangle block. In this case, the triangle blocks may be set to not use the same merge candidate. In an example, a merge candidate for a first triangle block may not be used as a merge candidate for a second triangle block. Thus, the merge index for the second triangle block may specify any one of the remaining merge candidates other than the merge candidate selected for the first triangle block.
[0306] The merge candidate may be derived based on a block having a predetermined shape or a predetermined size or larger. When the current block is not in the predetermined shape, or when the size of the current block is smaller than the predetermined size, a merge candidate of the current block is derived based on a block including the current block and in the predetermined shape or in the predetermined size or larger. The predetermined shape may be a square shape or a non-square shape.
[0307] When the predetermined shape is a square shape, a merge candidate of the non-square shaped coding unit is derived based on the square shaped coding units including the non-square shaped coding units.
[0308] Fig.21 is a diagram showing an example of deriving merging candidates of non-square blocks based on square blocks.
[0309] Merge candidates for non-square blocks can be derived based on square blocks including non-square blocks. For example, merge candidates for non-square shaped coding block 0 and non-square shaped coding block 1 can be derived based on square shaped blocks including coding block 0 and coding block 1. That is, the position of the spatially adjacent blocks can be determined based on the position, width / height or size of the square shaped blocks. Merge candidates for coding block 0 and coding block 1 can be derived based on at least one of the spatially adjacent blocks A0, A1, A2, A3 and A4 adjacent to the square shaped blocks.
[0310] Temporal merge candidates may be determined based on square-shaped blocks. That is, temporal neighboring blocks may be determined based on the position, width / height, or size of the square-shaped blocks. For example, merge candidates for coding block 0 and coding block 1 may be derived based on temporal neighboring blocks determined based on square-shaped blocks.
[0311] Alternatively, any one of the spatial merge candidate and the temporal merge candidate may be derived based on a square block, and the other merge candidate may be derived based on a non-square block. For example, a spatial merge candidate for coding block 0 may be derived based on a square block, and a temporal merge candidate for coding block 0 may be derived based on coding block 0.
[0312] A plurality of blocks included in a block of a predetermined shape or a predetermined size or larger may share a merge candidate. Fig.21In the example shown in , at least one of the spatial merging candidates and the temporal merging candidates of coding block 0 and coding block 1 may be the same.
[0313] The predetermined shape may be a non-square shape, such as 2NxN, Nx2N, etc. When the predetermined shape is a non-square shape, a merge candidate of the current block may be derived based on a non-square block including the current block. For example, when the current block is in a 2Nxn shape (herein, n is 1 / 2N), a merge candidate of the current block is derived based on a non-square block of a 2NxN shape. Alternatively, when the current block is in an nx2N shape, a merge candidate of the current block is derived based on a non-square block of an Nx2N shape.
[0314] Information indicating a predetermined shape or a predetermined size may be signaled through a bitstream. For example, information indicating any one of a non-square shape or a square shape may be signaled through a bitstream.
[0315] Alternatively, the predetermined shape or the predetermined size may be determined according to a rule predefined in an encoder and a decoder.
[0316] When the child node does not meet the predetermined condition, a merge candidate of the child node is derived based on the parent node that meets the predetermined condition. In this article, the predetermined condition may include at least one of the following: whether the block is a block generated due to quadtree partitioning, whether the size of the block is exceeded, the shape of the block and the picture boundary, and whether the depth difference between the child node and the parent node is equal to or greater than a predetermined value.
[0317] For example, the predetermined condition may include whether the block is a block generated due to quadtree partitioning, and whether the block is a coded block of a square shape of a predetermined size or larger. When the current block is generated by binary tree partitioning or ternary tree partitioning, a merge candidate of the current block is derived based on a high-level node block that includes the current block and satisfies the predetermined condition. When there is no high-level node block that satisfies the predetermined condition, a merge candidate of the current block is derived based on the current block, a block that includes the current block and is of a predetermined size or larger, or a high-level node block that includes the current block and has a depth difference of one with the current block.
[0318] Fig. 22 is a diagram showing an example of deriving merge candidates based on high-level node blocks.
[0319] Block 0 and block 1 are generated by partitioning a square block based on a binary tree. Merge candidates of block 0 and block 1 may be derived based on neighboring blocks (i.e., at least one of A0, A1, A2, A3, and A4) determined from high-level node blocks including block 0 and block 1. Therefore, block 0 and block 1 may use the same spatial merge candidate.
[0320] A high-level node block including block 2, block 3, and block 4 may be generated by partitioning a square block based on a binary tree. In addition, block 2 and block 3 may be generated by partitioning a non-square shaped block based on a binary tree. Merge candidates of non-square shaped blocks 2, block 3, and block 4 may be derived based on a high-level node block including non-square shaped blocks 2, block 3, and block 4. That is, a merge candidate may be derived based on an adjacent block (e.g., at least one of B0, B1, B2, B3, and B4) determined according to the position, width / height, or size of the square block including blocks 2, block 3, and block 4. Therefore, blocks 2, block 3, and block 4 may use the same spatial merge candidate.
[0321] Temporal merge candidates for non-square shaped blocks may be derived based on high-level node blocks. For example, temporal merge candidates for blocks 0 and 1 may be derived based on a square block including blocks 0 and 1. Temporal merge candidates for blocks 2, 3, and 4 may be derived based on a square block including blocks 2, 3, and 4. In addition, the same temporal merge candidates derived from temporal neighboring blocks determined based on each quadtree block may be used.
[0322] The lower-level node blocks included in the higher-level node block may share at least one of the spatial merge candidate and the temporal merge candidate. For example, the lower-level node blocks included in the higher-level node block may use the same merge candidate list.
[0323] Alternatively, at least one of the spatial merge candidate and the temporal merge candidate may be derived based on the low-level node block, and the other may be derived based on the high-level node block. For example, the spatial merge candidate of block 0 and block 1 may be derived based on the high-level node block. However, the temporal merge candidate of block 0 may be derived based on block 0, and the temporal merge candidate of block 1 may be derived based on block 1.
[0324] Alternatively, when the number of samples included in the low-level node block is less than a predefined number, a merge candidate is derived based on a high-level node block including a predefined number or more of samples. For example, when at least one of the following conditions is met: at least one of the low-level node blocks generated based on at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning is less than a preset size; at least one of the low-level node blocks is a non-square block; the high-level node block does not exceed the picture boundary; and the width or height of the high-level node block is equal to or greater than a predefined value, a merge candidate is derived based on a square or non-square high-level node block including a predefined number or more of samples (e.g., 64, 128, or 256 samples). The low-level node blocks included in the high-level node block can share the merge candidates derived based on the high-level node block.
[0325] A merge candidate can be derived based on any one of the low-level node blocks, and other low-level node blocks can be set to use the merge candidate. The low-level node blocks may be included in blocks of a predetermined shape or a predetermined size or larger. For example, the low-level node blocks may share a merge candidate list derived based on any one of the low-level node blocks. Information of the low-level node blocks serving as the basis for derivation of the merge candidate may be signaled via a bitstream. The information may be index information indicating any one of the low-level node blocks. Alternatively, the low-level node blocks serving as the basis for derivation of the merge candidate may be determined based on at least one of the position, size, shape, and scanning order of the low-level node blocks.
[0326] Information indicating whether a low-level node block shares a merge candidate list derived based on a high-level node block may be signaled through a bitstream. Based on this information, it may be determined whether a merge candidate for a block that is not in a predetermined shape or a block that is less than a predetermined size is derived based on a high-level node block including the block. Alternatively, it may be determined whether to derive a merge candidate based on a high-level node block according to rules predefined in an encoder and a decoder.
[0327] When there is a neighboring block adjacent to the current block within a predefined area, it is determined that the neighboring block is not available as a spatial merging candidate. The predefined area may be a parallel processing area defined for parallel processing between blocks. The parallel processing area may be referred to as a merge estimation region (MER). For example, when a neighboring block adjacent to the current block is included in the same merge estimation region as the current block, it is determined that the neighboring block is not available. A shift operation may be performed to determine whether the current block and the neighboring block are included in the same merge estimation region. Specifically, based on whether the value obtained by shifting the position of the upper left reference sample of the current block is the same as the value obtained by shifting the position of the upper left reference sample of the adjacent block, it may be determined whether the current block and the adjacent block are included in the same merge estimation region.
[0328] Fig.23 is a diagram showing an example of determining the availability of spatially neighboring blocks based on a merged estimated region.
[0329] exist Fig.23 In FIG. 5 , it is shown that the merged estimation area is in an Nx2N shape.
[0330] A merge candidate for block 1 may be derived based on a spatial neighboring block adjacent to block 1. The spatial neighboring blocks may include B0, B1, B2, B3, and B4. Herein, it may be determined that the spatial neighboring blocks B0 and B3 included in the same merge estimation region as block 1 are not available as merge candidates. Therefore, a merge candidate for block 1 may be derived from at least one of the spatial neighboring blocks B1, B2, and B4 other than the spatial neighboring blocks B0 and B3.
[0331] A merge candidate for block 3 may be derived based on a spatial neighboring block adjacent to block 3. The spatial neighboring blocks may include C0, C1, C2, C3, and C4. Herein, it may be determined that the spatial neighboring block C0 included in the same merge estimation region as block 3 is not available as a merge candidate. Therefore, a merge candidate for block 3 may be derived from at least one of the spatial neighboring blocks C1, C2, C3, and C4 other than the spatial neighboring block C0.
[0332] Based on at least one of the position, size, width, and height of the merge estimation region, a merge candidate of a block included in the merge estimation region may be derived. For example, a merge candidate of a plurality of blocks included in the merge estimation region may be derived from at least one of a spatial neighboring block and a temporal neighboring block determined based on at least one of the position, size, width, and height of the merge estimation region. The blocks included in the merge estimation region may share the same merge candidate.
[0333] Fig.24 is a diagram showing an example of deriving a merge candidate based on a merge estimation region.
[0334] When a plurality of coding units are included in a merge estimation region, merge candidates of the plurality of coding units may be derived based on the merge estimation region. That is, by using the merge estimation region as a coding unit, merge candidates may be derived based on the position, size, or width / height of the merge estimation region.
[0335] For example, merge candidates of both coding unit 0 (CU0) and coding unit 1 (CU1), which are both in (n / 2)xN (herein, n is N / 2) size and included in the merge estimation region in the (N / 2)xN size, may be derived based on the merge estimation region. That is, merge candidates of coding unit 0 and coding unit 1 may be derived based on at least one of neighboring blocks C0, C1, C2, C3, and C4 adjacent to the merge estimation region.
[0336] For example, merging candidates of coding unit 2 (CU2), coding unit 3 (CU3), coding unit 4 (CU4), and coding unit 5 (CU5) in an nxn size included in a merging estimation region in an NxN size may be derived based on the merging estimation region. That is, merging candidates of coding unit 2, coding unit 3, coding unit 4, and coding unit 5 may be derived based on at least one of neighboring blocks C0, C1, C2, C3, and C4 adjacent to the merging estimation region.
[0337] The shape of the merged estimation area may be a square shape or a non-square shape. For example, it may be determined that a square-shaped coding unit (or prediction unit) or a non-square-shaped coding unit (or prediction unit) is a merged estimation area. The ratio between the width and height of the merged estimation area may be limited to not exceed a predetermined range. For example, the merged estimation area cannot have a non-square shape with a ratio between width and height exceeding 2, or a non-square shape with a ratio between width and height less than 1 / 2. That is, the non-square merged estimation area may be in the shape of 2NxN or Nx2N. Information about the limitation on the ratio between width and height may be signaled via a bitstream. Alternatively, the limitation on the ratio between width and height may be predefined in the encoder and decoder.
[0338] At least one of information indicating the shape of the merged estimation region and information indicating the size of the merged estimation region may be signaled through the bitstream. For example, at least one of information indicating the shape of the merged estimation region and information indicating the size of the merged estimation region may be signaled through a slice header, a tile group header, a picture parameter, or a sequence parameter.
[0339] The shape of the merged estimation area or the size of the merged estimation area may be updated on a per-sequence basis, on a per-picture basis, on a per-slice basis, on a per-tile group basis, on a per-tile basis, or on a per-block (CTU) basis. When the shape of the merged estimation area or the size of the merged estimation area is different from the shape or size of the previous unit, information indicating the new shape of the merged estimation area or the new size of the merged estimation area is signaled through the bitstream.
[0340] At least one block may be included in the merged estimation area. The blocks included in the merged estimation area may be in a square shape or a non-square shape. The maximum number or minimum number of blocks that the merged estimation area can include may be determined. For example, three, four or more CUs may be included in the merged estimation area. The determination may be based on information signaled by the bitstream. Alternatively, the maximum number or minimum number of blocks that the merged estimation area can include may be predefined in the encoder and the decoder.
[0341] In at least one of a case where the number of blocks included in the merge estimation region is less than the maximum number and a case where the number is greater than the minimum number, parallel processing of the blocks may be allowed. For example, when the number of blocks included in the merge estimation region is equal to or less than the maximum number, or when the number of blocks included in the merge estimation region is equal to or greater than the minimum number, a merge candidate of the block is derived based on the merge estimation region. When the number of blocks included in the merge estimation region is greater than the maximum number, or when the number of blocks included in the merge estimation region is less than the minimum number, a merge candidate of each block in the block is derived based on the size, position, width, or height of each block in the block.
[0342] The information indicating the shape of the merged estimation region may include a 1-bit flag. For example, the syntax "isrectagular_mer_flag" may indicate a square-shaped or non-square-shaped merged candidate region. An isrectagular_mer_flag value of 1 may indicate a non-square-shaped merged estimation region, while an isrectagular_mer_flag value of 0 may indicate a square-shaped merged estimation region.
[0343] When the information indicates a non-square shaped merged estimation region, information indicating at least one of a width, a height, and a ratio between the width and the height of the merged estimation region is signaled through the bitstream. Based on this, the size and / or shape of the merged estimation region can be determined.
[0344] It is within the scope of the present invention to apply the embodiments described with emphasis on the decoding process or the encoding process to the encoding process or the decoding process. It is also within the scope of the present invention to change the embodiments described in a predetermined order to a different order.
[0345] Although the above embodiments have been described based on a series of steps or flow charts, the above embodiments are not intended to limit the time series order of the present invention, and can be executed simultaneously or in different orders. In addition, each of the components (e.g., units, modules, etc.) constituting the block diagram in the above embodiments can be implemented as a hardware device or software, and multiple components can be combined into one hardware device or software. The above embodiments can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. Computer-readable storage media can include program instructions, data files, data structures, etc., alone or in combination. Examples of computer-readable storage media include: magnetic recording media such as hard disks, floppy disks, and tapes; optical data storage media such as CD-ROMs or DVD-ROMs; magneto-optical media such as optical disks; and hardware devices such as read-only memories (ROMs), random access memories (RAMs), and flash memories, which are particularly structured to store and implement program instructions. Hardware devices can be configured to be operated by one or more software modules or software modules can be configured to be operated by one or more hardware devices to perform processing according to the present invention.
[0346] Industrial Applicability
[0347] The present invention can be applied to electronic devices capable of encoding / decoding images.
[0348] Additionally, the present technology may also be configured as follows.
[0349] (1) A method for decoding a video, the method comprising:
[0350] Deriving merge candidates based on neighboring blocks adjacent to the current block;
[0351] generating a first merge candidate list including the merge candidate,
[0352] Wherein, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, adding the merge candidates included in the second merge candidate list to the first merge candidate list;
[0353] decoding information for specifying one of the merge candidates included in the first merge candidate list; and
[0354] The motion information of the current block is derived according to the merge candidate assigned with the index determined by the information.
[0355] (2) The method according to (1), wherein the information includes an index prefix and an index suffix.
[0356] (3) The method according to (2), wherein when the value of the index prefix is less than a threshold value, the index is set to be the same as the index prefix.
[0357] (4) The method according to (3), wherein, when the value of the index prefix is greater than the threshold, the index is derived by adding the value of the index suffix to a value derived based on the index prefix.
[0358] (5) The method according to (3), wherein the threshold is determined based on the number of merge candidates included in the first merge candidate list.
[0359] (6) The method according to (1), wherein the second merge candidate list includes merge candidates derived from blocks that are not adjacent to the current block.
[0360] (7) The method according to (6), wherein the non-neighboring block is on the same line as a block adjacent to the current block.
[0361] (8) A method for encoding a video, the method comprising:
[0362] Deriving merge candidates based on neighboring blocks adjacent to the current block;
[0363] generating a first merge candidate list including the merge candidate,
[0364] Wherein, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, adding the merge candidates included in the second merge candidate list to the first merge candidate list;
[0365] encoding information for specifying one of the merge candidates included in the first merge candidate list; and
[0366] The motion information of the current block is derived according to the merge candidate assigned with the index determined by the information.
[0367] (9) The method according to (8), wherein the information includes an index prefix and an index suffix.
[0368] (10) The method according to (9), wherein when the value of the index prefix is less than a threshold value, the index is set to be the same as the index prefix.
[0369] (11) The method according to (10), wherein when the value of the index prefix is greater than the threshold, the index is derived by adding the value of the index suffix to a value derived based on the index prefix.
[0370] (12) The method according to (10), wherein the threshold is determined based on the number of merge candidates included in the first merge candidate list.
[0371] (13) The method according to (8), wherein the second merge candidate list includes merge candidates derived from blocks that are not adjacent to the current block.
[0372] (14) The method according to (13), wherein the non-neighboring block is on the same line as a block adjacent to the current block.
[0373] (15) A device for decoding an image, the device comprising:
[0374] a decoding unit that decodes information specifying one of the merge candidates included in the first merge candidate list; and
[0375] an inter-frame prediction unit that derives a merge candidate from a neighboring block adjacent to a current block, generates a first merge candidate list including the merge candidate, and derives motion information of the current block from the merge candidate to which an index determined by the information is assigned,
[0376] When the number of merge candidates included in the first merge candidate list is less than a predetermined value, the merge candidates included in the second merge candidate list are added to the first merge candidate list.
Claims
1. A method for decoding a video, the method comprises: deriving at least one spatial merge candidate of a current coding block, the spatial merge candidate being derived based on neighboring blocks adjacent to the current coding block; deriving a temporal merge candidate of the current coding block; generating a first candidate list including at least one spatial merge candidate and the temporal merge candidate of the current coding block, wherein, when the number of merge candidates included in the first candidate list is less than a predetermined value, adding motion information candidates included in a second candidate list to the first candidate list; selecting two merge candidates from multiple merge candidates included in the first candidate list based on two index information for two triangular shape partitions in the current coding block; obtaining a prediction block of the current coding block based on two motion information derived according to the two merge candidates; and reconstructing the current coding block based on the prediction block, wherein the motion information candidates included in the second candidate list are derived based on spatial blocks, the spatial blocks being decoded before decoding the current coding block and included in a current picture including the current coding block, wherein a first index information of the two index information specifies a first merge candidate among the multiple merge candidates, wherein a second index information of the two index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate, and wherein both the first index information and the second index information are composed of only a prefix part.
2. The method according to claim 1, wherein, updating the second candidate list with motion information of a block encoded by inter prediction coding, and wherein adding the motion information of the block as a new motion information candidate with the highest or lowest index to the second candidate list.
3. The method according to claim 2, wherein, if a first motion information candidate identical to the motion information of the block already exists in the second candidate list, deleting the first motion information candidate from the second candidate list and adding the motion information of the block as the new motion information candidate to the second candidate list.
4. The method according to claim 1, wherein, initializing the second candidate list in a predefined unit, and wherein the predefined unit is a slice, a tile, a coding tree unit or a row of coding tree units.
5. The method according to claim 1, wherein, at least one of multiple motion information candidates in the second candidate list is added to the first candidate list in descending order of indexes assigned to each of the multiple motion information candidates.
6. A method for encoding a video, the method comprises: deriving at least one spatial merge candidate of a current coding block, the spatial merge candidate being derived based on neighboring blocks adjacent to the current coding block; deriving a temporal merge candidate of the current coding block; generating a first candidate list including at least one spatial merge candidate and the temporal merge candidate of the current coding block, Wherein, when the number of merge candidates included in the first candidate list is less than a predetermined value, motion information candidates included in the second candidate list are added to the first candidate list; Select two merge candidates from among the multiple merge candidates included in the first candidate list for partitioning two triangular shapes in the current coding block; Obtain a prediction block for the current coding block based on two motion information derived from the two merge candidates; Obtain a residual block for the current coding block based on the prediction block; and Encode two pieces of index information, wherein the motion information candidates included in the second candidate list are derived from spatial blocks that are encoded before the current coding block and are included in the current picture including the current coding block, wherein a first index information among the two pieces of index information specifies a first merge candidate among the multiple merge candidates, wherein a second index information among the two pieces of index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate, and wherein both the first index information and the second index information are encoded only by a prefix part.
7. An apparatus for transmitting compressed video data, the apparatus comprising: one or more hardware units configured to: obtain the compressed video data; and transmit the compressed video data, wherein obtaining the compressed video data includes: deriving at least one spatial merge candidate for a current coding block, the spatial merge candidate being derived from neighboring blocks adjacent to the current coding block; deriving a temporal merge candidate for the current coding block; generating a first candidate list including at least one spatial merge candidate and a temporal merge candidate for the current coding block, wherein, when the number of merge candidates included in the first candidate list is less than a predetermined value, motion information candidates included in the second candidate list are added to the first candidate list; selecting two merge candidates from among the multiple merge candidates included in the first candidate list for partitioning two triangular shapes in the current coding block; obtaining a prediction block for the current coding block based on two motion information derived from the two merge candidates; obtaining a residual block for the current coding block based on the prediction block; and encoding two pieces of index information, wherein the motion information candidates included in the second candidate list are derived from spatial blocks that are encoded before the current coding block and are included in the current picture including the current coding block, wherein a first index information among the two pieces of index information specifies a first merge candidate among the multiple merge candidates, wherein a second index information among the two pieces of index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate, and wherein both the first index information and the second index information are encoded only by a prefix part.