Method for decoding and encoding video, and apparatus for transmitting compressed video data
By using multiple merge candidate lists for motion compensation and merge index encoding/decoding in video signal encoding/decoding, the problem of low inter prediction efficiency in the prior art is solved, and a more efficient encoding/decoding process is realized.
Patent Information
- Application Number
- CN202510425092.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-23
- Filing Date
- 2019-05-23
- Publication Date
- 2025-05-30
AI Technical Summary
In encoding/decoding a video signal, it is difficult for the prior art to effectively perform inter prediction, especially in motion compensation and merged index encoding/decoding.
Motion compensation is performed by using a plurality of merge candidate lists and a first merge candidate list and a second merge candidate list including merge candidates are generated. During the decoding process, merge candidates are derived from adjacent blocks adjacent to the current block, and motion information candidates are added to the merge candidate list to improve the efficiency of inter-frame prediction.
The efficiency of inter-frame prediction is improved, and through the use of multiple merge candidate lists, the merged index can be more efficiently encoded/decoded, reducing the cost in the encoding/decoding process.
Smart Images

Figure CN120075465A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application No. 201980034764.3, with the international application date of May 23, 2019, the international application number of PCT / KR2019 / 006216, and the invention name of "Method and Apparatus for Processing Video Signals", which entered the Chinese national phase on November 23, 2020. Technical Field
[0002] The present invention relates to a method and apparatus for processing video signals. Background Art
[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) images and ultra-high-definition (UHD) images, has increased in various application fields. However, compared with conventional image data, image data with higher resolution and quality has an increased data volume. Therefore, when transmitting image data by using a medium such as a conventional wired and wireless broadband network, or when storing image data by using a conventional storage medium, the cost of transmission and storage increases. To solve these problems that occur as the resolution and quality of image data increase, efficient image encoding / decoding techniques can be utilized.
[0004] Image compression techniques include various techniques, including: an inter-frame prediction technique for predicting pixel values included in a current picture based on a previous picture or a subsequent picture of the current picture; an intra-frame prediction technique for predicting pixel values included in a current picture by using pixel information in the current picture; an entropy encoding technique for assigning a short code to a value with a high occurrence frequency and assigning a long code to a value with a low occurrence frequency, and the like. Image data can be effectively compressed by using such image compression techniques and can be transmitted or stored.
[0005] Meanwhile, with the demand for high-resolution images, the demand for stereoscopic image content as a new image service has also increased. Video compression techniques for effectively providing stereoscopic image content with high resolution and ultra-high resolution are being studied. Summary of the Invention
[0006] Technical problem
[0007] The present invention provides a method and apparatus for effectively performing inter-frame prediction on an encoding / decoding target block when encoding / decoding a video signal.
[0008] The present invention provides a method and apparatus for performing motion compensation by using a plurality of merge candidate lists when encoding / decoding a video signal.
[0009] The present invention provides a method and apparatus for efficiently encoding / decoding a merge index when encoding / decoding a video signal.
[0010] Technical problems that can be obtained from the present invention are not limited to the above-mentioned technical tasks, and other unmentioned technical tasks can be clearly understood by those of ordinary skill in the technical field to which the present invention pertains based on the following description.
[0011] Technical solution
[0012] The video signal decoding method and apparatus according to the present invention can derive merge candidates based on neighboring blocks adjacent to a current block, generate a first merge candidate list including the merge candidates, decode information for specifying one of the merge candidates included in the first merge candidate list, and derive motion information of the current block based on the merge candidate assigned the index determined by the information. In this case, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, merge candidates included in a second merge candidate list can be added to the first merge candidate list.
[0013] The video signal encoding method and apparatus according to the present invention can derive merge candidates based on neighboring blocks adjacent to a current block, generate a first merge candidate list including the merge candidates, encode information for specifying one of the merge candidates included in the first merge candidate list, and derive motion information of the current block based on the merge candidate assigned the index determined by the information. In this case, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, merge candidates included in a second merge candidate list can be added to the first merge candidate list.
[0014] For the video signal encoding / decoding method and apparatus according to the present invention, the information can include an index prefix and an index suffix.
[0015] For the video signal encoding / decoding method and apparatus according to the present invention, when the value of the index prefix is less than a threshold, the index can be set to be the same as the index prefix.
[0016] For the video signal encoding / decoding method and apparatus according to the present invention, when the value of the index prefix is greater than the threshold, the index can be derived by adding the value of the index suffix to a value derived based on the index prefix.
[0017] For the video signal encoding / decoding method and apparatus according to the present invention, the threshold can be determined based on the number of merge candidates included in the first merge candidate list.
[0018] For the video signal encoding / decoding method and apparatus according to the present invention, the second merge candidate list can include merge candidates derived based on blocks not adjacent to the current block.
[0019] For the video signal encoding / decoding method and apparatus according to the present invention, non-neighboring blocks may be on the same line as blocks neighboring a current block.
[0020] The method for decoding video according to the present invention includes: deriving merge candidates for a current encoded block, the merge candidates including spatial merge candidates derived based on neighboring blocks adjacent to the current encoded block; generating a first candidate list including the merge candidates for the current encoded block, wherein when the number of merge candidates included in the first candidate list is less than a predetermined value, adding motion information candidates included in a second candidate list to the first candidate list; selecting two merge candidates from among the plurality of merge candidates included in the first candidate list for two triangular shape partitions in the current encoded block; obtaining a predicted block for the current encoded block based on two motion information derived from the two merge candidates; obtaining a residual block for the current encoded block by inverse quantization and inverse transformation; and reconstructing the current encoded block by summing the predicted block and the residual block. The motion information candidates included in the second candidate list are derived based on spatial blocks that are decoded before the current encoded block is decoded and are included in the current picture including the current encoded block. A first index information among the two index information specifies a first merge candidate among the plurality of merge candidates. A second index information among the two index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate. Both the first index information and the second index information are composed of only a prefix part.
[0021] The method for encoding video according to the present invention includes: deriving merge candidates for a current encoded block, the merge candidates including spatial merge candidates derived based on neighboring blocks adjacent to the current encoded block; generating a first candidate list including the merge candidates for the current encoded block, wherein when the number of merge candidates included in the first candidate list is less than a predetermined value, adding motion information candidates included in a second candidate list to the first candidate list; selecting two merge candidates from among the plurality of merge candidates included in the first candidate list for two triangular shape partitions in the current encoded block; obtaining a predicted block for the current encoded block based on the two motion information derived from the two merge candidates; obtaining a residual block for the current encoded block based on the predicted block; obtaining residual coefficients by performing transformation and quantization on the residual block; and encoding two index information. The motion information candidates included in the second candidate list are derived based on spatial blocks that are encoded before the current encoded block is encoded and are included in the current picture including the current encoded block. A first index information among the two index information specifies a first merge candidate among the plurality of merge candidates. A second index information among the two index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate. Both the first index information and the second index information are encoded by only a prefix part.
[0022] A device for transmitting compressed video data according to the present invention includes one or more hardware units configured to: obtain compressed video data; and transmit the compressed video data. Obtaining the compressed video data includes: deriving merge candidates for a current coding block, the merge candidates including spatial merge candidates derived based on neighboring blocks adjacent to the current coding block; generating a first candidate list including the merge candidates for the current coding block, wherein when the number of merge candidates included in the first candidate list is less than a predetermined value, adding motion information candidates included in a second candidate list to the first candidate list; selecting two merge candidates from among the plurality of merge candidates included in the first candidate list for two triangular-shaped partitions in the current coding block; obtaining a prediction block for the current coding block based on two motion information derived from the two merge candidates; obtaining a residual block for the current coding block based on the prediction block; obtaining residual coefficients by performing transformation and quantization on the residual block; and encoding two pieces of index information. The motion information candidates included in the second candidate list are derived based on spatial blocks that were encoded before the current coding block and are included in the current picture including the current coding block. The first index information among the two pieces of index information specifies a first merge candidate among the plurality of merge candidates. The second index information among the two pieces of index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate. Both the first index information and the second index information are encoded only by a prefix portion.
[0023] It should be understood that the features summarized above are exemplary aspects of the following detailed description of the present invention and do not limit the scope of the present invention.
[0024] Beneficial effect
[0025] According to the present invention, the efficiency of inter-frame prediction can be improved by performing motion compensation by using multiple merge candidate lists.
[0026] According to the present invention, the efficiency of inter-frame prediction can be improved by obtaining motion information based on multiple merge candidates.
[0027] According to the present invention, an efficient encoding / decoding method for merge indexes can be provided.
[0028] The effects that can be obtained from the present invention may not be limited to the effects mentioned above, and those skilled in the art to which the present invention pertains can clearly understand other unmentioned effects from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a block diagram showing a device for encoding video according to an embodiment of the present invention.
[0030] Figure 2 is a block diagram showing an apparatus for decoding a video according to an embodiment of the present invention.
[0031] Figure 3 is a diagram showing partition mode candidates that can be applied to an encoded block when encoding the encoded block by inter prediction.
[0032] Figure 4 shows an example of hierarchical partitioning of an encoded block based on a tree structure to which an embodiment of the present invention is applied.
[0033] Figure 5 is a diagram showing a partition shape that allows binary tree-based partitioning to which an embodiment of the present invention is applied.
[0034] Figure 6 shows a ternary tree partition shape.
[0035] Figure 7 is a diagram showing an example of allowing only a specific shape of binary tree-based partitioning.
[0036] Figure 8 is a diagram for describing an example in which information related to the number of times of allowing binary tree partitioning is encoded / decoded according to an embodiment of the present invention.
[0037] Figure 9 is a flowchart showing an inter prediction method to which an embodiment of the present invention is applied.
[0038] Figure 10 is a diagram showing a process of deriving motion information of a current block when a merge mode is applied to the current block.
[0039] Figure 11 is a diagram showing an example of spatially adjacent blocks.
[0040] Figure 12 is a diagram showing an example of deriving motion vectors of temporal merge candidates.
[0041] Figure 13 is a diagram showing positions of candidate blocks that can be used as co-located blocks.
[0042] Figure 14 is a diagram showing a process of deriving motion information of a current block when an AMVP mode is applied to the current block.
[0043] Figure 15 is a diagram showing an example of deriving a merge candidate according to a second merge candidate when a first merge candidate block is not available.
[0044] Figure 16It is a diagram showing an example of deriving a merge candidate based on a second merge candidate block that is on the same line as the first merge candidate block.
[0045] Figures 17 to 20 It is a diagram showing the order of searching for merge candidate blocks.
[0046] Figure 21 It is a diagram showing an example of deriving a merge candidate for a non-square block based on a square block.
[0047] Figure 22 It is a diagram showing an example of deriving a merge candidate based on a high-level node block.
[0048] Figure 23 It is a diagram showing an example of determining the availability of spatially adjacent blocks based on a merge estimation region.
[0049] Figure 24 It is a diagram showing an example of deriving a merge candidate based on a merge estimation region. Detailed Description
[0050] Various modifications can be made to the present invention, and there are various embodiments of the present invention. Examples of these embodiments will now be provided with reference to the accompanying drawings and described in detail. However, the present invention is not limited thereto, and the exemplary embodiments can be construed as including all modifications, equivalents, or alternatives within the technical concept and technical scope of the present invention. In the described drawings, like reference numerals refer to like elements.
[0051] The terms 'first','second', etc. used in the specification may be used to describe various components, but these components should not be construed as being limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the present invention, the 'first' component may be referred to as the'second' component, and the'second' component may also be similarly referred to as the 'first' component. The term 'and / or' includes combinations of multiple items or any one of the multiple items.
[0052] In the present disclosure, when an element is referred to as being 'connected' or 'coupled' to another element, it should be understood to include not only that the element is directly connected or coupled to the other element, but also that there may be another element between them. When an element is referred to as being 'directly connected' or 'directly coupled' to another element, it should be understood that there are no other elements between them.
[0053] The terms used in this specification are only for describing specific embodiments and are not intended to limit the present invention. Expressions used in the singular form include those in the plural form unless the expression has a clearly different meaning in the context. In this specification, it should be understood that terms such as "including" and "having" are intended to indicate the existence of features, numbers, steps, actions, elements, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, parts, or combinations thereof may exist or may be added.
[0054] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. Hereinafter, the same constituent elements in the drawings are denoted by the same reference numerals, and repeated descriptions of the same elements will be omitted.
[0055] Figure 1 is a block diagram showing an apparatus for encoding a video according to an embodiment of the present invention.
[0056] Referring to Figure 1 , the apparatus 100 for encoding a video may include: a picture partitioning module 110, prediction modules 120 and 125, a transform module 130, a quantization module 135, a rearrangement module 160, an entropy encoding module 165, an inverse quantization module 140, an inverse transform module 145, a filter module 150, and a memory 155.
[0057] Figure 1 The components shown in are independently shown to represent different characteristic functions from each other in the apparatus for encoding a video. Therefore, this does not mean that each component is composed of a separate hardware or software component unit. In other words, for convenience, each component includes each of the listed components. Therefore, at least two components in each component may be combined to form one component, or one component may be divided into multiple components to perform each function. Embodiments of combining each component and embodiments of dividing one component are also included in the scope of the present invention without departing from the essence of the present invention.
[0058] In addition, some of the components may not be essential components for performing the basic functions of the present invention, but are only optional components for improving the performance of the present invention. The present invention can be implemented by excluding the components for improving performance and only including the essential components for implementing the essence of the present invention. The structure of excluding only the optional components for improving performance and only including the essential components is also included in the scope of the present invention.
[0059] The picture partitioning module 110 may partition an input picture into one or more processing units. Herein, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture partitioning module 110 may partition a picture into a combination of multiple coding units, prediction units, and transform units, and may select a combination of a coding unit, a prediction unit, and a transform unit by using a predetermined criterion (e.g., a cost function) to encode the picture.
[0060] For example, a picture may be partitioned into multiple coding units. A recursive tree structure such as a quadtree structure may be used to partition the picture into coding units. A coding unit that is rooted at a picture or a largest coding unit and is partitioned into other coding units may be partitioned into child nodes corresponding to the number of the partitioned coding units. A coding unit that cannot be further partitioned according to a predetermined limit is used as a leaf node. That is, when it is assumed that only square partitioning is feasible for a coding unit, a coding unit may be partitioned into at most four other coding units.
[0061] Hereinafter, in an embodiment of the present invention, the coding unit may mean a unit that performs encoding or a unit that performs decoding.
[0062] The prediction unit may be one of the partitions that are partitioned into a square shape or a rectangular shape having the same size within a single coding unit, or the prediction unit may be one of the partitions that are partitioned within a single coding unit such that they have different shapes / sizes.
[0063] When a prediction unit that undergoes intra prediction is generated based on a coding unit and the coding unit is not the smallest coding unit, intra prediction may be performed without partitioning the coding unit into multiple prediction units NxN.
[0064] The prediction modules 120 and 125 may include an inter - prediction module 120 that performs inter - frame prediction and an intra - prediction module 125 that performs intra - frame prediction. It can be determined whether to perform inter - frame prediction or intra - frame prediction for a prediction unit, and the details according to each prediction method (e.g., intra - frame prediction mode, motion vector, reference picture, etc.) can be determined. Here, the processing unit subjected to prediction may be different from the processing unit for which the prediction method and details are determined. For example, the prediction method, prediction mode, etc. can be determined by the prediction unit, and the prediction can be performed by the transform unit. The residual value (residual block) between the generated prediction block and the original block can be input to the transform module 130. In addition, the prediction mode information, motion vector information, etc. for prediction can be encoded by the entropy encoding module 165 together with the residual value and can be transmitted to the device for decoding the video. When using a specific coding mode, the original block can be transmitted to the device for decoding the video by being encoded as it is without generating a prediction block through the prediction modules 120 and 125.
[0065] The inter - prediction module 120 may predict a prediction unit based on information of at least one of a previous picture or a subsequent picture of the current picture, or in some cases, may predict a prediction unit based on information of some encoded regions in the current picture. The inter - prediction module 120 may include a reference picture interpolation module, a motion prediction module, and a motion compensation module.
[0066] The reference picture interpolation module may receive reference picture information from the memory 155 and may generate pixel information of whole pixels or sub - pixels based on the reference picture. In the case of luminance pixels, a DCT - based 8 - tap interpolation filter with different filter coefficients may be used to generate pixel information of whole pixels or sub - pixels in units of 1 / 4 pixels. In the case of chrominance signals, a DCT - based 4 - tap interpolation filter with different filter coefficients may be used to generate pixel information of whole pixels or sub - pixels in units of 1 / 8 pixels.
[0067] The motion prediction module may perform motion prediction based on the reference picture interpolated by the reference picture interpolation module. Various methods such as a full - search - based block - matching algorithm (FBMA), three - step search (TSS), new three - step search algorithm (NTS), etc. can be used as the method for calculating the motion vector. Based on the interpolated pixels, the motion vector may have a motion vector value in units of 1 / 2 pixels or 1 / 4 pixels. The motion prediction module may predict the current prediction unit by changing the motion prediction method. Various methods such as a skip method, a merge method, an AMVP (advanced motion vector prediction) method, an intra - block copy method, etc. can be used as the motion prediction method.
[0068] The intra prediction module 125 can generate a prediction unit based on reference pixel information adjacent to the current block, and the reference pixel information is pixel information in the current picture. When an adjacent block of the current prediction unit is a block subjected to inter prediction and thus the reference pixel is a pixel subjected to inter prediction, the reference pixel information of an adjacent block subjected to intra prediction can be used to replace the reference pixel included in the block subjected to inter prediction. That is, when the reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information.
[0069] The prediction mode of intra prediction can include a directional prediction mode using reference pixel information according to a prediction direction and a non-directional prediction mode that does not use direction information when performing prediction. The mode for predicting luminance information can be different from the mode for predicting chrominance information, and in order to predict chrominance information, the intra prediction mode information for predicting luminance information or the predicted luminance signal information can be utilized.
[0070] When performing intra prediction, when the size of the prediction unit is the same as the size of the transform unit, intra prediction can be performed on the prediction unit based on the pixels located on the left, upper left, and upper sides of the prediction unit. However, when performing intra prediction, when the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed using the reference pixels based on the transform unit. In addition, the intra prediction using NxN partitioning can be used only for the smallest coding unit.
[0071] In the intra prediction method, according to the prediction mode, a prediction block can be generated after applying an AIS (Adaptive Intra Smoothing) filter to the reference pixels. The type of the AIS filter applied to the reference pixels can be different. In order to perform the intra prediction method, the intra prediction mode of the current prediction unit can be predicted according to the intra prediction mode of the prediction unit adjacent to the current prediction unit. When predicting the prediction mode of the current prediction unit by using the mode information predicted according to the adjacent prediction unit, when the intra prediction mode of the current prediction unit is the same as the intra prediction mode of the adjacent prediction unit, predetermined flag information can be used to transmit information indicating that the prediction mode of the current prediction unit and the intra prediction mode of the adjacent prediction unit are the same as each other. When the prediction mode of the current prediction unit is different from the intra prediction mode of the adjacent prediction unit, entropy coding can be performed to encode the prediction mode information of the current block.
[0072] In addition, a residual block including information about a residual value, which is the difference between the predicted prediction unit and the original block of the prediction unit, can be generated based on the prediction units generated by the prediction modules 120 and 125. The generated residual block can be input to the transform module 130.
[0073] The transformation module 130 can transform a residual block including information on the residual value between an original block and a prediction unit generated by the prediction modules 120 and 125 by using transformation methods such as discrete cosine transform (DCT), discrete sine transform (DST), and KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on the intra prediction mode information of the prediction unit used to generate the residual block.
[0074] The quantization module 135 can quantize the values transformed to the frequency domain by the transformation module 130. The quantization coefficients can vary according to the blocks or importance of the picture. The values calculated by the quantization module 135 can be provided to the inverse quantization module 140 and the rearrangement module 160.
[0075] The rearrangement module 160 can rearrange the coefficients of the quantized residual values.
[0076] The rearrangement module 160 can change the coefficients in the form of a two-dimensional block into the form of a one-dimensional vector by a coefficient scanning method. For example, the rearrangement module 160 can use a zigzag scanning method to scan from the DC coefficient to the coefficients in the high-frequency domain so as to change the coefficients into the form of a one-dimensional vector. According to the size of the transformation unit and the intra prediction mode, a vertical direction scan that scans the coefficients in the form of a two-dimensional block in the column direction or a horizontal direction scan that scans the coefficients in the form of a two-dimensional block in the row direction can be used instead of the zigzag scan. That is, which scanning method among the zigzag scan, the vertical direction scan, and the horizontal direction scan is used can be determined according to the size of the transformation unit and the intra prediction mode.
[0077] The entropy encoding module 165 can perform entropy encoding based on the values calculated by the rearrangement module 160. The entropy encoding can use various encoding methods such as exponential Golomb coding, context adaptive variable length coding (CAVLC), and context adaptive binary arithmetic coding (CABAC).
[0078] The entropy encoding module 165 can encode various information from the rearrangement module 160 and the prediction modules 120 and 125, such as the residual value coefficient information and block type information of the coding unit, prediction mode information, partitioning unit information, prediction unit information, transformation unit information, motion vector information, reference frame information, block interpolation information, filtering information, etc.
[0079] The entropy encoding module 165 can perform entropy encoding on the coefficients of the coding unit input from the rearrangement module 160.
[0080] The inverse quantization module 140 may inverse-quantize the values quantized by the quantization module 135, and the inverse transform module 145 may inverse-transform the values transformed by the transform module 130. The residual values generated by the inverse quantization module 140 and the inverse transform module 145 may be combined with the prediction units predicted by the motion estimation module, the motion compensation module, and the intra prediction module of the prediction modules 120 and 125, such that a reconstructed block may be generated.
[0081] The filter module 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).
[0082] The deblocking filter may remove block distortion that appears due to the boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, pixels in several rows or columns included in a block may be a basis for determining whether to apply the deblocking filter to the current block. When the deblocking filter is applied to a block, a strong filter or a weak filter may be applied according to the required deblocking filtering strength. In addition, when applying the deblocking filter, horizontal direction filtering and vertical direction filtering may be processed in parallel.
[0083] The offset correction module may correct an offset from an original picture in pixel units in a picture that has undergone deblocking. To perform offset correction on a specific picture, a method of applying an offset considering edge information of each pixel or the following method may be used: dividing pixels of the picture into a predetermined number of regions, determining regions to undergo offset application, and applying an offset to the determined regions.
[0084] Adaptive loop filtering (ALF) may be performed based on values obtained by comparing a filtered reconstructed picture with an original picture. Pixels included in a picture may be divided into predetermined groups, filters to be applied to each of the groups may be determined, and filtering may be performed separately for each group. Information on whether to apply ALF and a luminance signal may be transmitted through a coding unit (CU). The shape and filter coefficients of the filter for ALF may be different according to each block. In addition, a filter for ALF having the same shape (fixed shape) may be applied regardless of the characteristics of an application target block.
[0085] The memory 155 may store the reconstructed block or the reconstructed picture calculated by the filter module 150. The stored reconstructed block or reconstructed picture may be provided to the prediction modules 120 and 125 when performing inter prediction.
[0086] Figure 2 is a block diagram showing an apparatus for decoding video according to an embodiment of the present invention.
[0087] Refer to Figure 2The apparatus 200 for decoding a video may include: an entropy decoding module 210, a rearrangement module 215, an inverse quantization module 220, an inverse transform module 225, prediction modules 230 and 235, a filter module 240, and a memory 245.
[0088] When a video bitstream is input from an apparatus for encoding a video, the input bitstream may be decoded according to the inverse processing of the apparatus for encoding the video.
[0089] The entropy decoding module 210 may perform entropy decoding according to the inverse processing of the entropy encoding performed by the entropy encoding module of the apparatus for encoding the video. For example, corresponding to the method performed by the apparatus for encoding the video, various methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC) may be applied.
[0090] The entropy decoding module 210 may decode information regarding intra prediction and inter prediction performed by the apparatus for encoding the video.
[0091] The rearrangement module 215 may perform rearrangement on the bitstream entropy decoded by the entropy decoding module 210 based on the rearrangement method used in the apparatus for encoding the video. The rearrangement module may reconstruct and rearrange coefficients in the form of a one-dimensional vector into coefficients in the form of a two-dimensional block. The rearrangement module 215 may receive information related to coefficient scanning performed in the apparatus for encoding the video, and may perform rearrangement via a method of inverse scanning of the coefficients based on the scanning order performed in the apparatus for encoding the video.
[0092] The inverse quantization module 220 may perform inverse quantization based on the quantization parameter received from the apparatus for encoding the video and the coefficients of the rearranged blocks.
[0093] The inverse transform module 225 may perform an inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the quantization result of the apparatus for encoding the video. This inverse transform is the inverse processing of the transform, i.e., DCT, DST, and KLT, performed by the transform module. The inverse transform may be performed based on the transmission unit determined by the apparatus for encoding the video. The inverse transform module 225 of the apparatus for decoding the video may selectively perform transform schemes (e.g., DCT, DST, and KLT) according to multiple pieces of information such as the prediction method, the size of the current block, and the prediction direction.
[0094] The prediction modules 230 and 235 may generate a prediction block based on the information regarding the generation of the prediction block received from the entropy decoding module 210 and the previously decoded block or picture information received from the memory 245.
[0095] As described above, similar to the operation of the apparatus for encoding video, when performing intra prediction, when the size of the prediction unit is the same as the size of the transform unit, intra prediction can be performed on the prediction unit based on the pixels located on the left, upper left, and upper sides of the prediction unit. When performing intra prediction, when the size of the prediction unit is different from the size of the transform unit, intra prediction can be performed using the reference pixels based on the transform unit. In addition, intra prediction using NxN partitioning can be used only for the smallest coding unit.
[0096] The prediction modules 230 and 235 may include a prediction unit determination module, an inter prediction module, and an intra prediction module. The prediction unit determination module may receive various information such as prediction unit information, prediction mode information of the intra prediction method, information on motion prediction of the inter prediction method, etc. from the entropy decoding module 210, may divide the current coding unit into prediction units, and may determine whether to perform inter prediction or intra prediction on the prediction units. By using the information required for inter prediction of the current prediction unit received from the apparatus for encoding video, the inter prediction module 230 may perform inter prediction on the current prediction unit based on the information of at least one of the previous picture or the subsequent picture of the current picture including the current prediction unit. Alternatively, inter prediction may be performed based on the information of some pre-reconstructed regions in the current picture including the current prediction unit.
[0097] To perform inter prediction, it may be determined for the coding unit which of the skip mode, merge mode, AMVP mode, and inter block copy mode is used as the motion prediction method for the prediction units included in the coding unit.
[0098] The intra prediction module 235 may generate a prediction block based on the pixel information in the current picture. When the prediction unit is a prediction unit subject to intra prediction, intra prediction may be performed based on the intra prediction mode information of the prediction unit received from the apparatus for encoding video. The intra prediction module 235 may include an adaptive intra smoothing (AIS) filter, a reference pixel interpolation module, and a DC filter. The AIS filter filters the reference pixels of the current block and may determine whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering may be performed on the reference pixels of the current block by using the AIS filter information and the prediction mode of the prediction unit received from the apparatus for encoding video. When the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0099] When the prediction mode of the prediction unit is a prediction mode in which intra prediction is performed based on pixel values obtained by interpolating reference pixels, the reference pixel interpolation module may interpolate the reference pixels to generate reference pixels that are either integer pixels or sub-integer pixels. When the prediction mode of the current prediction unit is a prediction mode in which a prediction block is generated without interpolating the reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is the DC mode, the DC filter may generate a prediction block by filtering.
[0100] The reconstructed block or picture may be provided to the filter module 240. The filter module 240 may include a deblocking filter, an offset correction module, and an ALF.
[0101] Information on whether to apply the deblocking filter to a corresponding block or picture and information on which of a strong filter and a weak filter to apply when applying the deblocking filter may be received from a device for encoding video. The deblocking filter of a device for decoding video may receive information on the deblocking filter from the device for encoding video and may perform deblocking filtering on the corresponding block.
[0102] The offset correction module may perform offset correction on the reconstructed picture based on the type and offset value information of the offset correction applied to the picture during encoding.
[0103] The ALF may be applied to an encoding unit based on information on whether to apply the ALF received from a device for encoding video, ALF coefficient information, etc. The ALF information may be provided as being included in a specific parameter set.
[0104] The memory 245 may store the reconstructed picture or block to be used as a reference picture or block and may provide the reconstructed picture to the output module.
[0105] As described above, in an embodiment of the present invention, for ease of explanation, the encoding unit is used as a term representing a unit for encoding, but the encoding unit may be used as a unit that performs both decoding and encoding.
[0106] In addition, the current block may represent a target block to be encoded / decoded. And, depending on the encoding / decoding step, the current block may represent a coding tree block (or coding tree unit), a coding block (or coding unit), a transform block (or transform unit), a prediction block (or prediction unit), etc. In this specification, 'unit' may represent a basic unit for performing a specific encoding / decoding process, and 'block' may represent an array of samples of a predetermined size. Unless otherwise specified, 'block' and 'unit' may be used with the same meaning. For example, in the examples mentioned later, it may be understood that a coding block and a coding unit have the same meaning for each other.
[0107] A picture can be encoded / decoded by dividing it into basic blocks having a square shape or a non-square shape. At this time, the basic block can be referred to as an encoding tree unit. The encoding tree unit can be defined as the largest-sized encoding unit allowed within a sequence or slice. Information indicating whether the encoding tree unit has a square shape or a non-square shape, or information related to the size of the encoding tree unit, can be signaled by a sequence parameter set, a picture parameter set, or a slice header. The encoding tree unit can be divided into smaller-sized partitions. At this time, if it is assumed that the depth of the partition generated by dividing the encoding tree unit is 1, the depth of the partition generated by dividing the partition with a depth of 1 can be defined as 2. That is, the partition generated by dividing the partition with a depth of k in the encoding tree unit can be defined as having a depth of k + 1.
[0108] Any-sized partition generated by dividing the encoding tree unit can be defined as an encoding unit. The encoding unit can be recursively divided or split into basic units for performing prediction, quantization, transformation, loop filtering, etc. For example, any-sized partition generated by dividing the encoding unit can be defined as an encoding unit, or can be defined as a transformation unit or a prediction unit, which is a basic unit for performing prediction, quantization, transformation, loop filtering, etc.
[0109] Alternatively, a prediction block having the same size as or smaller than the encoding block can be determined by a prediction partition of the encoding block. For the prediction partition of the encoding block, any one of the partition mode (Part_mode) candidates indicating the partition shape of the encoding block can be specified. Information for determining the partition index indicating any one of the partition mode candidates can be signaled by a bitstream. Alternatively, the partition index of the encoding block can be determined based on at least one of the size, shape, or encoding mode of the encoding block. The size or shape of the prediction block can be determined based on the partition mode specified by the partition index. The partition mode candidates can include asymmetric partition shapes (e.g., nLx2N, nRx2N, 2NxnU, 2NxnD). The number or type of asymmetric partition mode candidates available for the encoding block can be determined based on at least one of the size, shape, or encoding mode of the encoding block.
[0110] Figure 3 is a diagram showing the partition mode candidates that can be applied to an encoding block when encoding the encoding block by inter-frame prediction.
[0111] When encoding an encoding block by inter-frame prediction, Figure 3 any one of the 8 partition mode candidates shown in
[0112] On the other hand, when encoding a coding block by intra prediction, only square partitioning can be applied to the coding block. In other words, when encoding a coding block by intra prediction, a partitioning mode, PART_2Nx2N or PART_NxN, can be applied to the coding block.
[0113] When the coding block has the minimum size, PART_NxN can be applied. Herein, the minimum size of the coding block can be predefined in the encoder and the decoder. Alternatively, information related to the minimum size of the coding block can be signaled through the bitstream. In an example, the minimum size of the coding block can be signaled through the header. Thus, the minimum size of the coding block can be determined differently for each slice.
[0114] In another example, the partitioning mode candidates available for a coding block can be determined differently according to at least one of the size or shape of the coding block. In an example, the number or type of the partitioning mode candidates available for a coding block can be determined differently according to at least one of the size or shape of the coding block.
[0115] Alternatively, the type or number of the asymmetric partitioning mode candidates available for a coding block can be determined based on the size or shape of the coding block. The number or type of the asymmetric partitioning mode candidates available for a coding block can be determined differently according to at least one of the size or shape of the coding block. In an example, when the coding block has a non-square shape with a width greater than the height, at least one of PART_2NxN, PART_2NxnU, or PART_2NxnD may not be used as a partitioning mode candidate for the coding block. When the coding block has a non-square shape with a height greater than the width, at least one of PART_Nx2N, PART_nLx2N, PART_nRx2N may not be used as a partitioning mode candidate for the coding block.
[0116] Generally, a prediction block can have a size from 4x4 to 64x64. However, when encoding a coding block by inter prediction, the prediction block can be restricted to not having a size of 4x4 to reduce the memory bandwidth during motion compensation.
[0117] Based on the partitioning mode, the coding block can be recursively partitioned. In other words, based on the partitioning mode determined by the partitioning index, the coding block can be partitioned and each partition generated by partitioning the coding block can be defined as a coding block.
[0118] In the following, a method of partitioning a coding unit will be described in more detail. In the examples mentioned later, the coding unit may refer to a coding tree unit or a coding unit included in the coding tree unit. Additionally, a "partition" generated by partitioning a coding block may refer to a "coding block". The partitioning methods mentioned later can be applied when partitioning a coding block into a plurality of prediction blocks or transform blocks.
[0119] A coding unit can be partitioned by at least one line. In this case, the angle of the line for partitioning the coding unit can be a value within the range of 0 to 360 degrees. For example, the angle of a horizontal line can be 0 degrees, the angle of a vertical line can be 90 degrees, the angle of a diagonal line in the upper right direction can be 45 degrees, and the angle of a diagonal line in the upper left can be 135 degrees.
[0120] When partitioning a coding unit by multiple lines, all the lines among the multiple lines can have the same angle. Alternatively, at least one of the multiple lines can have an angle different from that of the other lines. Alternatively, the multiple lines for partitioning the coding tree unit or the coding unit can have a predefined angle difference (e.g., 90 degrees).
[0121] Information related to the line for partitioning the coding unit can be determined by a partitioning pattern. Alternatively, information regarding at least one of the number, direction, angle, or position of the lines in the block can be encoded.
[0122] For ease of description, in the examples mentioned later, it is assumed that a coding unit is partitioned into a plurality of coding units by using at least one of a vertical line or a horizontal line.
[0123] The number of vertical or horizontal lines for partitioning the coding unit can be at least one or more. In the example, a coding unit can be partitioned into 2 partitions by using one vertical line or one horizontal line. Alternatively, a coding unit can be partitioned into 3 partitions by using two vertical lines or two horizontal lines. Alternatively, a coding unit can be partitioned into 4 partitions by using one vertical line or one horizontal line, and the width and height of the 4 partitions are each halved with respect to the coding unit.
[0124] When a coding unit is partitioned into a plurality of partitions by using at least one vertical line or at least one horizontal line, the partitions can have a uniform size. Alternatively, one partition can have a size different from that of the other partitions, or each partition can have a different size. In the example, when a coding unit is partitioned by two horizontal lines or two vertical lines, the coding unit can be partitioned into 3 partitions. In this case, the width ratio or height ratio of the 3 partitions can be n:2n:n, 2n:n:n, or n:n:2n.
[0125] In the examples mentioned later, dividing a coding block into four divisions is called quadtree-based division. Also, dividing a coding block into two divisions is called binary-tree-based division. Additionally, dividing a coding block into three divisions is called ternary-tree-based division.
[0126] In the accompanying drawings mentioned later, dividing a coding unit using a vertical line and / or a horizontal line will be shown, but it will be described that dividing a coding unit into more divisions than shown or dividing a coding unit into fewer divisions than shown by using more vertical lines and / or more horizontal lines than shown is also included within the scope of the present invention.
[0127] Figure 4 An example of hierarchical division of a coding block based on a tree structure as an embodiment applying the present invention is shown.
[0128] An input video signal is decoded in a predetermined block unit, and a basic unit for decoding the input video signal is called a coding block. A coding block can be a unit that performs intra / inter prediction, transformation, and quantization. Additionally, a prediction mode (e.g., intra prediction mode or inter prediction mode) can be determined in units of coding blocks, and prediction blocks included in a coding block can share the determined prediction mode. A coding block can be a square block or a non-square block of any size in the range of 8x8 to 64x64, or can be a square block or a non-square block having a size of 128x128, 256x256, or larger.
[0129] Specifically, a coding block can be hierarchically divided based on at least one of a quadtree division method, a binary-tree division method, or a ternary-tree division method. Quadtree-based division can mean a method of dividing a 2Nx2N coding block into four NxN coding blocks. Binary-tree-based division can mean a method of dividing a coding block into two coding blocks. Ternary-tree-based division can mean a method of dividing a coding block into three coding blocks. Even when ternary-tree-based or binary-tree-based division is performed, square coding blocks can exist at a lower depth.
[0130] Divisions generated by binary-tree-based division can be symmetric or asymmetric. Additionally, coding blocks based on binary-tree division can be square blocks or non-square blocks (e.g., rectangles).
[0131] Figure 5FIG. is a diagram showing the partitioning shape of a coding block based on binary tree partitioning. The partitioning shape of a coding block based on binary tree partitioning may include symmetric types such as 2NxN (non-square coding units in the horizontal direction) or Nx2N (non-square coding units in the vertical direction), or asymmetric types such as nLx2N, nRx2N, 2NxnU, or 2NxnD. Only one of the symmetric type or the asymmetric type may be allowed as the partitioning shape of the coding block.
[0132] The ternary tree partitioning shape may include at least one of a shape that divides the coding block into two vertical lines or a shape that divides the coding block into two horizontal lines. Three non-square partitions may be generated by ternary tree partitioning.
[0133] Figure 6 FIG. shows the ternary tree partitioning shape.
[0134] The ternary tree partitioning shape may include a shape that divides the coding block into two horizontal lines or a shape that divides the coding block into two vertical lines. The width ratio or height ratio of the partitions generated by partitioning the coding block may be n:2n:n, 2n:n:n, or n:n:2n.
[0135] The position of the partition having the maximum width or height among the three partitions may be predefined in the encoder and the decoder. Alternatively, information indicating the partition having the maximum width or height among the three partitions may be signaled in the bitstream.
[0136] For a coding unit, only a square shape or a non-square symmetric shape of partitioning may be allowed. In this case, partitioning the coding unit into square partitions may correspond to quadtree CU partitioning, and partitioning the coding unit into non-square partitions of symmetric shape may correspond to binary tree partitioning. Partitioning the coding tree unit into square partitions and non-square partitions of symmetric shape may correspond to quadtree and binary tree CU partitioning (QTBT).
[0137] Binary tree or ternary tree partitioning may be performed on a coding block for which quadtree-based partitioning is no longer performed. The coding blocks generated by binary tree or ternary tree partitioning may be divided into smaller coding blocks. In this case, at least one of quadtree partitioning, ternary tree partitioning, or binary tree partitioning may be set not to be applied to the coding block. Alternatively, for a coding block, binary tree partitioning in a predetermined direction or ternary tree partitioning in a predetermined direction may not be allowed. In an example, for a coding block generated by binary tree or ternary tree partitioning, quadtree partitioning and ternary tree partitioning may be set to be not allowed. For this coding block, only binary tree partitioning may be allowed.
[0138] Alternatively, only the largest coding block among the three coding blocks generated by the quadtree-based partitioning may be partitioned into smaller coding blocks. Alternatively, only the largest coding block among the three coding blocks generated by the quadtree-based partitioning may be allowed to be partitioned based on a binary tree or based on a quadtree.
[0139] The partitioning shape of a lower-depth partitioning may be determined accordingly based on the partitioning shape of a higher-depth partitioning. In an example, when the higher and lower partitionings are partitioned based on a binary tree, for the lower-depth partitioning, only a binary-tree partitioning having the same shape as the binary-tree partitioning shape of the higher-depth partitioning may be allowed. For example, when the binary-tree partitioning shape of the higher-depth partitioning is 2NxN, the binary-tree partitioning shape of the lower-depth partitioning may also be set to 2NxN. Alternatively, when the binary-tree partitioning shape of the higher-depth partitioning is Nx2N, the partitioning shape of the lower-depth partitioning may also be set to Nx2N.
[0140] Alternatively, for the largest partitioning among the partitionings generated by the quadtree-based partitioning, a binary-tree partitioning or a quadtree partitioning in the same partitioning direction as the higher-depth partitioning may be set to be not allowed.
[0141] Alternatively, the partitioning shape of a lower-depth partitioning may be determined by considering the partitioning shape of a higher-depth partitioning and the partitioning shapes of adjacent lower-depth partitionings. Specifically, if the higher-depth partitioning is partitioned based on a binary tree, the partitioning shape of the lower-depth partitioning may be determined such that a result the same as that of partitioning the higher-depth partitioning based on a quadtree does not occur. In an example, when the partitioning shape of the higher-depth partitioning is 2NxN and the partitioning shape of an adjacent lower-depth partitioning is Nx2N, the partitioning shape of the current lower-depth partitioning may not be set to Nx2N. This is because when the partitioning shape of the current lower-depth partitioning is Nx2N, it results in a result the same as that of partitioning the higher-depth partitioning based on a quadtree having an NxN shape. When the partitioning shape of the higher-depth partitioning is Nx2N and the partitioning shape of an adjacent lower-depth partitioning is 2NxN, the partitioning shape of the current lower-depth partitioning may not be set to 2NxN. In other words, when the binary-tree partitioning shape of the higher-depth partitioning is different from the binary-tree partitioning shape of an adjacent lower-depth partitioning, the binary-tree partitioning shape of the current lower-depth partitioning may be set to be the same as the binary-tree partitioning shape of the higher-depth partitioning.
[0142] Alternatively, the binary-tree partitioning shape of a lower-depth partitioning may be set to be different from the binary-tree partitioning shape of a higher-depth partitioning.
[0143] The allowable binary tree partitioning shapes can be determined in units of sequences, slices, or coding units. In an example, the allowable binary tree partitioning shapes of coding tree units can be restricted to 2NxN or Nx2N shapes. The allowable partitioning shapes can be predefined in the encoder or decoder. Alternatively, information regarding allowable partitioning shapes or non-allowable partitioning shapes can be encoded and signaled via a bitstream.
[0144] Figure 7 is a diagram showing an example of a binary tree-based partitioning that only allows specific shapes.
[0145] Figure 7 (a) represents an example that only allows partitioning based on a binary tree in the Nx2N shape, and Figure 7 (b) represents an example that only allows partitioning based on a binary tree in the 2NxN shape.
[0146] To represent various partitioning shapes, information regarding quadtree partitioning, information regarding binary tree partitioning, or information regarding ternary tree partitioning can be used. The information regarding quadtree partitioning can include at least one of information indicating whether quadtree-based partitioning is performed or information regarding the size / depth of coding blocks allowing quadtree-based partitioning. The information regarding binary tree partitioning can include at least one of the following: information indicating whether binary tree-based partitioning is performed, information regarding whether the binary tree-based partitioning is in the vertical direction or the horizontal direction, information regarding the size / depth of coding blocks allowing binary tree-based partitioning, or information regarding the size / depth of coding blocks not allowing binary tree-based partitioning. The information regarding ternary tree partitioning can include at least one of the following: information indicating whether ternary tree-based partitioning is performed, information regarding whether the ternary tree-based partitioning is in the vertical direction or the horizontal direction, information regarding the size / depth of coding blocks allowing ternary tree-based partitioning, or information regarding the size / depth of coding blocks not allowing ternary tree-based partitioning. The information regarding the size of a coding block can represent at least one minimum value or maximum value among the width, height, product of the width and height, or ratio of the width to the height of the coding block.
[0147] In an example, when the width or height of a coding block is less than the minimum size allowing binary tree partitioning, or when the partitioning depth of the coding block is greater than the maximum depth allowing binary tree partitioning, binary tree-based partitioning can be not allowed for that coding block.
[0148] In an example, when the width or height of a coding block is less than the minimum size allowing ternary tree partitioning, or when the partitioning depth of the coding block is greater than the maximum depth allowing ternary tree partitioning, ternary tree-based partitioning can be not allowed for that coding block.
[0149] Information about conditions allowing binary tree- or ternary tree-based partitioning can be signaled via a bitstream. The information can be encoded in units of a sequence, a picture, or a partial picture. The partial picture can mean at least one of a slice, a tile group, a tile, a brick, an encoded block, a prediction block, or a transform block.
[0150] In an example, the syntax'max_mtt_depth_idx_minus1' indicating the maximum depth allowing binary tree / ternary tree partitioning can be encoded / decoded via a bitstream. In this case, max_mtt_depth_idx_minus1 + 1 can indicate the maximum depth allowing binary tree / ternary tree partitioning.
[0151] In an example, at least one of the number of times allowing binary tree / ternary tree partitioning, the maximum depth allowing binary tree / ternary tree partitioning, or the number of depths allowing binary tree / ternary tree partitioning can be signaled at the sequence or slice level. Thus, for a first slice and a second slice, at least one of the number of times allowing binary tree / ternary tree partitioning, the maximum depth allowing binary tree / ternary tree partitioning, or the number of depths allowing binary tree / ternary tree partitioning can be different. In an example, while for the first slice, binary tree / ternary tree partitioning can be allowed only at one depth, for the second slice, binary tree / ternary tree partitioning can be allowed at two depths.
[0152] In Figure 8 the example shown, Figure 8 binary tree partitioning is shown to be performed on an encoding unit of depth 2 and an encoding unit of depth 3. Thus, at least one of the following can be encoded / decoded via a bitstream: information representing the number of times (2 times) of performing binary tree partitioning in an encoding tree unit, information representing the maximum depth (depth 3) of the partitions generated by binary tree partitioning in an encoding tree unit, or information representing the number of partition depths (2 depths, depth 2 and depth 3) to which binary tree partitioning is applied in an encoding tree unit.
[0153] Alternatively, the number of times allowing binary tree / ternary tree partitioning, the depth allowing binary tree / ternary tree partitioning, or the number of depths allowing binary tree / ternary tree partitioning can be predefined in an encoder and a decoder. Alternatively, the number of times allowing binary tree / ternary tree partitioning, the depth allowing binary tree / ternary tree partitioning, or the number of depths allowing binary tree / ternary tree partitioning can be determined based on at least one of an index of a sequence or a slice or the size / shape of an encoding unit. In an example, for a first slice, binary tree / ternary tree partitioning can be allowed at one depth, while for a second slice, binary tree / ternary tree partitioning can be allowed at two depths.
[0154] In another example, at least one of the number of times of allowing binary tree partitioning, the depth of allowing binary tree partitioning, or the number of depths of allowing binary tree partitioning may be set differently according to the temporal level identifier (TemporalID) of a slice or picture. Herein, the temporal level identifier (TemporalID) is used to identify each of multiple layers in an image having scalability in at least one or more of perspectives, space, time, or quality.
[0155] As Figure 4 shown, a first coding block 300 with a partitioning depth (split depth) of k may be partitioned into multiple second coding blocks based on a quadtree. For example, second coding blocks 310 to 340 may be square blocks with a width and height halved with respect to the first coding block, and the partitioning depth of the second coding blocks may be increased to k + 1.
[0156] The second coding block 310 with a partitioning depth of k + 1 may be partitioned into multiple third coding blocks with a partitioning depth of k + 2. The partitioning of the second coding block 310 may be performed by selectively using one of a quadtree or a binary tree according to a partitioning method. In this case, the partitioning method may be determined based on at least one of information indicating quadtree-based partitioning or information indicating binary tree-based partitioning.
[0157] When the second coding block 310 is partitioned based on a quadtree, the second coding block 310 may be partitioned into four third coding blocks 310a with a width and height halved with respect to the second coding block, and the partitioning depth of the third coding blocks 310a may be increased to k + 2. In other words, when the second coding block 310 is partitioned based on a binary tree, the second coding block 310 may be partitioned into two third coding blocks. In this case, each of the two third coding blocks may be a non-square block with one of the width and height halved with respect to the second coding block, and the partitioning depth may be increased to k + 2. The second coding block may be determined as a non-square block in a horizontal direction or a vertical direction according to a partitioning direction, and the partitioning direction may be determined based on information on whether binary tree-based partitioning is performed in the vertical direction or the horizontal direction.
[0158] Meanwhile, the second coding block 310 may be determined as a leaf coding block that is no longer partitioned based on a quadtree or a binary tree, and in this case, the corresponding coding block may be used as a prediction block or a transform block.
[0159] Similar to the partitioning of the second coding block 310, the third coding block 310a may be determined as a leaf coding block, or may be further partitioned based on a quadtree or a binary tree.
[0160] On the other hand, the third coding block 310b divided based on the binary tree can be further divided into a coding block 310b-2 in the vertical direction or a coding block 310b-3 in the horizontal direction based on the binary tree, and the division depth of the corresponding coding block can be increased to k+3. Alternatively, the third coding block 310b can be determined as a leaf coding block 310b-1 that is no longer divided based on the binary tree, and in this case, the corresponding coding block 310b-1 can be used as a prediction block or a transform block. However, the above-mentioned division process can be restrictively performed based on at least one of the following: information about the size / depth of the coding block that allows division based on the quadtree, information about the size / depth of the coding block that allows division based on the binary tree, or information about the size / depth of the coding block that does not allow division based on the binary tree.
[0161] The number of candidates representing the size of the coding block can be limited to a predetermined number, or the size of the coding block in a predetermined unit can have a fixed value. In an example, the size of the coding block in a sequence or a picture can be limited to any one of 256x256, 128x128, or 32x32. Information representing the size of the coding block in a sequence or a picture can be signaled in a sequence header or a picture header.
[0162] As a result of division based on the quadtree and the binary tree, the coding unit can be represented as a square or rectangular shape of any size.
[0163] As Figure 4 shown, the first coding block 300 with a division depth (split depth) of k can be divided into a plurality of second coding blocks based on the quadtree. For example, the second coding blocks 310 to 340 can be square blocks with a width and height halved with respect to the first coding block, and the division depth of the second coding block can be increased to k+1.
[0164] The second coding block 310 with a division depth of k+1 can be divided into a plurality of third coding blocks with a division depth of k+2. The division of the second coding block 310 can be performed by selectively using one of the quadtree or the binary tree according to the division method. In this case, the division method can be determined based on at least one of the information indicating division based on the quadtree or the information indicating division based on the binary tree.
[0165] When the second coding block 310 is partitioned based on a quadtree, the second coding block 310 can be partitioned into four third coding blocks 310a with widths and heights halved with respect to the second coding block, and the partition depth of the third coding blocks 310a can be increased to k + 2. In other words, when the second coding block 310 is partitioned based on a binary tree, the second coding block 310 can be partitioned into two third coding blocks. In this case, each of the two third coding blocks can be a non-square block with one of the width and height halved with respect to the second coding block, and the partition depth can be increased to k + 2. The second coding block can be determined as a non-square block in the horizontal or vertical direction according to the partition direction, and the partition direction can be determined based on the information on whether the binary-tree-based partition is performed in the vertical or horizontal direction.
[0166] Meanwhile, the second coding block 310 can be determined as a leaf coding block that is no longer partitioned based on a quadtree or a binary tree, and in this case, the corresponding coding block can be used as a prediction block or a transform block.
[0167] Similar to the partition of the second coding block 310, the third coding block 310a can be determined as a leaf coding block, or can be further partitioned based on a quadtree or a binary tree.
[0168] On the other hand, the third coding block 310b partitioned based on a binary tree can be further partitioned into coding blocks 310b-2 in the vertical direction or coding blocks 310b-3 in the horizontal direction based on the binary tree, and the partition depth of the corresponding coding blocks can be increased to k + 3. Alternatively, the third coding block 310b can be determined as a leaf coding block 310b-1 that is no longer partitioned based on the binary tree, and in this case, the corresponding coding block 310b-1 can be used as a prediction block or a transform block. However, the above-mentioned partitioning process can be restrictively performed based on at least one of the following: information on the size / depth of coding blocks allowing quadtree-based partitioning, information on the size / depth of coding blocks allowing binary-tree-based partitioning, or information on the size / depth of coding blocks not allowing binary-tree-based partitioning.
[0169] The number of candidates representing the size of coding blocks can be limited to a predetermined number, or the size of coding blocks in a predetermined unit can have a fixed value. In an example, the size of coding blocks in a sequence or a picture can be limited to any one of 256x256, 128x128, or 32x32. Information representing the size of coding blocks in a sequence or a picture can be signaled by a sequence header or a picture header.
[0170] As a result of partitioning based on a quadtree and a binary tree, coding units can be represented as square or rectangular shapes of arbitrary sizes.
[0171] The transform skip can be set not to be used for coding units generated by binary tree - based partitioning or ternary tree - based partitioning. Alternatively, the transform skip can be set to be applied to at least one of the vertical direction or the horizontal direction in non - square coding units. In an example, when the transform skip is applied to the horizontal direction, it means that scaling is performed only in the horizontal direction without performing transform / inverse transform, and transform / inverse transform using DCT or DST is performed in the vertical direction. When the transform skip is applied to the vertical direction, it means that scaling is performed only in the vertical direction without performing transform / inverse transform, and transform / inverse transform using DCT or DST is performed in the horizontal direction.
[0172] Information on whether to skip the inverse transform in the horizontal direction or information on whether to skip the inverse transform in the vertical direction can be signaled through the bitstream. In an example, the information on whether to skip the inverse transform in the horizontal direction can be a 1 - bit flag 'hor_transform_skip_flag', and the information on whether to skip the inverse transform in the vertical direction can be a 1 - bit flag'ver_transform_skip_flag'.
[0173] The encoder can determine whether to code 'hor_transform_skip_flag' or'ver_transform_skip_flag' based on the size and / or shape of the current block. In an example, when the current block has an Nx2N shape, hor_transform_skip_flag can be coded and the coding of ver_transform_skip_flag can be omitted. When the current block has a 2NxN shape, ver_transform_skip_flag can be coded, and hor_transform_skip_flag can be omitted.
[0174] Alternatively, based on the size and / or shape of the current block, it can be determined whether to perform transform skip in the horizontal direction or whether to perform transform skip in the vertical direction. In an example, when the current block has an Nx2N shape, the transform skip can be applied to the horizontal direction, and transform / inverse transform can be performed in the vertical direction. When the current block has a 2NxN shape, the transform skip can be applied to the vertical direction, and transform / inverse transform can be performed in the horizontal direction. Transform / inverse transform can be performed based on at least one of DCT or DST.
[0175] As a result of partitioning based on a quadtree, binary tree, or ternary tree, coding blocks that are no longer partitioned can be used as prediction blocks or transform blocks. In other words, coding blocks generated by quadtree partitioning or binary tree partitioning can be used as prediction blocks or transform blocks. In an example, a prediction image can be generated in units of coding blocks, and the difference between the residual signal, the original image, and the prediction image can be transformed in units of coding blocks. To generate a prediction image in units of coding blocks, motion information can be determined based on the coding blocks, or an intra prediction mode can be determined based on the coding blocks. Thus, a coding block can be coded by using at least one of a skip mode, intra prediction, or inter prediction.
[0176] Alternatively, a plurality of coding blocks generated by partitioning a coding block can be set to share at least one of motion information, merge candidates, reference samples, reference sample lines, or an intra prediction mode. In an example, when a coding block is partitioned by a ternary tree, the partitions generated by partitioning the coding block can share at least one of motion information, merge candidates, reference samples, reference sample lines, or an intra prediction mode according to the size or shape of the coding block. Alternatively, only a part of the plurality of coding blocks can be set to share information, and the remaining coding blocks can be set not to share information.
[0177] In another example, prediction blocks or transform blocks smaller than the coding block can be used by partitioning the coding block.
[0178] Hereinafter, a method of performing inter prediction on a coding block or a prediction block generated by partitioning a coding block will be described in detail.
[0179] Figure 9 is a flowchart showing an inter prediction method according to an embodiment to which the present invention is applied.
[0180] Referring to Figure 9 , motion information S910 of the current block can be determined. The motion information of the current block can include at least one of a motion vector of the current block, a reference picture index of the current block, an inter prediction direction of the current block, or a weight of weighted prediction. The weight of weighted prediction can represent the weight applied to the L0 reference block and the weight applied to the L1 reference block.
[0181] The motion vector of the current block can be determined based on information signaled through a bitstream. The precision of the motion vector represents the basic unit for expressing the motion vector of the current block. For example, the precision of the motion vector of the current block can be determined to be one of integer pixels, 1 / 2 pixels, 1 / 4 pixels, or 1 / 8 pixels. The precision of the motion vector can be determined on a per picture basis, on a per slice basis, on a per tile group basis, on a per tile basis, or on a per block basis. A block can represent a coding tree unit, a coding unit, a prediction unit, or a transform unit.
[0182] The motion information of a current block can be obtained based on at least one of information signaled through a bitstream or motion information of neighboring blocks adjacent to the current block.
[0183] Figure 10 FIG. is a diagram illustrating a process of deriving motion information of a current block when a merge mode is applied to the current block.
[0184] The merge mode represents a method of deriving motion information of a current block based on neighboring blocks.
[0185] When the merge mode is applied to a current block, a spatial merge candidate S1010 can be derived based on spatial neighboring blocks of the current block. The spatial neighboring blocks can include at least one of blocks neighboring the upper side boundary, left side boundary, or corners (e.g., at least one of the upper left corner, upper right corner, or lower left corner) of the current block.
[0186] Figure 11 FIG. is a diagram illustrating an example of spatial neighboring blocks.
[0187] As Figure 11 shown in the example of, the spatial neighboring blocks can include at least one of the following: neighboring block A neighboring the left side of the current block 1 , neighboring block B1 neighboring the upper side of the current block, neighboring block A neighboring the lower left corner of the current block 0 , neighboring block B neighboring the upper right corner of the current block 0 and neighboring block B neighboring the upper left corner of the current block. 2 . For example, assume that the position of the upper left corner sample of the current block is (0, 0), the width of the current block is W, and the height of the current block is H. Block A 1 can include samples at position (-1, H - 1). Block B 1 can include samples at position (W - 1, -1). Block A 0 can include samples at position (-1, H). Block B 0 can include samples at position (W, -1). Block B 2 can include samples at position (-1, -1).
[0188] Further expand Figure 11In an example, the spatial merge candidate can be derived based on a block adjacent to the upper-left sample of the current block or a block adjacent to the upper-center sample of the current block. For example, a block adjacent to the upper-left sample of the current block can include at least one of a block including a sample at position (0, -1) or a block including a sample at position (-1, 0). Alternatively, the spatial merge candidate can be derived based on at least one of a block adjacent to the upper-center sample of the current block or a block adjacent to the left-center sample of the current block. For example, a block adjacent to the upper-center sample of the current block can include a sample at position (W / 2, -1). A block adjacent to the left-center sample of the current block can include a sample at position (-1, H / 2).
[0189] Based on the size and / or shape of the current block, the positions of the upper adjacent block and / or the left adjacent block for deriving the spatial merge candidate can be determined. In an example, when the size of the current block is greater than a threshold, the spatial merge candidate can be derived based on a block adjacent to the upper-center sample of the current block and a block adjacent to the left-center sample of the current block. On the other hand, when the size of the current block is less than the threshold, the spatial merge candidate can be derived based on a block adjacent to the upper-right sample of the current block and a block adjacent to the lower-left sample of the current block. Herein, the size of the current block can be expressed based on at least one of width, height, the sum of width and height, the product of width and height, or the ratio of width to height. The threshold can be an integer, such as 2, 4, 8, 16, 32, or 128.
[0190] Based on the shape of the current block, the availability of the extended spatial adjacent block can be determined. In an example, when the current block is a non-square block with a width greater than the height, it can be determined that a block adjacent to the upper-left sample of the current block, a block adjacent to the left-center sample of the current block, or a block adjacent to the lower-left sample of the current block is unavailable. Meanwhile, when the current block is a block with a height greater than the width, it can be determined that a block adjacent to the upper-left sample of the current block, a block adjacent to the upper-center sample of the current block, or a block adjacent to the upper-right sample of the current block is unavailable.
[0191] The motion information of the spatial merge candidate can be set to be the same as the motion information of the spatial adjacent block.
[0192] The spatial merge candidate can be determined by searching adjacent blocks in a predetermined order. In an example, in the example shown in Figure 11 , the search can be performed in the order of block A 1 , B 1 , B 0 , A 0 and B 2 for determining the spatial merge candidate. Herein, when there are no remaining blocks (i.e., A 1 , B 1 , B 0 and A 0when encoding at least one of them or performing intra prediction mode on at least one of them, block B can be used 2 .
[0193] The order for searching for spatial merge candidates can be predefined in the encoder / decoder. Alternatively, the order for searching for spatial merge candidates can be adaptively determined according to the size or shape of the current block. Alternatively, the order for searching for spatial merge candidates can be determined based on the information signaled via the bitstream.
[0194] Temporal merge candidate S1020 can be derived based on the temporally adjacent blocks of the current block. Temporally adjacent blocks can mean co-located blocks included in the co-located picture. The co-located picture has a different POC from the current picture including the current block. The co-located picture can be determined as the picture with a predefined index in the reference picture list or the picture with the smallest POC difference from the POC of the current picture. Alternatively, the co-located picture can be determined by the information signaled via the bitstream. The information signaled via the bitstream can include at least one of the following: information indicating the reference picture list (e.g., L0 reference picture list or L1 reference picture list) including the co-located picture and the index of the co-located picture within the reference picture list. The information for determining the co-located picture can be signaled in at least one of the picture parameter set, the slice header, and the block level.
[0195] The motion information regarding the temporal merge candidate can be determined based on the motion information of the co-located block. In an example, the motion vector of the temporal merge candidate can be determined based on the motion vector of the co-located block. For example, the motion vector of the temporal merge candidate can be set to be the same as the motion vector of the co-located block. Alternatively, the motion vector of the temporal merge candidate can be derived by scaling the motion vector of the co-located block based on at least one of the POC difference between the current picture of the current block and the reference picture and the POC difference between the co-located picture of the co-located block and the reference picture.
[0196] Figure 12 is a diagram showing an example of deriving the motion vector of the temporal merge candidate.
[0197] In Figure 12 the example shown, tb represents the POC difference between the current picture curr_pic and the reference picture curr_ref of the current picture, and td represents the POC difference between the co-located picture col_pic of the co-located block and the reference picture col_ref. The motion vector of the temporal merge candidate can be derived by scaling the motion vector of the co-located block col_PU based on tb and / or td.
[0198] Alternatively, considering whether a collocated block is available, the motion vector of the collocated block and the motion vector obtained by scaling the motion vector of the collocated block can be used as the motion vector of the temporal merge candidate. In an example, the motion vector of the collocated block is set as the motion vector of the first temporal merge candidate, and the value obtained by scaling the motion vector of the collocated block can be set as the motion vector of the second temporal merge candidate.
[0199] The inter prediction direction of the temporal merge candidate can be set to be the same as the inter prediction direction of the temporally adjacent block. However, the reference picture index of the temporal merge candidate can have a fixed value. In an example, the reference picture index of the temporal merge candidate can be set to "0". Alternatively, the reference picture index of the temporal merge candidate can be adaptively determined based on at least one of the reference picture index of the spatial merge candidate and the reference picture index of the current picture.
[0200] A specific block within the collocated picture that has the same position and size as the current block, or a block adjacent to a block that is adjacent to a block having the same position and size as the current block can be determined as the collocated block.
[0201] Figure 13 is a diagram showing the positions of candidate blocks that can be used as collocated blocks.
[0202] The candidate block can include at least one of the following: a block within the collocated picture that is adjacent to the position of the upper left corner of the current block, a block within the collocated picture that is adjacent to the position of the center sample of the current block, and a block within the collocated picture that is adjacent to the position of the lower left corner of the current block.
[0203] In an example, the candidate block can include at least one of the following: block TL within the collocated picture that includes the position of the upper left sample of the current block, block BR within the collocated picture that includes the position of the lower right sample of the current block, block H within the collocated picture that is adjacent to the lower right corner of the current block, block C3 within the collocated picture that includes the position of the center sample of the current block, and block C0 within the collocated picture that is adjacent to the center sample of the current block (e.g., a block that includes the position of a sample that is (-1, -1) away from the center sample of the current block).
[0204] In addition to Figure 13 the examples shown in, a block within the collocated picture that includes the position of an adjacent block adjacent to a predetermined boundary of the current block can be selected as the collocated block.
[0205] The number of temporal merge candidates can be 1 or more. In an example, at least one temporal merge candidate can be derived based on at least one collocated block.
[0206] Information about the maximum number of temporal merge candidates can be encoded by an encoder and signaled. Alternatively, the maximum number of temporal merge candidates can be derived based on the maximum number of merge candidates that may be included in a merge candidate list and / or the maximum number of spatial merge candidates. Alternatively, the maximum number of temporal merge candidates can be determined based on the number of available co-located blocks.
[0207] Whether a candidate block is available can be determined according to a predetermined priority, and at least one co-located block can be determined based on the above determination and the maximum number of temporal merge candidates. In an example, when block C3 including the position of the central sample of the current block and block H adjacent to the lower right corner of the current block are candidate blocks, either block C3 or block H can be determined as the co-located block. When block H is available, block H can be determined as the co-located block. However, when block H is not available (e.g., when block H is encoded by intra prediction, when block H is not available, or when block H is outside the largest coding unit (LCU), etc.), block C3 can be determined as the co-located block.
[0208] In another example, when at least one of a plurality of blocks adjacent to the lower right corner position of the current block within a co-located picture is not available (e.g., block H and / or block BR), the unavailable block can be replaced with another available block. Another available block for replacing the unavailable block can include at least one block adjacent to the central sample position of the current block within the co-located picture (e.g., C0 and / or C3) and a block adjacent to the lower left corner of the current block within the co-located picture (e.g., TL).
[0209] When at least one of a plurality of blocks adjacent to the central sample position of the current block within a co-located picture is not available, or when at least one of a plurality of blocks adjacent to the upper left corner position of the current block within a co-located picture is not available, the unavailable block can be replaced with another available block.
[0210] Subsequently, a merge candidate list including spatial merge candidates and temporal merge candidates can be generated (S1030). When configuring the merge candidate list, merge candidates having the same motion information as existing merge candidates can be removed from the merge candidate list.
[0211] Information about the maximum number of merge candidates can be signaled via a bitstream. In an example, information indicating the maximum number of merge candidates can be signaled via sequence parameters or picture parameters. In an example, when the maximum number of merge candidates is six, a total of six can be selected from spatial merge candidates and temporal merge candidates. For example, five spatial merge candidates can be selected from five merge candidates, and one temporal merge candidate can be selected from two temporal merge candidates.
[0212] Alternatively, the maximum number of merge candidates can be predefined in the encoder and decoder. For example, the maximum number of merge candidates can be 2, 3, 4, 5, or 6. Alternatively, the maximum number of merge candidates can be determined based on at least one of whether merge with MVD (MMVD) is performed, whether combined prediction is performed, or whether triangular partitioning is performed.
[0213] If the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, the merge candidates included in the second merge candidate list can be added to the merge candidate list.
[0214] The second merge candidate list can include merge candidates derived based on the motion information of the blocks encoded / decoded by inter prediction before the current block. In an example, if motion compensation of a block with an inter prediction encoding mode is performed, the merge candidates derived based on the motion information of the block can be added to the second merge candidate list. If the encoding / decoding of the current block is completed, the motion information of the current block can be added to the second merge candidate list for inter prediction of subsequent blocks.
[0215] The second merge candidate list can be initialized in units of CTU, tile, or slice. The maximum number of merge candidates that can be included in the second merge candidate list can be predefined in the encoder and decoder. Alternatively, the information indicating the maximum number of merge candidates that can be included in the second merge candidate list can be signaled through the bitstream.
[0216] The index of the merge candidate included in the second merge candidate list can be determined based on the order in which it is added to the second merge candidate list. In an example, the index assigned to the Nth merge candidate added to the second merge candidate list can have a smaller value than the index assigned to the (N + 1)th merge candidate added to the second merge candidate list. For example, the index of the (N + 1)th merge candidate can be set to a value increased by 1 relative to the index of the Nth merge candidate. Alternatively, the index of the Nth merge candidate can be set to the index of the (N + 1)th merge candidate, and the value of the index of the Nth merge candidate can be subtracted by 1.
[0217] Alternatively, the index assigned to the Nth merge candidate added to the second merge candidate list can have a larger value than the index assigned to the (N + 1)th merge candidate added to the second merge candidate list. For example, the index of the Nth merge candidate can be set to the index of the (N + 1)th merge candidate, and the value of the index of the Nth merge candidate can be increased by 1.
[0218] Whether the motion information of the block for which motion compensation is performed is the same as the motion information of the merge candidate included in the second merge candidate list can determine whether the merge candidate derived from the block is added to the second merge candidate list. In an example, when a merge candidate having the same motion information as the block is included in the second merge candidate list, the merge candidate derived based on the motion information of the block may not be added to the second merge candidate list. Alternatively, when a merge candidate having the same motion information as the block is included in the second merge candidate list, the merge candidate may be deleted from the second merge candidate list, and the merge candidate derived based on the motion information of the block may be added to the second merge candidate list.
[0219] When the number of merge candidates included in the second merge candidate list is the same as the maximum number of merge candidates, the merge candidate with the lowest index or the merge candidate with the highest index may be deleted from the second merge candidate list, and the merge candidate derived based on the motion information of the block may be added to the second merge candidate list. In other words, after deleting the oldest merge candidate among the merge candidates included in the second merge candidate list, the merge candidate derived based on the motion information of the block may be added to the second merge candidate list.
[0220] When the number of merge candidates included in the merge candidate list has not reached the maximum number of merge candidates, a combined merge candidate obtained by combining two or more merge candidates or a merge candidate having a (0, 0) motion vector (zero motion vector) may be included in the merge candidate list.
[0221] Alternatively, an average merge candidate obtained by averaging the motion vectors of two or more merge candidates found can be added to the merge candidate list. The average merge candidate can be derived by finding the average of the motion vectors of two or more merge candidates included in the merge candidate list. In an example, when a first merge candidate and a second merge candidate are added to the merge candidate list, the average of the motion vector of the first merge candidate and the motion vector of the second merge candidate can be calculated to obtain the average merge candidate. Specifically, the L0 motion vector of the average merge candidate can be derived by calculating the average of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate, and the L1 motion vector of the average merge candidate can be derived by calculating the average of the L1 motion vector of the first merge candidate and the L1 motion vector of the second merge candidate. When bi-directional prediction is applied to any one of the first merge candidate and the second merge candidate and uni-directional prediction is performed on the other merge candidate, the motion vector of the bi-directional merge candidate can be set as the L0 motion vector or the L1 motion vector of the average merge candidate as it is. In an example, when LO direction prediction and L1 direction prediction are performed on the first merge candidate but LO direction prediction is performed on the second merge candidate, the L0 motion vector of the average merge candidate can be derived by calculating the average of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate. At the same time, the L1 motion vector of the average merge candidate can be derived as the L1 motion vector of the first merge candidate.
[0222] When the reference picture of the first merge candidate is different from that of the second merge candidate, the motion vector of the first merge candidate or the second merge candidate can be scaled according to the distance (i.e., POC difference) between the reference picture of each merge candidate and the current picture. For example, after scaling the motion vector of the second merge candidate, the average merge candidate can be derived by calculating the average of the motion vector of the first merge candidate and the scaled motion vector of the second merge candidate. Herein, the priority can be set based on the value of the reference picture index of each merge candidate, the distance between the reference picture of each merge candidate and the current block, or whether bi-directional prediction is applied, and the motion vector of the merge candidate with high (or low) priority can be scaled.
[0223] The reference picture index of the average merge candidate can be set to indicate a reference picture at a specific position within the reference picture list. In an example, the reference picture index of the average merge candidate can indicate the first reference picture or the last reference picture within the reference picture list. Alternatively, the reference picture index of the average merge candidate can be set to be the same as the reference picture index of the first merge candidate or the second merge candidate. In an example, when the reference picture index of the first merge candidate is the same as that of the second merge candidate, the reference picture index of the average merge candidate can be set to be the same as the reference picture indices of the first merge candidate and the second merge candidate. When the reference picture index of the first merge candidate is different from that of the second merge candidate, the priority can be set based on the value of the reference picture index of each merge candidate, the distance between the reference picture of each merge candidate and the current block, or whether bidirectional prediction is applied, and the reference picture index of the merge candidate with high (or low) priority can be set as the reference picture index of the average merge candidate. In an example, when bidirectional prediction is applied to the first merge candidate and unidirectional prediction is applied to the second merge candidate, the reference picture index of the first merge candidate to which bidirectional prediction is applied can be determined as the reference picture index of the average merge candidate.
[0224] Based on the priority among combinations of merge candidates, the sequence for generating the combination of the average merge candidate can be determined. The priority can be predefined in the encoder and the decoder. Alternatively, the sequence of the combination can be determined based on whether bidirectional prediction of the merge candidate is performed. For example, the combination of merge candidates encoded using bidirectional prediction can be set to have a higher priority than the combination of merge candidates encoded using unidirectional prediction. Alternatively, the sequence of the combination can be determined based on the reference pictures of the merge candidates. For example, the combination of merge candidates with the same reference picture can have a higher priority than the combination of merge candidates with different reference pictures.
[0225] The merge candidates can be included in the merge candidate list according to the predefined priority. The merge candidate with high priority can be assigned a small index value. In an example, the spatial merge candidate can be added to the merge candidate list before the temporal merge candidate. Additionally, the spatial merge candidates can be added to the merge candidate list in the order of the spatial merge candidate of the left adjacent block, the spatial merge candidate of the upper adjacent block, the spatial merge candidate of the block adjacent to the upper right corner, the spatial merge candidate of the block adjacent to the lower left corner, and the spatial merge candidate of the block adjacent to the upper left corner. Alternatively, it can be set such that the spatial merge candidate derived from the adjacent block ([ Figure 11 B2) adjacent to the upper left corner of the current block is added to the merge candidate list later than the temporal merge candidate.
[0226] In another example, the priority among merge candidates can be determined according to the size or shape of the current block. In the example, when the current block has a rectangular shape with a width greater than the height, the spatial merge candidate of the left adjacent block can be added to the merge candidate list before the spatial merge candidate of the upper adjacent block. In other words, when the current block has a rectangular shape with a height greater than the width, the spatial merge candidate of the upper adjacent block can be added to the merge candidate list before the spatial merge candidate of the left adjacent block.
[0227] In another example, the priority among merge candidates can be determined according to the motion information of each merge candidate. In the example, a merge candidate with bidirectional motion information can have a higher priority than a merge candidate with unidirectional motion information. Therefore, a merge candidate with bidirectional motion information can be added to the merge candidate list before a merge candidate with unidirectional motion information.
[0228] In another example, a merge candidate list can be generated according to a predefined priority, and then the merge candidates can be rearranged. The rearrangement can be performed based on the motion information of the merge candidates. In the example, the rearrangement can be performed based on whether the merge candidate has bidirectional motion information, the magnitude of the motion vector, the accuracy of the motion vector, or the POC difference between the reference picture of the merge candidate and the current picture. Specifically, a merge candidate with bidirectional motion information can be rearranged to have a higher priority than a merge candidate with unidirectional motion information. Alternatively, a merge candidate with a motion vector having a fractional pixel accuracy value can be rearranged to have a higher priority than a merge candidate with a motion vector having an integer pixel accuracy.
[0229] When generating the merge candidate list, at least one of the merge candidates included in the merge candidate list can be specified based on the merge candidate index (S1040).
[0230] The motion information of the current block can be set to be the same as the motion information of the merge candidate specified by the merge candidate index (S1050). In the example, when selecting a spatial merge candidate through the merge candidate index, the motion information of the current block can be set to be the same as the motion information of the spatially adjacent block. Alternatively, when selecting a temporal merge candidate through the merge candidate index, the motion information of the current block can be set to be the same as the motion information of the temporally adjacent block.
[0231] Figure 14 FIG. is a diagram illustrating a process of deriving motion information of a current block when the AMVP mode is applied to the current block.
[0232] When the AMVP mode is applied to a current block, at least one of an inter prediction direction and a reference picture index of the current block can be decoded from a bitstream S1410. In other words, when the AMVP mode is applied, at least one of the inter prediction direction and the reference picture index of the current block can be determined based on information encoded through the bitstream.
[0233] A spatial motion vector candidate can be determined based on motion vectors of spatially adjacent blocks of the current block S1420. The spatial motion vector candidate can include at least one of the following: a first spatial motion vector candidate derived from an upper adjacent block of the current block and a second spatial motion vector candidate derived from a left adjacent block of the current block. In this document, the upper adjacent block can include at least one of blocks adjacent to the upper side and the upper right corner of the current block, and the left adjacent block of the current block includes at least one of blocks adjacent to the left side and the lower left corner of the current block. A block adjacent to the upper left corner of the current block can be used as an upper adjacent block or can be used as a left adjacent block.
[0234] Alternatively, a spatial motion vector candidate can be derived based on spatially non - adjacent blocks not adjacent to the current block. In an example, a spatial motion vector candidate of the current block can be derived by using at least one of the following: a block on the same vertical line as a block adjacent to the upper side, the upper right corner, or the upper left corner of the current block; a block on the same horizontal line as a block adjacent to the left side, the lower left corner, or the upper left corner of the current block; and a block on the same diagonal line as a block adjacent to a corner of the current block. When spatially adjacent blocks are not available, a spatial motion vector candidate can be derived by using spatially non - adjacent blocks.
[0235] In another example, at least two spatial motion vector candidates can be derived by using spatially adjacent blocks and spatially non - adjacent blocks. In the example, a first spatial motion vector candidate and a second spatial motion vector candidate can be derived by using adjacent blocks adjacent to the current block. Meanwhile, a third spatial motion vector candidate and / or a fourth spatial motion vector candidate can be derived based on blocks not adjacent to the current block but adjacent to the above - mentioned adjacent blocks.
[0236] When the current block is different from a spatially adjacent block on a reference picture, a spatial motion vector can be obtained by scaling the motion vector of the spatially adjacent block. A temporal motion vector candidate can be determined based on motion vectors of temporally adjacent blocks of the current block S1430. When the current block is different from a temporally adjacent block on a reference picture, a temporal motion vector can be obtained by scaling the motion vector of the temporally adjacent block. In this document, when the number of spatial motion vector candidates is equal to or less than a predetermined number, a temporal motion vector candidate can be derived.
[0237] A motion vector candidate list including spatial motion vector candidates and temporal motion vector candidates can be generated S1440.
[0238] When generating a motion vector candidate list, at least one motion vector candidate included in the motion vector candidate list can be specified based on information specifying at least one of the motion vector candidates S1450.
[0239] The motion vector candidate specified by the information can be set as a predicted value of the motion vector of the current block, and the motion vector of the current block can be obtained by adding the residual value of the motion vector and the predicted value of the motion vector S1460. Herein, the residual value of the motion vector can be parsed through a bitstream.
[0240] When obtaining the motion information of the current block, motion compensation of the current block can be performed based on the obtained motion information S920. Specifically, motion compensation of the current block can be performed based on the inter prediction direction, the reference picture index, and the motion vector of the current block. The inter prediction direction indicates whether to perform L0 prediction, L1 prediction, or bi-prediction. When encoding the current block by bi-prediction, the predicted block of the current block can be obtained based on a weighted sum operation or an average operation of the L0 reference block and the L1 reference block.
[0241] When obtaining predicted samples by performing motion compensation, the current block can be reconstructed based on the generated predicted samples. Specifically, the reconstructed samples can be obtained by adding the predicted samples of the current block and the residual samples.
[0242] As in the above example, based on the motion information of the blocks encoded / decoded using inter prediction before the current block, the merge candidates of the current block can be derived. For example, based on the motion information of adjacent blocks at predefined positions adjacent to the current block, the merge candidates of the current block can be derived. Examples of adjacent blocks can include at least one of the following: the block adjacent to the left side of the current block, the block adjacent to the upper side of the current block, the block adjacent to the upper left corner of the current block, the block adjacent to the upper right corner of the current block, and the block adjacent to the lower left corner of the current block.
[0243] The merge candidates of the current block can be derived based on the motion information of blocks other than the adjacent blocks. For ease of description, the adjacent blocks at predefined positions adjacent to the current block are referred to as first merge candidate blocks, and the blocks at positions different from the first merge candidate blocks are referred to as second merge candidate blocks.
[0244] The second merge candidate blocks can include at least one of the following: the blocks encoded / decoded using inter prediction before the current block, the blocks adjacent to the first merge candidate blocks, or the blocks on the same line as the first merge candidate blocks. Figure 15 The second merge candidate blocks adjacent to the first merge candidate blocks are shown, and Figure 16 The second merge candidate blocks on the same line as the first merge candidate blocks are shown.
[0245] When the first merge candidate block is not available, a merge candidate derived based on the motion information of the second merge candidate block is added to the merge candidate list. Alternatively, even if at least one of the spatial merge candidate and the temporal merge candidate is added to the merge candidate list, when the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, a merge candidate derived based on the motion information of the second merge candidate block is added to the merge candidate list.
[0246] Figure 15 It is a diagram showing an example of deriving a merge candidate from a second merge candidate block when the first merge candidate block is not available.
[0247] When the first merge candidate block AN (where N ranges from 0 to 4 in this document) is not available, a merge candidate for the current block is derived based on the motion information of the second merge candidate block BM (where M ranges from 0 to 6 in this document). That is, a merge candidate for the current block can be derived by replacing the unavailable first merge candidate block with the second merge candidate block.
[0248] Among the blocks adjacent to the first merge candidate block, the block placed from the first merge candidate block in a predefined direction can be set as the second merge candidate block. The predefined direction can be a left direction, a right direction, an up direction, a down direction, or a diagonal direction. A predefined direction can be set for each first merge candidate block. For example, the predefined direction of the first merge candidate block adjacent to the left side of the current block can be the left direction. The predefined direction of the first merge candidate block adjacent to the upper side of the current block can be the up direction. The predefined direction of the first merge candidate block adjacent to the corner of the current block can include at least one of the left direction, the up direction, or the diagonal direction.
[0249] For example, when A0 adjacent to the left side of the current block is not available, a merge candidate for the current block is derived based on B0 adjacent to A1. When A1 adjacent to the upper side of the current block is not available, a merge candidate for the current block is derived based on B1 adjacent to A1. When A2 adjacent to the upper right corner of the current block is not available, a merge candidate for the current block is derived based on B2 adjacent to A2. When A3 adjacent to the lower left corner of the current block is not available, a merge candidate for the current block is derived based on B3 adjacent to A3. When A4 adjacent to the upper left corner of the current block is not available, a merge candidate for the current block is derived based on at least one of B4 to B6 adjacent to A4.
[0250] Figure 15 The examples shown are only for describing the embodiments of the present invention and do not limit the present invention. The position of the second merge candidate block can be set to be Figure 15The samples shown in are different. For example, a second merge candidate block adjacent to the first merge candidate block may be located in the upward or downward direction of the first merge candidate block, and the first merge candidate block is adjacent to the left side of the current block. Alternatively, a second merge candidate block adjacent to the first merge candidate block may be located in the leftward or rightward direction of the first merge candidate block, and the first merge candidate block is adjacent to the upper side of the current block.
[0251] Figure 16 FIG. is a diagram showing an example of deriving a merge candidate based on a second merge candidate block located on the same line as the first merge candidate block.
[0252] Blocks located on the same line as the first merge candidate block may include at least one of the following: a block located on the same horizontal line as the first merge candidate block, a block located on the same vertical line as the first merge candidate block, or a block located on the same diagonal line as the first merge candidate block. Blocks located on the same horizontal line have the same y - coordinate position. Blocks located on the same vertical line have the same x - coordinate position. The difference between the x - coordinate positions of blocks located on the same diagonal line is the same as the difference between the y - coordinate positions.
[0253] Assume that the upper - left sample of the current block is located at (0, 0), and the width and height of the current block are W and H, respectively. In Figure 18 FIG. , the positions of second merge candidate blocks (e.g., B4, C6) located on the same vertical line as the first merge candidate block are determined based on the right - most block on the upper side of the coded block (e.g., block A1 including coordinates (W - 1, - 1)). Additionally, in Figure 18 FIG. , the positions of second merge candidate blocks (e.g., B1, C1) located on the same horizontal line as the first merge candidate block are determined based on the lowest block on the left side of the coded block (e.g., block A0 including coordinates (- 1, H - 1)).
[0254] In another example, the position of the second merge candidate block may be determined based on the left - most block on the upper side of the coded block (e.g., a block including coordinates (0, - 1)) or a block located at the center of the upper side of the coded block (e.g., a block including coordinates (W / 2, - 1)). Additionally, the position of the second merge candidate block may be determined based on the uppermost block on the left side of the coded block (e.g., a block including coordinates (- 1, 0)) or a block located at the center of the left side of the coded block (e.g., a block including coordinates (- 1, H / 2)).
[0255] In another example, when there are multiple upper adjacent blocks adjacent to the upper side of the current block, all or part of the multiple upper adjacent blocks can be used to determine the second merge candidate block. In the example, the second merge candidate block can be determined by using a block at a specific position among the multiple upper adjacent blocks (for example, at least one of the upper adjacent blocks located at the leftmost side, the upper adjacent block located at the rightmost side, or the upper adjacent block located at the center). The number of upper adjacent blocks among the multiple upper adjacent blocks used to determine the second merge candidate block can be 1, 2, 3, or more. Additionally, when there are multiple left adjacent blocks adjacent to the left side of the current block, all or part of the multiple left adjacent blocks can be used to determine the second merge candidate block. In the example, the second merge candidate block can be determined by using a block at a specific position among the multiple left adjacent blocks (for example, at least one of the left adjacent blocks located at the lowermost side, the left adjacent block located at the uppermost side, or the left adjacent block located at the center). The number of left adjacent blocks among the multiple left adjacent blocks used to determine the second merge candidate block can be 1, 2, 3, or more.
[0256] According to the size and / or shape of the current block, the positions and / or numbers of the upper adjacent block and / or the left adjacent block used to determine the second merge candidate block can be determined differently. In the example, when the size of the current block is greater than a threshold, the second merge candidate block can be determined based on the upper center block and / or the left center block. On the other hand, when the size of the current block is less than the threshold, the second merge candidate block can be determined based on the uppermost right block and / or the lowermost left block. The threshold can be an integer, such as 8, 16, 32, 64, or 128.
[0257] A first merge candidate list and a second merge candidate list can be constructed, and motion compensation for the current block can be performed based on at least one of the first merge candidate list or the second merge candidate list.
[0258] The first merge candidate list can include at least one of the following: a spatial merge candidate derived based on the motion information of adjacent blocks at predefined positions adjacent to the current block, or a temporal merge candidate derived based on the motion information of co-located blocks.
[0259] The second merge candidate list can include merge candidates derived based on the motion information of the second merge candidate blocks.
[0260] As an embodiment of the present invention, the first merge candidate list can be constructed to include merge candidates derived from the first merge candidate blocks, and the second merge candidate list can be constructed to include merge candidates derived from the second merge candidate blocks. In the example, Figure 15In the example shown, merge candidates derived from blocks A0 to A4 can be added to the first merge candidate list, and merge candidates derived from blocks B0 to B6 can be added to the second merge candidate list. In the example, in Figure 16 In the example shown, merge candidates derived from blocks A0 to A4 can be added to the first merge candidate list, and merge candidates derived from blocks B0 to B5 and C0 to C7 can be added to the second merge candidate list.
[0261] Alternatively, the second merge candidate list may include merge candidates derived based on the motion information of blocks that use inter prediction encoding / decoding before the current block. For example, when performing motion compensation for a block with an inter prediction encoding mode, the merge candidates derived based on the motion information of the block are added to the second merge candidate list. When the encoding / decoding of the current block is completed, the motion information of the current block is added to the second merge candidate list for inter prediction of subsequent blocks.
[0262] The index of the merge candidates included in the second merge candidate list can be determined based on the order in which the merge candidates are added to the second merge candidate list. For example, the index assigned to the Nth merge candidate added to the second merge candidate list may have a lower value than the index assigned to the (N + 1)th merge candidate added to the second merge candidate list. For example, the index of the (N + 1)th merge candidate can be set to have a value 1 higher than the index of the Nth merge candidate. Alternatively, the index of the Nth merge candidate can be set to the index of the (N + 1)th merge candidate, and the value of the index of the Nth merge candidate is decreased by 1.
[0263] Alternatively, the index assigned to the Nth merge candidate added to the second merge candidate list may have a higher value than the index assigned to the (N + 1)th merge candidate added to the second merge candidate list. For example, the index of the Nth merge candidate can be set to the index of the (N + 1)th merge candidate, and the value of the index of the Nth merge candidate is increased by 1.
[0264] Based on whether the motion information of the block undergoing motion compensation is the same as the motion information of the merge candidates included in the second merge candidate list, it can be determined whether to add the merge candidates derived from the block to the second merge candidate list. For example, when a merge candidate with the same motion information as the block is included in the second merge candidate list, the merge candidates derived based on the motion information of the block are not added to the second merge candidate list. Alternatively, when a merge candidate with the same motion information as the block is included in the second merge candidate list, the merge candidate is deleted from the second merge candidate list, and the merge candidates derived based on the motion information of the block are added to the second merge candidate list.
[0265] When the number of merge candidates included in the second merge candidate list is the same as the maximum number of merge candidates, a merge candidate with the lowest index or a merge candidate with the highest index is detected from the second merge candidate list, and a merge candidate derived based on the motion information of the block is added to the second merge candidate list. That is, after deleting the oldest merge candidate among the merge candidates included in the second merge candidate list, a merge candidate derived based on the motion information of the block can be added to the second merge candidate list.
[0266] The second merge candidate list can be initialized in units of CTUs, tiles, or slices. In other words, blocks different from the current block included in a CTU, tile, or slice can be set to be unavailable as second merge candidate blocks. The maximum number of merge candidates that can be included in the second merge candidate list can be predefined in the encoder and decoder. Alternatively, information representing the maximum number of merge candidates that can be included in the second merge candidate list can be signaled via a bitstream.
[0267] Either the first merge candidate list or the second merge candidate list can be selected, and the selected merge candidate list can be used to perform inter prediction for the current block. Specifically, based on the index information, any one of the merge candidates included in the merge candidate list can be selected, and the motion information of the current block can be obtained from the merge candidate.
[0268] Information specifying the first merge candidate list or the second merge candidate list can be signaled via a bitstream. The decoder can select the first merge candidate list or the second merge candidate list based on this information.
[0269] Alternatively, among the first merge candidate list and the second merge candidate list, the merge candidate list including a larger number of available merge candidates can be selected.
[0270] Alternatively, the first merge candidate list or the second merge candidate list can be selected based on at least one of the size, shape, and partition depth of the current block.
[0271] Alternatively, the merge candidate list is configured by adding (or appending) one of the first merge candidate list and the second merge candidate list to the other.
[0272] For example, inter prediction can be performed based on a merge candidate list having at least one merge candidate included in the first merge candidate list and at least one merge candidate included in the second merge candidate list.
[0273] For example, a merge candidate included in the second merge candidate list can be added to the first merge candidate list. Alternatively, a merge candidate included in the first merge candidate list can be added to the second merge candidate.
[0274] When the number of merge candidates included in the first merge candidate list is less than the maximum number, or when the first merge candidate block is unavailable, the merge candidates included in the second merge candidate list are added to the first merge candidate list.
[0275] Alternatively, when the first merge candidate block is unavailable, a merge candidate derived from a block adjacent to the first merge candidate block among the merge candidates included in the second merge candidate list is added to the first merge candidate list. Refer to Figure 15 , when A0 is unavailable, a merge candidate derived from the motion information of B0 among the merge candidates included in the second merge candidate list is added to the first merge candidate list. When A1 is unavailable, a merge candidate derived from the motion information of B1 among the merge candidates included in the second merge candidate list is added to the first merge candidate list. When A2 is unavailable, a merge candidate derived from the motion information of B2 among the merge candidates included in the second merge candidate list is added to the first merge candidate list. When A3 is unavailable, a merge candidate derived from the motion information of B3 among the merge candidates included in the second merge candidate list is added to the first merge candidate list. When A4 is unavailable, a merge candidate derived from the motion information of B4, B5, or B6 among the merge candidates included in the second merge candidate list is added to the first merge candidate list.
[0276] Alternatively, the merge candidate to be added to the first merge candidate list can be determined according to the priority of the merge candidates included in the second merge candidate list. The priority can be determined based on the index value assigned to each merge candidate. For example, when the number of merge candidates included in the first merge candidate list is less than the maximum number, or when the first merge candidate block is unavailable, the merge candidate with the smallest index value or the largest index value among the merge candidates included in the second merge candidate list is added to the first merge candidate list.
[0277] When there is a merge candidate in the first merge candidate list that has the same motion information as the merge candidate with the highest priority among the merge candidates included in the second merge candidate list, the merge candidate with the highest priority may not be added to the first merge candidate list. Additionally, it can be determined whether a merge candidate with the next priority (for example, a merge candidate assigned an index value that is 1 greater than the index value assigned to the merge candidate with the highest priority or a merge candidate assigned an index value that is 1 less than the index value assigned to the merge candidate with the highest priority) can be added to the first merge candidate list.
[0278] Alternatively, a merge candidate list including a merge candidate derived from motion information of a first merge candidate block and a merge candidate derived from motion information of a second merge candidate block may be generated. The merge candidate list may be a combination of a first merge candidate list and a second merge candidate list.
[0279] For example, according to a predetermined search order, a merge candidate list may be generated by searching a first merge candidate block and a second merge candidate block.
[0280] Figures 17 to 20 is a diagram showing an order of searching merge candidate blocks.
[0281] Figures 17 to 20 The order of searching merge candidates is shown as follows.
[0282] A0→A1→A2→A3→A4→B0→B1→B2→B3→B4→(B5)→(B6).
[0283] Searching of blocks B5 and B6 is performed only when block B4 is unavailable or when the number of merge candidates included in the merge candidate list is equal to or less than a preset number.
[0284] A search order different from the example shown in Figures 17 to 20 may be set.
[0285] A merge candidate list may be generated, which includes a combination of at least one merge candidate included in a first merge candidate list and at least one merge candidate included in a second merge candidate list. For example, the combined merge candidate list may include N merge candidates among the merge candidates included in the first merge candidate list and M merge candidates among the merge candidates included in the second merge candidate list. The letters N and M may represent the same number or different numbers. Alternatively, at least one of N and M may be determined based on at least one of the number of merge candidates included in the first merge candidate list and the number of merge candidates included in the second merge candidate list. Alternatively, information for determining at least one of N and M may be signaled through a bitstream. Any one of N and M may be derived by subtracting the other of N and M from the maximum number of merge candidates in the combined merge candidate list.
[0286] A merge candidate to be added to the combined merge candidate list may be determined according to a predefined priority. The predefined priority may be determined based on an index assigned to the merge candidate.
[0287] Alternatively, a merge candidate to be added to the list of merge candidates for the combination can be determined based on the association between merge candidates. For example, when A0 included in the first merge candidate list is added to the list of merge candidates for the combination, a merge candidate (e.g., B0) at a position adjacent to A0 is not added to the combination list.
[0288] When the number of merge candidates included in the first merge candidate list is less than N, more than M merge candidates among the merge candidates included in the second merge candidate list are added to the list of merge candidates for the combination. For example, when N is 4 and M is 2, 4 merge candidates among the merge candidates included in the first merge candidate list are added to the list of merge candidates for the combination, and 2 merge candidates among the merge candidates included in the second merge candidate list are added to the list of merge candidates for the combination. When the number of merge candidates included in the first merge candidate list is less than 4, two or more merge candidates among the merge candidates included in the second merge candidate list are added to the list of merge candidates for the combination. When the number of merge candidates included in the second merge candidate list is less than 2, four or more merge candidates among the merge candidates included in the first merge candidate list are added to the list of merge candidates for the combination.
[0289] That is, the values of N or M can be adjusted according to the number of merge candidates included in each merge candidate list. By adjusting the values of N or M, the total number of merge candidates included in the list of merge candidates for the combination can be fixed. When the total number of merge candidates included in the list of merge candidates for the combination is less than the maximum number of merge candidates, combination merge candidates, average merge candidates, or zero motion vector candidates are added.
[0290] Motion compensation for the current block can be performed by using at least one of the merge candidates included in the first merge candidate list and the second merge candidate list. The encoder can encode index information for specifying any one of the multiple merge candidates. In an example, `merge_idx` can specify any one of the multiple merge candidates. In an example, Table 1 shows the merge indices for each merge candidate derived from the first merge candidate block and the second merge candidate block shown in Figure 16 the first merge candidate block and the second merge candidate block shown in
[0291]
Table 1
[0292]
[0293]
[0294] However, as the number of merge candidates included in the merge candidate list increases, the codewords for encoding the merge index become longer. Therefore, there is a problem of reducing the encoding / decoding efficiency. To reduce the length of the codewords, the merge index can be determined by using a prefix and a suffix. In the example, the merge index can be determined by using merge_idx_prefix representing the prefix of the merge index and merge_idx_suffix representing the suffix of the merge index.
[0295] Table 2 shows the merge index prefix values and merge index suffix values for each merge index, and Table 3 shows the process of determining the merge index based on the merge index prefix values and merge index suffix values.
[0296]
Table 2
[0297]
[0298]
[0299]
Table 3
[0300]
[0301] As shown in Tables 2 and 3, when the merge index prefix value is less than the threshold, the merge index can be set to be the same as the merge index prefix value. On the other hand, when the merge index prefix value is greater than the threshold, the merge index can be determined by subtracting a reference value from the merge index prefix and adding the merge index suffix to the shifted value of the result. The reference value can be the threshold, or can be a value obtained by subtracting 1 from the threshold.
[0302] Tables 2 and 3 show that the threshold is 4. The threshold can be determined based on at least one of the number of merge candidates included in the merge candidate list, the number of second merge candidate blocks, or the number of lines including the second merge candidate blocks. Alternatively, the threshold can be predefined in the encoder and decoder.
[0303] It can be determined whether to use a prefix and a suffix to determine the merge index according to the number of merge candidates included in the merge candidate list or the maximum number of merge candidates that can be included in the merge candidate list. In the example, when the maximum number of merge candidates that can be included in the merge candidate list is greater than the threshold, the merge index prefix and merge index suffix for determining the merge index can be signaled. On the other hand, when the maximum number of merge candidates is less than the threshold, the merge index can be signaled.
[0304] A rectangular block can be divided into a plurality of triangular blocks. The merge candidates for the triangular blocks can be derived based on the rectangular block including the triangular blocks. The triangular blocks can share the same merge candidates.
[0305] A merge index can be signaled for each triangular block. In this case, the triangular blocks can be set such that the same merge candidates are not used. In the example, the merge candidates for the first triangular block can be not used as the merge candidates for the second triangular block. Thus, the merge index for the second triangular block can specify any one of the remaining merge candidates other than the merge candidates selected for the first triangular block.
[0306] Merge candidates can be derived based on a block having a predetermined shape or a predetermined size or larger. When the current block does not have the predetermined shape, or when the size of the current block is smaller than the predetermined size, the merge candidates for the current block are derived based on a block that includes the current block and has the predetermined shape or has the predetermined size or larger. The predetermined shape can be a square shape or a non-square shape.
[0307] When the predetermined shape is a square shape, the merge candidates for a non-square-shaped coding unit are derived based on a square-shaped coding unit that includes the non-square-shaped coding unit.
[0308] Figure 21 FIG. is an example showing the derivation of the merge candidates for a non-square block based on a square block.
[0309] The merge candidates for a non-square block can be derived based on a square block that includes the non-square block. For example, the merge candidates for a non-square-shaped coding block 0 and a non-square-shaped coding block 1 can be derived based on a square-shaped block that includes coding block 0 and coding block 1. That is, the positions of spatially adjacent blocks can be determined based on the position, width / height, or size of the square-shaped block. The merge candidates for coding block 0 and coding block 1 can be derived based on at least one of the spatially adjacent blocks A0, A1, A2, A3, and A4 adjacent to the square-shaped block.
[0310] The temporal merge candidates can be determined based on a square-shaped block. That is, the temporally adjacent blocks can be determined based on the position, width / height, or size of the square-shaped block. For example, the merge candidates for coding block 0 and coding block 1 can be derived based on the temporally adjacent blocks determined based on the square-shaped block.
[0311] Alternatively, any one of the spatial merge candidates and the temporal merge candidates can be derived based on a square block, and the other merge candidate can be derived based on a non-square block. For example, the spatial merge candidate for coding block 0 can be derived based on a square block, while the temporal merge candidate for coding block 0 can be derived based on coding block 0.
[0312] Multiple blocks included in a block having a predetermined shape or a predetermined size or larger can share merge candidates. For example, in Figure 21In the example shown, at least one of the spatial merge candidates and the temporal merge candidates for coding block 0 and coding block 1 may be the same.
[0313] The predetermined shape may be a non-square shape, such as 2NxN, Nx2N, etc. When the predetermined shape is a non-square shape, the merge candidates for the current block may be derived based on the non-square block including the current block. For example, when the current block is in a 2Nxn shape (here, n is 1 / 2N), the merge candidates for the current block are derived based on the non-square block of 2NxN shape. Alternatively, when the current block is in an nx2N shape, the merge candidates for the current block are derived based on the non-square block of Nx2N shape.
[0314] Information indicating the predetermined shape or the predetermined size may be signaled via a bitstream. For example, information indicating either a non-square shape or a square shape may be signaled via a bitstream.
[0315] Alternatively, the predetermined shape or the predetermined size may be determined according to rules predefined in the encoder and the decoder.
[0316] When a child node does not satisfy a predetermined condition, the merge candidates for the child node are derived based on the parent node that satisfies the predetermined condition. Here, the predetermined condition may include at least one of the following: whether the block is a block generated by quadtree partitioning, whether the size of the block is exceeded, the shape of the block and the picture boundary, and whether the difference in depth between the child node and the parent node is equal to or greater than a predetermined value.
[0317] For example, the predetermined condition may include whether the block is a block generated by quadtree partitioning and whether the block is a coded block of a predetermined size or larger square shape. When the current block is generated by binary tree partitioning or ternary tree partitioning, the merge candidates for the current block are derived based on the high-level node block including the current block and satisfying the predetermined condition. When there is no high-level node block satisfying the predetermined condition, the merge candidates for the current block are derived based on the current block, the block including the current block and being of a predetermined size or larger, or the high-level node block including the current block and having a depth difference of one from the current block.
[0318] Figure 22 is a diagram showing an example of deriving merge candidates based on a high-level node block.
[0319] Blocks 0 and 1 are generated by partitioning a square block based on a binary tree. The merge candidates for blocks 0 and 1 may be derived based on the neighboring blocks determined according to the high-level node block including blocks 0 and 1 (i.e., at least one of A0, A1, A2, A3, and A4). Therefore, blocks 0 and 1 may use the same spatial merge candidates.
[0320] An advanced node block including block 2, block 3, and block 4 can be generated by dividing a square block based on a binary tree. Additionally, block 2 and block 3 can be generated by dividing a non-square shaped block based on a binary tree. A merge candidate for the non-square shaped blocks 2, 3, and 4 can be derived based on the advanced node block including the non-square shaped blocks 2, 3, and 4. That is, a merge candidate can be derived based on an adjacent block (e.g., at least one of B0, B1, B2, B3, and B4) determined according to the position, width / height, or size of the square block including blocks 2, 3, and 4. Therefore, blocks 2, 3, and 4 can use the same spatial merge candidate.
[0321] A temporal merge candidate for the non-square shaped blocks can be derived based on the advanced node block. For example, a temporal merge candidate for blocks 0 and 1 can be derived based on a square block including blocks 0 and 1. A temporal merge candidate for blocks 2, 3, and 4 can be derived based on a square block including blocks 2, 3, and 4. Additionally, the same temporal merge candidate derived from the temporally adjacent blocks determined based on each quadtree block can be used.
[0322] The low-level node blocks included in the advanced node block can share at least one of the spatial merge candidate and the temporal merge candidate. For example, the low-level node blocks included in the advanced node block can use the same merge candidate list.
[0323] Alternatively, at least one of the spatial merge candidate and the temporal merge candidate can be derived based on the low-level node block, and the other can be derived based on the advanced node block. For example, a spatial merge candidate for blocks 0 and 1 can be derived based on the advanced node block. However, a temporal merge candidate for block 0 can be derived based on block 0, and a temporal merge candidate for block 1 can be derived based on block 1.
[0324] Alternatively, when the number of samples included in the low-level node block is less than a predefined number, a merge candidate is derived based on the advanced node block including a predefined number or more of samples. For example, when at least one of the following conditions is satisfied: at least one of the low-level node blocks generated based on at least one of quadtree division, binary tree division, and ternary tree division is smaller than a preset size; at least one of the low-level node blocks is a non-square block; the advanced node block does not exceed the picture boundary; and the width or height of the advanced node block is equal to or greater than a predefined value, a merge candidate is derived based on a square or non-square shaped advanced node block including a predefined number or more of samples (e.g., 64, 128, or 256 samples). The low-level node blocks included in the advanced node block can share the merge candidate derived based on the advanced node block.
[0325] Merge candidates can be derived based on any one of the low-level node blocks, and the other low-level node blocks can be set to use the merge candidate. The low-level node blocks can be included in blocks of a predetermined shape or a predetermined size or larger. For example, the low-level node blocks can share a list of merge candidates derived based on any one of the low-level node blocks. Information on the low-level node block that is the basis for the derivation of the merge candidate can be signaled via a bitstream. This information can be index information indicating any one of the low-level node blocks. Alternatively, the low-level node block that is the basis for the derivation of the merge candidate can be determined based on at least one of the position, size, shape, and scan order of the low-level node blocks.
[0326] Information indicating whether the low-level node blocks share a list of merge candidates derived based on the high-level node blocks can be signaled via a bitstream. Based on this information, it can be determined whether the merge candidates for blocks that are not in a predetermined shape or have a size smaller than a predetermined size are derived based on the high-level node blocks that include the blocks. Alternatively, according to rules predefined in the encoder and decoder, it can be determined whether to derive the merge candidates based on the high-level node blocks.
[0327] When there are neighboring blocks adjacent to the current block within a predefined region, it is determined that the neighboring block is not available as a spatial merge candidate. The predefined region can be a parallel processing region defined for parallel processing between blocks. The parallel processing region can be referred to as a merge estimation region (MER). For example, when a neighboring block adjacent to the current block is included in the same merge estimation region as the current block, it is determined that the neighboring block is not available. A shift operation can be performed to determine whether the current block and the neighboring block are included in the same merge estimation region. Specifically, based on whether the value obtained by shifting the position of the upper-left reference sample of the current block is the same as the value obtained by shifting the position of the upper-left reference sample of the neighboring block, it can be determined whether the current block and the neighboring block are included in the same merge estimation region.
[0328] Figure 23 is a diagram showing an example of determining the availability of spatial neighboring blocks based on the merge estimation region.
[0329] In Figure 23 it is shown that the merge estimation region is in the shape of Nx2N.
[0330] The merge candidate for block 1 can be derived based on the spatial neighboring blocks adjacent to block 1. The spatial neighboring blocks can include B0, B1, B2, B3, and B4. Here, it can be determined that the spatial neighboring blocks B0 and B3 included in the same merge estimation region as block 1 are not available as merge candidates. Therefore, the merge candidate for block 1 can be derived from at least one of the spatial neighboring blocks B1, B2, and B4 other than the spatial neighboring blocks B0 and B3.
[0331] The merge candidates for block 3 can be derived based on spatially adjacent blocks adjacent to block 3. The spatially adjacent blocks can include C0, C1, C2, C3, and C4. In this document, it can be determined that the spatially adjacent block C0 included in the same merge estimation region as block 3 is not available as a merge candidate. Therefore, the merge candidates for block 3 can be derived from at least one of the spatially adjacent blocks C1, C2, C3, and C4 other than the spatially adjacent block C0.
[0332] Based on at least one of the position, size, width, and height of the merge estimation region, the merge candidates for the blocks included in the merge estimation region can be derived. For example, the merge candidates for the multiple blocks included in the merge estimation region can be derived from at least one of the spatially adjacent blocks and temporally adjacent blocks determined based on at least one of the position, size, width, and height of the merge estimation region. The blocks included in the merge estimation region can share the same merge candidates.
[0333] Figure 24 is a diagram showing an example of deriving merge candidates based on a merge estimation region.
[0334] When multiple coding units are included in the merge estimation region, the merge candidates for the multiple coding units can be derived based on the merge estimation region. That is, by using the merge estimation region as a coding unit, the merge candidates can be derived based on the position, size, or width / height of the merge estimation region.
[0335] For example, the merge candidates for both coding unit 0 (CU0) and coding unit 1 (CU1) that are both of size (n / 2)xN (where n is N / 2 in this document) and included in a merge estimation region of size (N / 2)xN can be derived based on the merge estimation region. That is, the merge candidates for coding unit 0 and coding unit 1 can be derived according to at least one of the adjacent blocks C0, C1, C2, C3, and C4 adjacent to the merge estimation region.
[0336] For example, the merge candidates for coding unit 2 (CU2), coding unit 3 (CU3), coding unit 4 (CU4), and coding unit 5 (CU5) of size nxn included in a merge estimation region of size NxN can be derived based on the merge estimation region. That is, the merge candidates for coding unit 2, coding unit 3, coding unit 4, and coding unit 5 can be derived according to at least one of the adjacent blocks C0, C1, C2, C3, and C4 adjacent to the merge estimation region.
[0337] The shape of the merged estimation region can be a square shape or a non-square shape. For example, a coding unit (or prediction unit) of square shape or a coding unit (or prediction unit) of non-square shape can be determined as the merged estimation region. The ratio between the width and height of the merged estimation region can be restricted to not exceed a predetermined range. For example, the merged estimation region cannot have a non-square shape with a ratio between the width and height exceeding 2, or a non-square shape with a ratio between the width and height less than 1 / 2. That is, the non-square merged estimation region can be in the shape of 2NxN or Nx2N. Information regarding the restriction on the ratio between the width and height can be signaled via the bitstream. Alternatively, the restriction on the ratio between the width and height can be predefined in the encoder and decoder.
[0338] At least one of the information indicating the shape of the merged estimation region and the information indicating the size of the merged estimation region can be signaled via the bitstream. For example, at least one of the information indicating the shape of the merged estimation region and the information indicating the size of the merged estimation region can be signaled via the sequence header, slice group header, picture parameter, or sequence parameter.
[0339] The shape or size of the merged estimation region can be updated on a per-sequence basis, per-picture basis, per-slice basis, per-slice group basis, per-tile basis, or per-block (CTU) basis. When the shape or size of the merged estimation region is different from the shape or size of the previous unit, information indicating the new shape or new size of the merged estimation region is signaled via the bitstream.
[0340] At least one block can be included in the merged estimation region. The blocks included in the merged estimation region can be in a square shape or a non-square shape. The maximum number or minimum number of blocks that the merged estimation region can include can be determined. For example, three, four, or more CUs can be included in the merged estimation region. This determination can be based on the information signaled via the bitstream. Alternatively, the maximum number or minimum number of blocks that the merged estimation region can include can be predefined in the encoder and decoder.
[0341] Parallel processing of blocks may be allowed in at least one of a case where the number of blocks included in the merge estimation region is less than the maximum number and a case where the number is greater than the minimum number. For example, when the number of blocks included in the merge estimation region is equal to or less than the maximum number, or when the number of blocks included in the merge estimation region is equal to or greater than the minimum number, merge candidates for the blocks are derived based on the merge estimation region. When the number of blocks included in the merge estimation region is greater than the maximum number, or when the number of blocks included in the merge estimation region is less than the minimum number, merge candidates for each block in the blocks are derived based on the size, position, width, or height of each block in the blocks.
[0342] Information indicating the shape of the merge estimation region may include a 1-bit flag. For example, the syntax “isrectagular_mer_flag” may indicate a merge candidate region having a square shape or a non-square shape. A value of isrectagular_mer_flag being 1 may indicate a non-square-shaped merge estimation region, while a value of isrectagular_mer_flag being 0 may indicate a square-shaped merge estimation region.
[0343] When the information indicates a non-square-shaped merge estimation region, information indicating at least one of the width, height, and ratio between the width and height of the merge estimation region is signaled through the bitstream. Based on this, the size and / or shape of the merge estimation region can be determined.
[0344] Applying embodiments described with emphasis on decoding processing or encoding processing to encoding processing or decoding processing is included in the scope of the present invention. Changing the order of embodiments described in a predetermined order is also included in the scope of the present invention.
[0345] Although the above embodiments have been described based on a series of steps or flowcharts, the above embodiments are not intended to limit the time series order of the present invention, and can be executed simultaneously or in a different order. Additionally, each of the components (e.g., units, modules, etc.) that make up the block diagrams in the above embodiments can be implemented as a hardware device or software, and multiple components can be combined into one hardware device or software. The above embodiments can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable storage medium can include program instructions, data files, data structures, etc. individually or in combination. Examples of computer-readable storage media include: magnetic recording media such as hard disks, floppy disks, and magnetic tapes; optical data storage media such as CD-ROMs or DVD-ROMs; magneto-optical media such as optical discs; and hardware devices such as read-only memories (ROMs), random access memories (RAMs), and flash memories, which are specifically structured to store and implement program instructions. The hardware device can be configured to be operated by one or more software modules or the software module can be configured to be operated by one or more hardware devices to perform the processing according to the present invention.
[0346] Industrial Applicability
[0347] The present invention can be applied to an electronic device capable of encoding / decoding an image.
[0348] Additionally, the present technology can also be configured as follows.
[0349] (1) A method for decoding a video, the method comprising:
[0350] Deriving merge candidates based on neighboring blocks adjacent to a current block;
[0351] Generating a first merge candidate list including the merge candidates,
[0352] wherein, when the number of merge candidates included in the first merge candidate list is less than a predetermined value, adding the merge candidates included in a second merge candidate list to the first merge candidate list;
[0353] Decoding information for specifying one of the merge candidates included in the first merge candidate list; and
[0354] Deriving motion information of the current block based on the merge candidate assigned an index determined by the information.
[0355] (2) The method according to (1), wherein the information includes an index prefix and an index suffix.
[0356] (3) The method according to (2), wherein when the value of the index prefix is less than a threshold, the index is set to be the same as the index prefix.
[0357] (4) The method according to (3), wherein when the value of the index prefix is greater than the threshold, the index is derived by adding the value of the index suffix to a value derived based on the index prefix.
[0358] (5) The method according to (3), wherein the threshold is determined based on the number of merge candidates included in the first merge candidate list.
[0359] (6) The method according to (1), wherein the second merge candidate list includes merge candidates derived from blocks not adjacent to the current block.
[0360] (7) The method according to (6), wherein the non - adjacent blocks are on the same line as the blocks adjacent to the current block.
[0361] (8) A method for encoding a video, the method comprising:
[0362] Deriving merge candidates from adjacent blocks adjacent to a current block;
[0363] Generating a first merge candidate list including the merge candidates,
[0364] wherein when the number of merge candidates included in the first merge candidate list is less than a predetermined value, adding the merge candidates included in the second merge candidate list to the first merge candidate list;
[0365] Encoding information for specifying one of the merge candidates included in the first merge candidate list; and
[0366] Deriving motion information of the current block based on the merge candidate assigned the index determined by the information.
[0367] (9) The method according to (8), wherein the information includes an index prefix and an index suffix.
[0368] (10) The method according to (9), wherein when the value of the index prefix is less than a threshold, the index is set to be the same as the index prefix.
[0369] (11) The method according to (10), wherein when the value of the index prefix is greater than the threshold, the index is derived by adding the value of the index suffix to a value derived based on the index prefix.
[0370] (12) The method according to (10), wherein the threshold is determined based on the number of merge candidates included in the first merge candidate list.
[0371] (13) The method according to (8), wherein the second merge candidate list includes merge candidates derived from blocks not adjacent to the current block.
[0372] (14) The method according to (13), wherein the non - adjacent blocks are on the same line as the blocks adjacent to the current block.
[0373] (15) An apparatus for decoding an image, the apparatus comprising:
[0374] a decoding unit that decodes information specifying one of the merge candidates included in the first merge candidate list; and
[0375] an inter - prediction unit that derives merge candidates from adjacent blocks adjacent to the current block, generates a first merge candidate list including the merge candidates, and derives motion information of the current block from the merge candidate assigned an index determined by the information,
[0376] wherein when the number of merge candidates included in the first merge candidate list is less than a predetermined value, the merge candidates included in the second merge candidate list are added to the first merge candidate list.
Claims
1. A method for decoding a video, the method comprises: Deriving merge candidates for a current coded block, the merge candidates including spatial merge candidates derived according to neighboring blocks adjacent to the current coded block; Generating a first candidate list including the merge candidates for the current coded block, wherein, when the number of the merge candidates included in the first candidate list is less than a predetermined value, adding motion information candidates included in a second candidate list to the first candidate list; Based on two index information, selecting two merge candidates from the multiple merge candidates included in the first candidate list for two triangular shape partitions in the current coded block; Obtaining a predicted block of the current coded block based on two motion information derived according to the two merge candidates; Obtaining a residual block of the current coded block through inverse quantization and inverse transformation; and Reconstructing the current coded block by summing the predicted block and the residual block, wherein, the motion information candidates included in the second candidate list are derived according to spatial blocks, the spatial blocks are decoded before decoding the current coded block and are included in the current picture including the current coded block, wherein, the first index information in the two index information specifies a first merge candidate among the multiple merge candidates, wherein, the second index information in the two index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate, and wherein, both the first index information and the second index information are only composed of a prefix part.
2. The method according to claim 1, wherein, Updating the second candidate list with motion information of a block coded by inter prediction, and wherein, adding the motion information of the block as a new motion information candidate with the highest or lowest index to the second candidate list.
3. The method according to claim 2, wherein, If a first motion information candidate identical to the motion information of the block already exists in the second candidate list, deleting the first motion information candidate from the second candidate list, and adding the motion information of the block as the new motion information candidate to the second candidate list.
4. The method according to claim 1, wherein, Initializing the second candidate list in a predefined unit, and wherein, the predefined unit is a slice, a tile, a coding tree unit or a row of coding tree units.
5. The method according to claim 1, wherein, At least one of the multiple motion information candidates in the second candidate list is added to the first candidate list in descending order of the index assigned to each of the multiple motion information candidates.
6. A method for encoding a video, the method comprises: Deriving merge candidates for a current coded block, the merge candidates including spatial merge candidates derived according to neighboring blocks adjacent to the current coded block; Generating a first candidate list including the merge candidates for the current coded block, Wherein, when the number of the merge candidates included in the first candidate list is less than a predetermined value, the motion information candidates included in the second candidate list are added to the first candidate list; Select two merge candidates from the multiple merge candidates included in the first candidate list for triangular shape partitioning of two in the current coding block; Obtain a prediction block of the current coding block based on two motion information derived from the two merge candidates; Obtain a residual block of the current coding block based on the prediction block; Obtain residual coefficients by performing transformation and quantization on the residual block; and Encode two pieces of index information, wherein the motion information candidates included in the second candidate list are derived from spatial blocks, the spatial blocks are encoded before encoding the current coding block and are included in the current picture including the current coding block, wherein the first index information of the two pieces of index information specifies a first merge candidate among the multiple merge candidates, wherein the second index information of the two pieces of index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate, and wherein both the first index information and the second index information are encoded only by a prefix part.
7. An apparatus for transmitting compressed video data, the apparatus comprising: one or more hardware units configured to: obtain the compressed video data; and transmit the compressed video data, wherein obtaining the compressed video data includes: deriving merge candidates of a current coding block, the merge candidates including spatial merge candidates derived from adjacent blocks adjacent to the current coding block; generating a first candidate list including the merge candidates of the current coding block, wherein, when the number of the merge candidates included in the first candidate list is less than a predetermined value, the motion information candidates included in the second candidate list are added to the first candidate list; selecting two merge candidates from the multiple merge candidates included in the first candidate list for triangular shape partitioning of two in the current coding block; obtaining a prediction block of the current coding block based on two motion information derived from the two merge candidates; obtaining a residual block of the current coding block based on the prediction block; obtaining residual coefficients by performing transformation and quantization on the residual block; and encoding two pieces of index information, wherein the motion information candidates included in the second candidate list are derived from spatial blocks, the spatial blocks are encoded before encoding the current coding block and are included in the current picture including the current coding block, wherein the first index information of the two pieces of index information specifies a first merge candidate among the multiple merge candidates, wherein the second index information of the two pieces of index information specifies a second merge candidate among the remaining merge candidates other than the first merge candidate, and wherein both the first index information and the second index information are encoded only by a prefix part.