Inter prediction method and apparatus using the same

KR103025467B1Active Publication Date: 2026-09-29SK TELECOM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020190060392
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-06-18
Filing Date
2019-05-23
Publication Date
2026-09-29
Estimated Expiration
2039-05-23

Smart Images

  • Figure 112019052795989-PAT00006_ABST
    Figure 112019052795989-PAT00006_ABST
Patent Text Reader

Abstract

An inter prediction method and an image decoding device using the same are disclosed. According to one embodiment of the present invention, an inter prediction method is provided, comprising: a step of selecting a group indicated by group information decoded from a bitstream within a list of merge candidates in which merge candidates are classified into a plurality of groups; a step of selecting a merge candidate corresponding to a merge index decoded from the bitstream within the selected group; and a step of deriving movement information of a current block based on movement information of the selected merge candidate.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to the encoding and decoding of images, and more specifically, to an inter-prediction method that improves the efficiency of encoding and decoding by applying a new method for representing motion information, and an image decoding device using the same. Background Technology

[0002] Since video data contains a large amount of data compared to audio or still image data, storing or transmitting it as is without compression processing requires significant hardware resources, including memory.

[0003] Therefore, typically, when storing or transmitting video data, an encoder is used to compress the video data for storage or transmission, and a decoder receives the compressed video data, decompresses it, and plays it. Such video compression technologies include H.264 / AVC, as well as HEVC (High Efficiency Video Coding), which improves coding efficiency by about 40% compared to H.264 / AVC.

[0004] However, as video size, resolution, and frame rates are gradually increasing, and the amount of data that needs to be encoded is also growing accordingly, a new compression technology is required that offers better encoding efficiency and higher image quality improvement effects than existing compression technologies. The problem to be solved

[0005] To meet these requirements, the present invention aims to provide improved image encoding and decoding technology. In particular, one aspect of the present invention relates to a technology that improves the efficiency of encoding and decoding by increasing the number of merge candidates to improve the accuracy of the merge mode and reducing the number of bits required to represent motion information. means of solving the problem

[0006] One aspect of the present invention provides an inter prediction method characterized by comprising: a step of selecting a group indicated by group information decoded from a bitstream within a list of merge candidates in which merge candidates are classified into a plurality of groups; a step of selecting a merge candidate corresponding to a merge index decoded from the bitstream within the selected group; and a step of deriving movement information of a current block based on movement information of the selected merge candidate.

[0007] Another aspect of the present invention provides an image decoding device characterized by comprising: a group selection unit for selecting a group indicated by group information decoded from a bitstream within a merge candidate list in which merge candidates are classified into a plurality of groups; a candidate selection unit for selecting a merge candidate corresponding to a merge index decoded from the bitstream within the selected group; and a derivation unit for deriving movement information of a current block based on movement information of the selected merge candidate. Effects of the invention

[0008] As described above, according to one embodiment of the present invention, the optimal merge candidate can be selected using a merge candidate list consisting of an increased number of merge candidates, thereby improving the accuracy of the prediction.

[0009] In addition, according to another embodiment of the present invention, bit efficiency can be improved by applying a new binarization method that can reduce the number of bits required to represent an increased number of merge candidates. Brief explanation of the drawing

[0010] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the technologies of the present disclosure. Figure 2 is a diagram illustrating a method for dividing blocks using a QTBTTT structure. Figure 3 is a diagram illustrating multiple intra-prediction modes. FIG. 4 is an exemplary block diagram of an image decoding device capable of implementing the technologies of the present disclosure. Figure 5 is a diagram illustrating the locations of candidate blocks included in the candidate list. FIG. 6 is an exemplary block diagram of an inter-prediction unit capable of implementing the technologies of the present disclosure. Figure 7 is a flowchart illustrating an example of predicting the current block using a merge candidate list and a merge index. FIGS. 8 and 9 are drawings for illustrating various embodiments of the present invention that represent merge candidates by applying a new binarization method. FIG. 10 is a flowchart illustrating an embodiment of the present invention for selecting an optimal merge candidate based on a list of merge candidates classified into multiple groups. FIG. 11 is a diagram illustrating an embodiment of the present invention that binarizes a list of merge candidates by classifying merge candidates into multiple groups. FIG. 12 is a drawing for explaining an embodiment of the present invention regarding candidate blocks used in configuring a merge candidate list. FIG. 13 is a flowchart illustrating a conventional method for distinguishing prediction modes. FIG. 14 is a flowchart illustrating an embodiment of the present invention for distinguishing prediction modes. Specific details for implementing the invention

[0011] Hereinafter, some embodiments of the present invention will be described in detail with reference to exemplary drawings. It should be noted that in assigning identification symbols to the components of each drawing, the same components are assigned the same symbol whenever possible, even if they are shown in different drawings. Furthermore, in describing the present invention, if it is determined that a detailed description of related known components or functions could obscure the essence of the invention, such detailed description is omitted.

[0013] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the technologies of the present disclosure. Hereinafter, the image encoding device and its sub-components will be described with reference to FIG. 1.

[0014] The video encoding device may be configured to include a block division unit (110), a prediction unit (120), a subtractor (130), a conversion unit (140), a quantization unit (145), an encoding unit (150), an inverse quantization unit (160), an inverse conversion unit (165), an adder (170), a filter unit (180), and a memory (190).

[0015] Each component of the video encoding device may be implemented in hardware or software, or as a combination of hardware and software. Additionally, the function of each component may be implemented in software, and a microprocessor may be implemented to execute the software function corresponding to each component.

[0016] A single image (video) consists of multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles or / and slices. Here, one or more tiles can be defined as a tile group. Each tile or / slice is divided into one or more Coding Tree Units (CTUs). And each CTU is divided into one or more Coding Units (CUs) by a tree structure. Information applicable to each CU is encoded as the syntax of the CU, and information applicable to all CUs included in a single CTU is encoded as the syntax of the CTU. Additionally, information applicable to all blocks within a single tile is encoded as the syntax of the tile or as the syntax of the tile group, which is a collection of multiple tiles, and information applicable to all blocks constituting a single picture is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information that is commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). And, information that is commonly referenced by one or more SPSs is encoded in a Video Parameter Set (VPS).

[0017] The block division unit (110) determines the size of the Coding Tree Unit (CTU). Information regarding the size of the CTU (CTU size) is encoded as a syntax of SPS or PPS and transmitted to an image decoder.

[0018] The block division unit (110) divides each picture constituting the image into multiple Coding Tree Units (CTUs) having a predetermined size, and then recursively divides the CTUs using a tree structure. The leaf nodes in the tree structure become the coding units (CUs) which are the basic units of coding.

[0019] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of equal size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that combines two or more of these QT, BT, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT combined can be referred to as an MTT (Multiple-Type Tree).

[0020] FIG. 2 shows a QTBTTT splitting tree structure. As seen in FIG. 2, the CTU can first be split into a QT structure. Quadtree splitting can be repeated until the size of the splitting block reaches the minimum block size of the leaf node allowed in QT (MinQTSize). A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the encoding unit (150) and signaled to the image decoder. If the leaf node of the QT is not larger than the maximum block size of the root node allowed in BT (MaxBTSize), it can be further split into one or more of the BT structure or TT structure. In the BT structure and / or TT structure, multiple splitting directions may exist. For example, there may be two directions in which the block of the corresponding node is split horizontally and vertically. As shown in FIG. 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether the nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (Binary or Ternary) are encoded by the encoding unit (150) and signaled to the image decoding device.

[0021] As another example of a tree structure, when splitting a block using a QTBTTT structure, information such as a CU split flag (split_cu_flag) indicating that it has been split and a QT split flag (split_qt_flag) indicating whether the split type is QT split is encoded by the encoding unit (150) and signaled to the video decoder. If the CU split flag (split_cu_flag) value does not indicate that it has not been split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (coding unit), which is the basic unit of encoding. If the CU split flag (split_cu_flag) value does not indicate that it has been split, the split type is distinguished as QT or MTT through the QT split flag (split_qt_flag) value. When the split type is QT, there is no further additional information, and when the split type is MTT, a flag (mtt_split_cu_vertical_flag) indicating the MTT split direction (vertical or horizontal) and / or a flag (mtt_split_cu_binary_flag) indicating the MTT split type (binary or ternary) is additionally encoded by the encoding unit (150) and signaled to the image decoder.

[0022] When QTBT is used as another example of a tree structure, there may be two types: a type that divides the block of the corresponding node horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that divides it vertically (i.e., symmetric vertical splitting). A splitting flag (split_flag) indicating whether each node of the BT structure is divided into a block of a lower layer and splitting type information indicating the type of splitting are encoded by the encoding unit (150) and transmitted to the image decoding device. Meanwhile, there may also be an additional type that divides the block of the corresponding node into two blocks of an asymmetric shape. The asymmetric shape may include a shape that divides the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that divides the block of the corresponding node diagonally.

[0023] A CU can have various sizes depending on the QTBT or QTBTTT partition from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of QTBTTT) is referred to as the 'current block'.

[0024] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).

[0025] Generally, current blocks within a picture can each be predictively coded. Typically, the prediction of a current block can be performed using an intra-prediction technique (using data from the picture containing the current block) or an inter-prediction technique (using data from a picture coded prior to the picture containing the current block). Inter-prediction includes both unidirectional and bidirectional prediction.

[0026] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture containing the current block. Multiple intra prediction modes exist depending on the prediction direction. For example, as shown in FIG. 3, multiple intra prediction modes may include non-directional modes including planar mode and DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently for each prediction mode.

[0027] The intra prediction unit (122) can determine the intra prediction mode to use for encoding the current block. In some examples, the intra prediction unit (122) may encode the current block using several intra prediction modes and select an appropriate intra prediction mode to use from the tested modes. For example, the intra prediction unit (122) may calculate rate-distortion values ​​using rate-distortion analysis of several tested intra prediction modes and select an intra prediction mode that has the best rate-distortion features among the tested modes.

[0028] The intra prediction unit (122) selects one intra prediction mode among a plurality of intra prediction modes and predicts the current block using a calculation formula and surrounding pixels (reference pixels) determined according to the selected intra prediction mode. Information regarding the selected intra prediction mode is encoded by the encoding unit (150) and transmitted to the image decoding device.

[0029] The inter-prediction unit (124) generates a prediction block for the current block through a motion compensation process. It searches for the block most similar to the current block within a reference picture that is encoded and decoded before the current picture, and uses the searched block to generate a prediction block for the current block. Then, it generates a motion vector corresponding to the displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chroma component. Motion information, including information about the reference picture used to predict the current block and information about the motion vector, is encoded by the encoding unit (150) and transmitted to the image decoding device.

[0030] The subtractor (130) generates a residual block by subtracting the prediction block generated by the intra prediction unit (122) or the inter prediction unit (124) from the current block.

[0031] The conversion unit (140) converts residual signals within a residual block having pixel values ​​in a spatial domain into conversion coefficients in the frequency domain. The conversion unit (140) can convert residual signals within the residual block using the entire size of the residual block as the conversion unit, or it can divide the residual block into two sub-blocks, a conversion area and a non-conversion area, and convert residual signals using only the conversion area sub-block as the conversion unit. Here, the conversion area sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only the sub-block has been converted, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the encoding unit (150) and signaled to an image decoder. In addition, the size of the conversion area subblock may have a size ratio of 1:3 based on the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the division is additionally encoded by the encoding unit (150) and signaled to the image decoding device.

[0032] The quantization unit (145) quantizes the conversion coefficients output from the conversion unit (140) and outputs the quantized conversion coefficients to the encoding unit (150).

[0033] The encoding unit (150) generates a bitstream by encoding quantized transform coefficients using an encoding method such as CABAC (Context-based Adaptive Binary Arithmetic Code). The encoding unit (150) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division direction, and MTT division type, so that the video decoder can divide blocks in the same way as the video encoding device.

[0034] Additionally, the encoding unit (150) encodes information about a prediction type indicating whether the current block is encoded by intra prediction or by inter prediction, and encodes intra prediction information (i.e., information about the intra prediction mode) or inter prediction information (information about the reference picture and motion vector) according to the prediction type.

[0035] The inverse quantization unit (160) inversely quantizes the quantized transformation coefficients output from the quantization unit (145) to generate transformation coefficients. The inverse transformation unit (165) converts the transformation coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.

[0036] The adder (170) restores the current block by adding the restored residual block and the prediction block generated by the prediction unit (120). The pixels within the restored current block are used as reference pixels when intra-predicting the next block in sequence.

[0037] The filter unit (180) performs filtering on the restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. caused by block-based prediction and transformation / quantization. The filter unit (180) may include a deblocking filter (182) and a SAO (Sample Adaptive Offset) filter (184).

[0038] The deblocking filter (180) filters the boundaries between restored blocks to remove blocking artifacts caused by block-unit encoding / decoding, and the SAO filter (184) performs additional filtering on the deblocking filtered image. The SAO filter (184) is a filter used to compensate for the difference between the restored pixel and the original pixel caused by lossy coding.

[0039] The restored blocks filtered through the deblocking filter (182) and the SAO filter (184) are stored in memory (190). When all blocks within a picture are restored, the restored picture is used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0041] FIG. 4 is an exemplary block diagram of an image decoding device capable of implementing the technologies of the present disclosure. Hereinafter, the image decoding device and its sub-components will be described with reference to FIG. 4.

[0042] The image decoding device may be configured to include a decoding unit (410), an inverse quantization unit (420), an inverse transformation unit (430), a prediction unit (440), an adder (450), a filter unit (460), and a memory (470).

[0043] Similar to the image encoding device of FIG. 1, each component of the image decoding device may be implemented in hardware or software, or in combination of hardware and software. Additionally, the function of each component may be implemented in software, and a microprocessor may be implemented to execute the function of the software corresponding to each component.

[0044] The decoding unit (410) decodes the bitstream received from the video encoding device to extract information related to block division, thereby determining the current block to be decoded, and extracts prediction information and information regarding residual signals necessary to restore the current block.

[0045] The decoding unit (410) extracts information about the CTU size from the SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) to determine the size of the CTU and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and divides the CTU using the tree structure by extracting division information for the CTU.

[0046] For example, when splitting a CTU using a QTBTTT structure, first, a first flag (QT_split_flag) related to QT splitting is extracted to split each node into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, a second flag (MTT_split_flag) related to MTT splitting and splitting direction (vertical / horizontal) and / or splitting type (binary / ternary) information are extracted to split the corresponding leaf nodes into an MTT structure. Through this, each node below the leaf nodes of QT is recursively split into a BT or TT structure.

[0047] As another example, when splitting a CTU using the QTBTTT structure, a CU split flag (split_cu_flag) indicating whether the CU is split is first extracted, and if the block is split, a QT split flag (split_qt_flag) is extracted. If the split type is not QT but MTT, a flag (mtt_split_cu_vertical_flag) indicating the MTT split direction (vertical or horizontal) and / or a flag (mtt_split_cu_binary_flag) indicating the MTT split type (Binary or Ternary) are additionally extracted. During the splitting process, each node may undergo zero or more iterative MTT splits following zero or more iterative QT splits. For example, a CTU may undergo MTT splitting immediately, or conversely, only multiple QT splits.

[0048] As another example, when splitting a CTU using a QTBT structure, a first flag (QT_split_flag) related to the splitting of QT is extracted to split each node into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and split direction information are extracted.

[0049] Meanwhile, when the decoding unit (410) determines the current block to be decoded through the division of the tree structure, it extracts information regarding the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-predicted, the decoding unit (410) extracts syntax elements for the intra-predicted information (intra-predicted mode) of the current block. If the prediction type information indicates inter-predicted, the decoding unit (410) extracts syntax elements for the inter-predicted information, namely information indicating the motion vector and the reference picture that the motion vector refers to.

[0050] Meanwhile, the decoding unit (410) extracts information about the quantized transformation coefficients of the current block as information about the residual signal.

[0051] The inverse quantization unit (420) inversely quantizes the quantized transformation coefficients, and the inverse transformation unit (430) inversely transforms the inversely quantized transformation coefficients from the frequency domain to the spatial domain to restore the residual signals, thereby generating a residual block for the current block.

[0052] Additionally, when the inverse transformation unit (430) inversely transforms only a part of the transformation block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transformation block has been transformed, information on the directionality (vertical / horizontal) of the sub-block (cu_sbt_horizontal_flag) and / or information on the position of the sub-block (cu_sbt_pos_flag), restores residual signals by inversely transforming the transformation coefficients of the corresponding sub-block from the frequency domain to the spatial domain, and generates a final residual block for the current block by filling the area that has not been inversely transformed with a "0" value as the residual signal.

[0053] The prediction unit (440) may include an intra prediction unit (442) and an inter prediction unit (444). The intra prediction unit (442) is activated when the prediction type of the current block is an intra prediction, and the inter prediction unit (444) is activated when the prediction type of the current block is an inter prediction.

[0054] The intra prediction unit (442) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the decoding unit (410), and predicts the current block using reference pixels around the current block according to the intra prediction mode.

[0055] The inter prediction unit (444) determines the motion vector of the current block and the reference picture that the motion vector refers to using the syntax elements for the intra prediction mode extracted from the decoding unit (410), and predicts the current block using the motion vector and the reference picture.

[0056] The adder (450) restores the current block by adding the residual block output from the inverse transformation unit and the prediction block output from the inter prediction unit or the intra prediction unit. The pixels within the restored current block are used as reference pixels when intra-predicting the block to be decoded later.

[0057] The filter unit (460) may include a deblocking filter (462) and an SAO filter (464). The deblocking filter (462) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-unit decoding. The SAO filter (464) performs additional filtering on the restored blocks after deblocking filtering to compensate for the difference between the restored pixels and the original pixels caused by lossy coding. The restored blocks filtered through the deblocking filter (462) and the SAO filter (464) are stored in memory (470). When all blocks within a picture are restored, the restored picture is used as a reference picture to inter-predict blocks within the picture to be encoded later.

[0059] Inter-frame prediction encoding / decoding methods (inter-prediction methods) can be broadly classified into skip mode, merge mode, and AMVP (adaptive (or advanced) motion vector predictor) mode.

[0060] In skip mode, motion information of one of the motion information candidates of surrounding blocks is transmitted from the video encoder to the video decoder. In merge mode, motion information of one of the motion information candidates of surrounding blocks and information encoded with the prediction residual are transmitted. In AMVP mode, motion information of the current block and information encoded with the prediction residual are transmitted.

[0061] Motion information for Skip Mode and Merge Mode is represented by an index value indicating one of the candidates included in the merge candidate list, and motion information for AMVP Mode is represented by the difference value (mvd, motion vector difference) between the motion information of surrounding blocks and the motion information of the current block.

[0062] A method for constructing a merge candidate list for skip mode and merge mode may involve steps such as searching for spatial candidates and adding them to the merge candidate list, searching for temporal candidates and adding them to the merge candidate list, combining the candidates added to the merge candidate list (combined bi-directional candidate), and adding zero motion vector candidates to the merge candidate list.

[0063] The locations of candidate blocks for spatial and temporal candidates added to the merge candidate list are shown in FIG. 5. FIG. 5 (A) shows the locations of candidate blocks for spatial candidates, and FIG. 5 (B) shows the locations of candidate blocks for temporal candidates.

[0064] The video encoding device and the video decoder can search candidate blocks in the order A1→B1→B0→A0→B2 to include up to four spatial candidates in the merge candidate list. Additionally, the video encoding device and the video decoder can search candidate blocks in the order BR→CT to include up to one temporal candidate in the merge candidate list. During the process of setting the merge candidate list, if there is duplicate motion information, that motion information is not added to the list. In other words, duplication of motion information is not allowed.

[0065] Information regarding the number of candidates that can be included in the merge candidate list is recorded in the slice header, and up to five candidates can be specified based on this information. The TU (truncated unary) method is used as the binaryization method for the candidate indices, and Table 1 below shows the codewords representing each candidate index when the maximum number of candidates is five.

[0066] Index Codeword Bits 0 0 1 1 10 2 2 110 3 3 1110 4 4 1111 4

[0068] FIG. 6 is an exemplary block diagram of an inter prediction unit (444) capable of implementing the technologies of the present disclosure. As illustrated in FIG. 6, the inter prediction unit (444) may be configured to include a selection unit (610), a derivation unit (620), a prediction execution unit (630), a selection unit (640), and a determination unit (650), and the selection unit (610) may be configured to include a group selection unit (612) and a candidate selection unit (614).

[0069] The selection unit (640) and the determination unit (650) determine a merge candidate list consisting of one or more merge candidates using information regarding the number of candidates, which is the number of merge candidates (S710). Here, information regarding the number of candidates can be obtained through a decoding process in the decoding unit (410) after being transmitted from the video encoding device. The video encoding device and the video decoding device can determine or configure the merge candidate list through the same method or process.

[0070] The selection unit (610) selects a merge candidate (indicated by the merge index) corresponding to the merge index from the merge candidate list (S720). The merge index is information indicating one of the merge candidates included in the merge candidate list, and the merge candidate indicated by the merge index is the optimal merge candidate among the merge candidates included in the merge candidate list that has the most similar movement information to the current block.

[0071] The derivation unit (620) derives movement information of the current block based on the movement information of the selected merge candidate (S730), and the prediction execution unit (630) generates a prediction block for the current block using the derived movement information of the current block (S740).

[0072] Conventional video encoding devices or video decoding devices are configured to determine a merge candidate list consisting of up to five merge candidates. The video encoding device and video decoding device of the present invention may be configured to determine a merge candidate list with five or fewer merge candidates, or to determine a merge candidate list with six or more merge candidates, in the same manner as the conventional method. That is, in the present invention, information regarding the number of candidates transmitted from the video encoding device to the video decoding device may have a natural number of six or more, and the video decoding device may determine a merge candidate list with a number of merge candidates corresponding (identical) to the number of transmitted candidates.

[0073] In this way, when the present invention determines or configures a merge candidate list with a larger number of merge candidates, the range of choices for deriving movement information of the current block becomes relatively wider compared to conventional methods. Accordingly, the present invention enables the optimal movement information that is more closely matched with the actual movement information of the current block to be used as the movement information of the current block, thereby allowing the movement information of the current block to be estimated more accurately.

[0074] Meanwhile, conventional image encoding and image decoding devices formed a merge candidate list with five or fewer merge candidates, and at the same time, used a truncated unary (TU) binary conversion method to represent the indices (merge indices) of the merge candidates included in the merge candidate list.

[0075] Conventional methods construct a merge candidate list using a relatively small number (maximum 5) of merge candidates, so even when using a TU binaryization method in which the number of bits required to represent the merge index increases proportionally to the number of candidates, not many bits are required to represent or signal the merge index, so no major problem arises in terms of bit efficiency.

[0076] However, as previously mentioned, if a relatively large number (6 or more) of merge candidates are used as a merge candidate list to facilitate accurate estimation of motion information, the number of bits required for the merge index representation increases due to the characteristics of the TU binarization method, which may reduce bit efficiency.

[0077] Accordingly, the present invention presents a new method that can simultaneously achieve accurate estimation of motion information and bit efficiency by applying a binarization method other than the TU binarization method in addition to the TU binarization method in the aforementioned embodiment that increases the number of merge candidates, or by applying a combination of the TU binarization method and a binarization method other than the TU binarization method.

[0078] Other binarization methods may include various binarization methods such as the EG (exponential golomb) binarization method, the TR (truncated rice) binarization method, and the FL (fixed length) binarization method. Hereinafter, a encoding method including one or more of these new binarization methods will be referred to as the second method, and the TU binarization method corresponding to the conventional binarization method will be referred to as the first method.

[0080] FIGS. 8 and 9 are drawings illustrating various embodiments of the present invention for representing merge candidates by applying a new binarization method. FIG. 8 shows a table comparing the number of bits when the index of motion information is encoded by the TU binarization method, i.e., the first method, and the number of bits when encoded by the EG binarization method among the second methods. Here, the value of k for EG is “0”.

[0081] Regarding the internal structure of the table, the codeword encoded by the TU binary method and the codeword encoded by the EG binary method are shown, the number of bits required to represent each codeword is shown for TU and EG respectively, and the difference between the number of bits by the TU binary method and the number of bits by the EG binary method is shown.

[0082] As shown in FIG. 8, there is no difference in the number of bits between the TU binary method and the EG binary method at indices 0, 2, and 4 (0), but the number of bits of the EG binary method increases at indices 1 and 3 (+1), and the number of bits of the EG binary method decreases at the remaining indices (index 5 or higher). When the present invention is configured to form a merge candidate list with six or more merge candidates, the effect of decreasing the number of bits for all indices after index 4 (index 5 or higher) can be obtained.

[0083] That is, when a merge index of index 5 or higher is transmitted from an image encoding device to an image decoding device, an embodiment of the present invention that encodes the merge index using the EG binary method can reduce the number of bits required to represent the merge index compared to a conventional method that encodes using the TU binary method.

[0084] FIG. 9 shows a table for comparing the number of bits when the index of motion information is encoded using the TU binary method and the number of bits when encoded using the TR binary method. The internal configuration of the table shown in FIG. 9 is identical to the internal configuration of the table shown in FIG. 8. Here, the cRiceParam value for TR is “1” and the cMax value is “10”.

[0085] As shown in Fig. 9, when the TR binary conversion method is applied, there is no difference between the number of bits representing indices 1 and 2 (0), the number of bits representing index 0 increases (+1), and the number of bits representing the remaining indices decreases.

[0086] Looking at an embodiment of the present invention that forms a merge candidate list with six or more merge candidates, the effect of reducing the number of bits for all indices after index 4 is shown. In addition, even when a merge candidate list is formed with five or fewer merge candidates, the number of bits required to represent indices 3 and 4 is reduced.

[0087] In the above, the results of applying the EG binarization method and the TR binarization method among the encoding methods were explained, but various binarization methods such as the FL binarization method can be applied as new binarization methods.

[0088] As described using FIGS. 8 and 9, in the case of an embodiment in which a merge index is encoded using the second method, which is a new binarization method, the reduction in the number of bits can be said to occur variably depending on the number of merge indices included in the merge candidate list. For example, when the new binarization method EG binarization method is applied, the effect of reducing the number of bits can be obtained only when 6 or more merge indices are included in the merge candidate list. In addition, when the TR binarization method is applied, the effect of reducing the number of bits can be obtained only when 4 or more merge candidates are included in the merge candidate list.

[0089] Based on this, the present invention can variably determine the method of encoding the merge index according to the number of merge indices included in the merge candidate list, that is, the number of merge candidates indicated by information regarding the number of candidates.

[0090] For example, the present invention may encode the merge index using any one of the second methods when the number of merge candidates included in the merge candidate list is 6 or more, and may encode the merge index using the TU binary method when the number of merge candidates included in the merge candidate list is 5 or fewer. As another example, the present invention may encode the merge index using the TR binary method when the number of merge candidates included in the merge candidate list is 4 to 5.

[0092] FIG. 10 is a flowchart illustrating an embodiment of the present invention for selecting an optimal merge candidate based on a list of merge candidates classified into multiple groups, FIG. 11 is a diagram illustrating an embodiment of the present invention for binarizing a list of merge candidates by classifying the merge candidates into multiple groups, and FIG. 12 is a diagram illustrating an embodiment of the present invention regarding candidate blocks used in configuring a list of merge candidates.

[0093] The present invention can be configured to select the optimal merge candidate by classifying merge candidates included in a merge candidate list into multiple groups and additionally signaling information distinguishing these multiple groups.

[0094] The video encoding device divides the merge candidates included in the merge candidate list into multiple groups, and when the optimal merge candidate is selected, it transmits information (group information) indicating which group the candidate belongs to, and then transmits the index value (merge index) for the optimal merge candidate within that group.

[0095] The decoding unit (410) parses and decodes group information from the bitstream (S1030), and the group selection unit (612) included in the selection unit (610) determines or selects the group indicated by the decoded group information (S1040). Here, the group information is information indicating one of a plurality of groups, and the group indicated by this group information corresponds to a group containing a merge candidate (corresponding to the merge index) to be used to derive movement information of the current block.

[0096] When the merge candidate list consists of two groups, the group information may consist of a flag (MPC_flag) indicating the first of the two groups, as shown in FIG. 10. The merge index may consist of a form (MPC_idx) indicating one of the merge candidates classified into this first group and a form (other_idx) indicating a merge candidate classified into the other group.

[0097] If the merge candidate list consists of three or more groups, the group information may consist of index information indicating any one of the three or more groups, and the merge index may be in a form indicating any one of the merge candidates classified into the group indicated by the group information.

[0098] When the group indicated by the group information is determined, the decoding unit (410) parses and decodes the merge index from the bitstream (S1050, S1070), and the candidate selection unit (614) included in the selection unit (610) selects a merge candidate corresponding to the merge index from the indicated group.

[0099] In an embodiment where the merge candidate list is composed of two groups, if the group information indicates the first group (MPC_flag=1), the candidate selection unit (614) selects a merge candidate corresponding to the merge index (MPC_idx) among the merge candidates classified into this first group. Conversely, if the group information does not indicate the first group (MPC_flag≠1), the candidate selection unit (614) selects a merge candidate corresponding to the merge index (other_idx) among the merge candidates classified into the other group.

[0100] The derivation unit (620) derives movement information of the current block based on the movement information of the selected merge candidate, and the prediction execution unit (630) generates a prediction block for the current block using the derived movement information of the current block.

[0101] As described below, the present invention proposes a new method for distinguishing prediction modes, and among the entire process shown in FIG. 10, process S1010, process S1020, process S1060, and process S1080 correspond to the results to which this new method is applied. A detailed description of the new method for distinguishing prediction modes will be provided below.

[0102] The table in Fig. 11 shows a list of merge candidates composed of two groups and a comparison of the number of bits between the case where the merge index is encoded using only the TU binary method and the case where the merge index is encoded using a combination of group information consisting of flags and the TU and FL binary methods.

[0103] The flag bit value “1” in the codeword of index 0 to 2 indicates that the corresponding index, i.e., the corresponding merge candidate, belongs to Group I, and the flag bit value “0” in the codeword of index 3 to 10 indicates that the corresponding merge candidate belongs to Group II.

[0104] As shown in Fig. 11, if merge candidates are divided into two groups and either the TU or FL binarization method is applied to each group, the same or relatively fewer bits (including flag bits) are required in Group II compared to when only the TU binarization method is applied.

[0105] Although only an embodiment in which multiple groups are encoded using a combination of the TU binary method and the FL binary method has been described through FIG. 11, the aforementioned EG binary method or TR binary method may also be applied. To generalize this point, the merge index can be encoded using different binary methods for each of the multiple groups included in the merge candidate list (for each group indicated by the group information).

[0106] Meanwhile, regarding surrounding blocks (candidate blocks) that may be included in the merge candidate list, the conventional method used all or part of the left block, top block, top-right block, bottom-left block, and top-left block adjacent to the current block within the current picture as candidate blocks. Additionally, the conventional method used collocated blocks located within a collocated reference picture, rather than the current picture containing the current block, as candidate blocks. For example, it further used blocks located at the same point as the current block within the collocated reference picture (collocated block) or blocks adjacent to this block at the same location as candidate blocks. Here, information regarding the collocated reference picture can be transmitted in the upper header (e.g., slice header).

[0107] The present invention proposes a new example (new surrounding block) of surrounding blocks that can be included in a merge candidate list. In the present invention, an example of a surrounding block that can be used to construct a merge candidate list including the new surrounding blocks is illustrated in FIG. 12.

[0108] FIG. 12 (a) is an example of a spatial surrounding block, FIG. 12 (b) is an example of a temporal surrounding block, FIG. 12 (c) is an example of a central surrounding block, and FIG. 12 (d) is an example of a non-adjacent surrounding block.

[0109] Spatial surrounding blocks refer to surrounding blocks located adjacent to the current block (1200) and located at the edges of the corners constituting the current block (1200). Specifically, as illustrated in FIG. 12 (a), surrounding blocks (A1) located to the left of the current block (1200), surrounding blocks (B1) located at the top, surrounding blocks (B0) located at the upper right, surrounding blocks (A0) located at the lower left, and surrounding blocks (B2) located at the upper left may be included in the spatial surrounding blocks.

[0110] The temporal peripheral block refers to a peripheral block located in the collocated reference picture. As illustrated in FIG. 11 (b), this temporal peripheral block may include a peripheral block (CT) located within the collocated block (1210), a peripheral block (BR) located at the lower right of the collocated block (1210), a peripheral block (TR) located to the right of the collocated block (1110), and a peripheral block (BL) located at the bottom of the collocated block (1110), centered around the collocated block (1210) located at the same point as the current block within the collocated reference picture.

[0111] The central peripheral block refers to a peripheral block located adjacent to the current block (1200) but located in the center of the corner constituting the current block (1200). As illustrated in FIG. 12 (c), this central peripheral block may include blocks (T1, T2) located in the upper center of the current block (1200) and blocks (L1, L2) located in the left center.

[0112] Non-adjacent surrounding blocks refer to surrounding blocks located at a preset distance from the current block (1200). As illustrated in FIG. 12 (d), non-adjacent surrounding blocks may include blocks (N0, N1, N2, and N3) located at a preset distance from the current block (1200) among blocks located not adjacent to the current block (1200). For example, if the size of the current block (1200) is 16x16 and motion information is stored in memory (190, 470) in 4x4 units, one or more of the surrounding blocks (N0, N1, N2, and N3) located at a distance of 4 pixels from the current block (1200) may be used as non-adjacent surrounding blocks. Here, the preset distance between the current block (1200) and the non-adjacent blocks may be stored in memory (190, 470).

[0113] The selection unit (640) selects one or more surrounding blocks (merge candidates) from spatial surrounding blocks, temporal surrounding blocks, central surrounding blocks and non-adjacent surrounding blocks, and can select a number of surrounding blocks equal to the number of candidates transmitted from the image encoding device.

[0114] When the selection of surrounding blocks is completed, the decision unit (650) classifies the selected merge candidates into multiple groups to determine the merge candidate list.

[0115] The process of classifying surrounding blocks into multiple groups can be performed by distinguishing spatial surrounding blocks, temporal surrounding blocks, central surrounding blocks, and non-adjacent surrounding blocks. Specifically, if the merge candidate list consists of two groups, merge candidates selected from spatial surrounding blocks may be classified into Group I, and merge candidates selected from temporal surrounding blocks, merge candidates selected from central surrounding blocks, and merge candidates selected from non-adjacent surrounding blocks may be classified into Group II. As another example, merge candidates selected from spatial surrounding blocks and merge candidates selected from temporal surrounding blocks may be classified into Group I, and merge candidates selected from central surrounding blocks and merge candidates selected from non-adjacent surrounding blocks may be classified into Group II. As yet another example, instead of classifying the four types of surrounding blocks into groups, a certain number of surrounding blocks among the total selected surrounding blocks may be classified into Group I, and the remainder into Group II.

[0117] FIG. 13 is a flowchart illustrating a conventional method for distinguishing prediction modes, and FIG. 14 is a flowchart illustrating an embodiment of the present invention for distinguishing prediction modes.

[0118] As illustrated in FIG. 13, a conventional method for distinguishing prediction modes has flags that distinguish between skip mode and merge mode, and signals a merge index according to the values ​​of these flags. The method of constructing a merge candidate list or the method of signaling a merge index can be performed identically in skip mode and merge mode. However, in the case of skip mode, the size of the corresponding block is 2Nx2N, whereas in the case of merge mode, the size of the corresponding block may be not only 2Nx2N but also 2NxN, Nx2N, or asymmetric partition.

[0119] If the block has a size of 2Nx2N and possesses all-zero conversion factors, the block is classified as skip mode. Therefore, for the block with a size of 2Nx2N to be classified as merge mode, it must have at least one non-zero conversion factor. Whether the block has one or more non-zero conversion factors can be indicated via rqt_root_cbf. In skip mode, rqt_root_cbf is not signaled and its value is set to “0”; in merge mode, if the block has a size of 2Nx2N, rqt_root_cbf is not signaled and its value is set to “1”. In merge mode, if the block does not have a size of 2Nx2N, rqt_root_cbf is explicitly signaled.

[0120] The flag and merge index values ​​for the skip mode and merge mode used in the conventional method are shown in Tables 2 and 3 below.

[0121] coding_unit( x0, y0, log2CbSize ) { Descriptor if( slice_type != I ) cu_skip_flag[ x0 ][ y0 ] ae(v) if( cu_skip_flag[ x0 ][ y0 ] ) prediction_unit( x0, y0, nCbS, nCbS ) else { if( slice_type != I ) pred_mode_flag ae(v) … if( CuPredMode[ x0 ][ y0 ] != MODE_INTRA | | log2CbSize = = MinCbLog2SizeY ) part_mode ae(v) if( CuPredMode[ x0 ][ y0 ] = = MODE_INTRA ) { … } else { if( PartMode = = PART_2Nx2N ) prediction_unit( x0, y0, nCbS, nCbS ) … } … } …

[0122] prediction_unit( x0, y0, nPbW, nPbH ) { Descriptor if( cu_skip_flag[ x0 ][ y0 ] ) { if( MaxNumMergeCand > 1 ) merge_idx[ x0 ][ y0 ] ae(v) } else { / * MODE_INTER * / merge_flag[ x0 ][ y0 ] ae(v) if( merge_flag[ x0 ][ y0 ] ) { if( MaxNumMergeCand > 1 ) merge_idx[ x0 ][ y0 ] ae(v) } else { … }

[0123] As illustrated in FIG. 13, a conventional method for distinguishing prediction modes uses a flag (skip_flag) indicating whether it corresponds to a skip mode and a flag (merge_flag) indicating whether it corresponds to a merge mode. Accordingly, processing (S1310, S1320, S1344, and S1346) for decoding and analyzing the flags is implemented separately.

[0124] In addition, in the conventional method, to determine whether the current block is predicted to be in merge mode, the process involves decoding and analyzing a flag (skip_flag) indicating whether it corresponds to skip mode (S1310, S1320), decoding and analyzing a flag (pred_mode_flag) indicating whether it corresponds to inter mode (S1330, S1340), and decoding and analyzing a flag (merge_flag) indicating whether it corresponds to merge mode (S1344, S1346).

[0125] As such, the conventional method is configured to distinguish and determine prediction modes through relatively more processing compared to the present invention described below, and thus has the problem of reduced efficiency in image encoding and decoding.

[0126] The present invention can solve these problems by configuring it to determine, through a single processing step, whether the current block corresponds to skip mode or merge mode, and to determine the prediction mode of the current block through a relatively small number of steps. This specification is based on the assumption that the size of the block to be encoded / decoded is MxN (where M and N may be the same), and the size of the target block is constant and does not vary depending on the prediction mode.

[0127] First, as shown in FIG. 14, the merge_flag transmitted from the video encoding device is parsed and decoded (S1410). Here, the merge_flag indicates the mode used for predicting the current block among the first mode, which integrates the skip mode and the merge mode, and the second mode, which integrates the inter mode and the intra mode.

[0128] As a result of analyzing or interpreting the decrypted merge_flag (S1420), if the merge_flag indicates a first mode, the merge_idx, which is an index indicating one of the merge candidates included in the merge candidate list, is parsed and decrypted (S1430). In addition, to determine which prediction mode among the first modes (skip mode and merge mode) the current block is predicted to be, cu_cbf, which indicates whether all transformation coefficients correspond to 0 (whether non-zero transformation coefficients exist), is parsed and decrypted (S1432), and the decrypted cu_cbf is analyzed (S1434).

[0129] If cu_cbf indicates a transform factor that is all zero (0), this means that the current block is predicted to be in skip mode by the video encoding device. Therefore, the current block is predicted to be in skip mode. In contrast, if cu_cbf indicates a non-zero transform factor (1), this means that the current block is predicted to be in merge mode by the video encoding device. Therefore, the current block is predicted to be in merge mode.

[0130] In this way, the present invention is configured to distinguish the prediction mode through more simplified processes, such as the process of determining the prediction mode of the current block among the second mode and the first mode using merge_flag, and the process of distinguishing the skip mode and the merge mode using cu_cbf, so that the efficiency of video encoding and decoding can be improved.

[0131] Again, going back to the step of analyzing merge_flag (S1420), if merge_flag indicates a second mode in which the inter (AMVP) mode and intra mode are integrated, a flag (pred_mode_flag) indicating which prediction mode the current block was predicted to be between the inter mode and the intra mode is parsed and decoded (S1440), and the decoded pred_mode_flag is analyzed (S1442).

[0132] If pred_mode_flag indicates intra mode, cu_cbf is set to 1 (S1446), and the current block is predicted to be in intra mode. In contrast, if pred_mode_flag indicates inter mode, after cu_cbf is parsed and decoded (S1444), the current block is predicted to be in AMVP mode.

[0133] An example of the combination of the “method for integrally distinguishing prediction modes” described above and the “method for grouping merge candidate lists” described earlier is explained again below using FIG. 10. The example illustrated in FIG. 10 corresponds to an example in which merge candidate lists are classified into two groups.

[0134] First, a process of decoding and analyzing merge_flag to distinguish between the first mode and the second mode (S1010, S1020) is performed. If the prediction mode of the current block is determined to be the first mode, MPC_flag, which is group information indicating the first group among two groups (group information indicating whether the optimal merge candidate belongs to group I), is decoded and analyzed (S1030, S1040).

[0135] If the optimal merge candidate belongs to group I (MPC_flag=1), the merge index MPC_idx, which indicates the optimal merge candidate, is parsed and decoded (S1050), and the optimal merge candidate (the merge candidate corresponding to the merge index) is selected from group I using MPC_idx. In contrast, if the merge candidate belongs to another group (group II) (MPC_flag=0), the merge index (other_idx), which indicates the optimal merge candidate (indicating any one of the merge candidates belonging to group II), is parsed and decoded (S1070), and the optimal merge candidate is selected from group II using other_idx.

[0136] When the optimal merge candidate is selected, cu_cbf is decoded and analyzed (S1060, S1080), and a process of distinguishing between skip mode and merge mode is performed.

[0138] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment. Explanation of the symbols

[0139] 120, 440: Prediction unit 130: Subtractor 170: Adder 180, 460: Filter section

Claims

Claim 1 An inter-prediction method comprising: a step of encoding first information; a step of encoding a merge index for indicating a merge candidate of a current block within a merge candidate group corresponding to the first information among a first merge candidate group and a second merge candidate group composed of merge candidates in a merge candidate list; and a step of predicting the current block based on movement information of the merge candidate of the current block, wherein the merge index is encoded using a different binarization method for each merge candidate group. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 A video decoding method comprising: decoding first information from a bitstream and selecting one of a first merge candidate group and a second merge candidate group formed from merge candidates in a merge candidate list based on the first information; decoding a merge index for indicating a merge candidate of the current block within the merge candidate group corresponding to the first information; and deriving movement information of the current block based on movement information of a merge candidate corresponding to the merge index, wherein the merge index is encoded using a different binary method for each merge candidate group. Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 A method for providing image data to a decoding device, comprising: a step of encoding the image data to generate a bitstream; and a step of transmitting the bitstream to the decoding device, wherein the step of generating the bitstream comprises: a step of encoding first information; a step of encoding a merge index for indicating a merge candidate of a current block within a merge candidate group corresponding to the first information among a first merge candidate group and a second merge candidate group composed of merge candidates in a merge candidate list; and a step of predicting the current block based on movement information of the merge candidate of the current block, wherein the merge index is encoded using a different binaryization method for each merge candidate group.

Citation Information

Patent Citations

  • Spatial prediction based intra coding

    KR1020050007607A

  • Video Coding Method and Apparatus using Improved Merge

    KR1020150122106A