Image encoding / decoding method and device for transmitting compressed video data

By constructing an improved motion information merge list with additional candidates and deriving motion information from multiple lists, the method addresses the challenge of high-resolution video data volume and quality, enhancing encoding/decoding efficiency and prediction accuracy.

WO2026063733A1PCT designated stage Publication Date: 2026-03-26KT CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video data leads to higher data volumes, increasing transmission and storage costs, and existing video compression technologies struggle to efficiently handle stereoscopic video content.

Method used

A method for constructing an improved motion information merge list by adding additional motion information merging candidates and deriving motion information based on multiple motion information merge lists, using a combination of motion information candidates to enhance inter-prediction.

Benefits of technology

Improves encoding/decoding efficiency and prediction accuracy by utilizing additional motion information merging candidates and selecting optimal combinations for inter-prediction, reducing data volume and enhancing video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025014673_26032026_PF_FP_ABST
    Figure KR2025014673_26032026_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method according to the present disclosure may comprise the steps of: configuring a motion information merge list for the current block; deriving motion information about the current block on the basis of the motion information merge list; and acquiring a prediction block of the current block on the basis of the motion information about the current block. Here, the motion information merge list may be derived by adding a first additional motion information merge candidate to an initial motion information merge list, the first additional motion information merge candidate may be derived on the basis of motion information merge candidates included in the initial motion information merge list, and a reference picture index for a prescribed direction of the first additional motion information merge candidate may be set to indicate that a template cost is the smallest among reference pictures for the prescribed direction of the motion information merge candidates.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device for transmitting compressed video data

[0001] The present disclosure relates to a video signal processing method and apparatus.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields. As video data becomes higher in resolution and quality, the relative volume of data increases compared to conventional video data; consequently, transmission and storage costs increase when video data is transmitted using existing wired or wireless broadband lines or stored using conventional storage media. To address these issues arising from the increase in video data resolution and quality, high-efficiency video compression technologies can be utilized.

[0003] Various video compression technologies exist, such as inter-frame prediction technology that predicts pixel values ​​in the current picture from previous or subsequent pictures, intra-frame prediction technology that predicts pixel values ​​in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values ​​with high frequency and long codes to values ​​with low frequency; by utilizing these video compression technologies, video data can be effectively compressed for transmission or storage.

[0004] Meanwhile, along with the increasing demand for high-resolution video, the demand for stereoscopic video content as a new video service is also rising. Discussions are underway regarding video compression technologies to effectively provide high-resolution and ultra-high-resolution stereoscopic video content.

[0005] The present disclosure aims to provide a method for constructing an improved motion information merge list and an apparatus for the same.

[0006] The present disclosure aims to provide a method for inducing additional motion information merging candidates and an apparatus for the same by utilizing a plurality of motion information merging candidates already inserted in a motion information merging list.

[0007] The present disclosure aims to provide a method and apparatus for performing inter-prediction based on selecting one of a plurality of combinations, each comprising a plurality of motion information candidates, and based on the selected combination.

[0008] The present disclosure aims to provide a method for deriving motion information of a current block based on a plurality of motion information merge lists and an apparatus for the same.

[0009] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.

[0010] A video decoding method according to the present disclosure may include: a step of configuring a motion information merging list for a current block; a step of deriving motion information for the current block based on the motion information merging list; and a step of obtaining a predicted block of the current block based on the motion information of the current block. In this case, the motion information merging list is derived by adding a first additional motion information merging candidate to an initial motion information merging list, and the first additional motion information merging candidate is derived based on motion information merging candidates included in the initial motion information merging list, and the reference picture index for a predetermined direction of the first additional motion information merging candidate may be set to indicate the one with the smallest template cost among the reference pictures for a predetermined direction of the motion information merging candidates.

[0011] A video encoding method according to the present disclosure may include: a step of configuring a motion information merging list for a current block; a step of deriving motion information for the current block based on the motion information merging list; and a step of obtaining a predicted block of the current block based on the motion information of the current block. In this case, the motion information merging list is derived by adding a first additional motion information merging candidate to an initial motion information merging list, and the first additional motion information merging candidate is derived based on motion information merging candidates included in the initial motion information merging list, and the reference picture index for a predetermined direction of the first additional motion information merging candidate may be set to indicate the one with the smallest template cost among the reference pictures for a predetermined direction of the motion information merging candidates.

[0012] In the image encoding / decoding method according to the present disclosure, the motion vector for the predetermined direction of the first additional motion information merging candidate is derived by averaging the motion vectors for the predetermined direction of the plurality of motion information merging candidates, and the predetermined direction may be the L0 direction or the L1 direction.

[0013] In the image encoding / decoding method according to the present disclosure, if at least one of the plurality of motion information merging candidates does not have motion information for the predetermined direction, the at least one motion information merging candidate may not be used to derive the motion vector for the predetermined direction of the first additional motion information merging candidate.

[0014] In the image encoding / decoding method according to the present disclosure, if at least one of the plurality of motion information merging candidates does not have motion information for the predetermined direction, the motion vector for the predetermined direction of the at least one motion information merging candidate is considered to be 0, and the motion vector for the predetermined direction of the first additional motion information merging candidate can be derived.

[0015] In the image encoding / decoding method according to the present disclosure, the template cost of each of the reference pictures can be derived based on a current template composed of a restored area around the current block and a reference template at a position spaced apart from the current template within the reference picture by the motion vector of the first additional motion information merging candidate.

[0016] In the image encoding / decoding method according to the present disclosure, the number of the plurality of motion information merging candidates may be greater than 2.

[0017] In the image encoding / decoding method according to the present disclosure, the motion information merging list further includes a second additional motion information merging candidate, and the number of motion information merging candidates used to derive the first additional motion information merging candidate and the number of motion information merging candidates used to derive the second additional motion information merging candidate may be different.

[0018] In the image encoding / decoding method according to the present disclosure, the first additional motion information merging candidate may not be available for use in deriving the second additional motion information merging candidate.

[0019] In the image encoding / decoding method according to the present disclosure, the motion information merging list may be divided into a first motion information merging list and a second motion information merging list.

[0020] In the image encoding / decoding method according to the present disclosure, the motion information is derived from one selected combination among a plurality of combinations, and the combination may consist of a first motion information merging candidate belonging to the first motion information merging list and a second motion information merging candidate belonging to the second motion information merging list.

[0021] In the image encoding / decoding method according to the present disclosure, the prediction block may be obtained by weighting the sum of a first prediction block derived based on the first motion information of the first motion information merging candidate and a second prediction block derived based on the second motion information of the second motion information merging candidate.

[0022] In the image encoding / decoding method according to the present disclosure, each of the first motion information merging list and the second motion information merging list can be rearranged by a template cost.

[0023] In the image encoding / decoding method according to the present disclosure, the first motion information merging list may be composed of motion information merging candidates in which the index within the motion information merging list is even, and the second motion information merging list may be composed of motion information merging candidates in which the index within the motion information merging list is odd.

[0024] According to the present disclosure, a computer-readable recording medium for storing a bitstream generated by an image encoding method may be provided.

[0025] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0026] According to the present disclosure, encoding / decoding efficiency can be improved by providing a method for configuring an improved motion information merge list.

[0027] According to the present disclosure, prediction accuracy can be improved by inducing additional motion information merging candidates using a plurality of motion information merging candidates already inserted in a motion information merging list.

[0028] According to the present disclosure, prediction accuracy can be improved by selecting one of a plurality of combinations, each comprising a plurality of motion information candidates, and performing inter-prediction based on the selected combination.

[0029] According to the present disclosure, the prediction accuracy can be improved by providing a method for deriving motion information of the current block based on a plurality of motion information merge lists.

[0030] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0031] FIG. 1 is a block diagram showing an image encoding device according to one embodiment of the present disclosure.

[0032] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.

[0033] Figure 3 is a diagram illustrating the process of performing inter-prediction in the encoder and decoder.

[0034] Figure 4 shows an example where motion estimation is performed.

[0035] Figures 5 and 6 show examples of how a predicted block of the current block is generated based on motion information generated through motion estimation.

[0036] Figure 7 shows the location referenced to derive the motion vector prediction value.

[0037] Figure 8 is a diagram illustrating a template-based motion estimation method.

[0038] Figure 9 shows examples of template configurations.

[0039] Figure 10 is a diagram illustrating a motion estimation method based on a two-way matching method.

[0040] Figure 11 is a diagram illustrating a motion estimation method based on a unidirectional matching method.

[0041] Figures 12 and 13 illustrate examples in which prediction blocks are generated according to the precision of the motion vectors.

[0042] Figure 14 shows an example in which motion compensation based on a translational model and a zooming model is performed for the current block.

[0043] Figure 15 shows an example in which motion compensation based on a translational model and a rotational model is performed for the current block.

[0044] Figures 16 and 17 show an example of generating a prediction block for the current block using control point motion vectors.

[0045] Figure 18 shows an example of generating a prediction block for the current block using three control point motion vectors.

[0046] Figure 19 shows an example in which motion vectors are derived in sub-block units.

[0047] Figures 20 and 21 show an example where motion vectors are induced in sub-block units within the current block when SbTMVP is applied.

[0048] Figures 22 and 23 are diagrams illustrating examples in which a prediction block is derived according to the precision of the motion vector.

[0049] FIGS. 24 and FIGS. 25 are diagrams illustrating the process of encoding and decoding motion vector difference values ​​when the AMVR method is applied, respectively.

[0050] Figure 26 shows an example of generating a new motion information merge candidate using motion information merge candidates that are already inserted into the motion information merge list.

[0051] Figure 27 shows an example in which additional motion information merging candidates are derived based on three motion information merging candidates.

[0052] Figure 28 shows an example in which additional motion information merging candidates are generated up to the maximum number of combinations of motion information merging candidates already inserted in the motion information merging list.

[0053] Figure 29 shows an example in which a candidate for 4th movement information merging is derived.

[0054] Figure 30 illustrates an example of the configuration of a motion information merging list.

[0055] Figure 31 shows an example in which a unique index is assigned to a combination of multiple motion information merging candidates.

[0056] FIG. 32 is a diagram illustrating candidate blocks for inducing motion information merging candidates.

[0057] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.

[0058] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0059] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0060] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0061] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.

[0062] FIG. 1 is a block diagram showing an image encoding device according to one embodiment of the present disclosure.

[0063] Referring to FIG. 1, the image encoding device (100) may include a picture splitting unit (110), a prediction unit (120, 125), a conversion unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse conversion unit (145), a filter unit (150), and a memory (155).

[0064] Each component shown in FIG. 1 is depicted independently to represent different characteristic functions of the image encoding device and does not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for convenience of explanation, but at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform functions, and such integrated and separated embodiments of each component are included within the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure.

[0065] Additionally, some components may not be essential components performing an essential function in the present disclosure, but may be optional components merely for enhancing performance. The present disclosure may be implemented by including only the components essential to embody the essence of the present disclosure, excluding components used merely for enhancing performance, and a structure including only the essential components, excluding optional components used merely for enhancing performance, is also included within the scope of the rights of the present disclosure.

[0066] The picture segmentation unit (110) can divide an input picture into at least one processing unit. At this time, the processing unit may be a Prediction Unit (PU), a Transform Unit (TU), or a Coding Unit (CU). The picture segmentation unit (110) can divide a picture into a combination of multiple coding units, prediction units, and transformation units, and can encode the picture by selecting one combination of coding units, prediction units, and transformation units based on a predetermined criterion (e.g., a cost function).

[0067] For example, a single picture can be divided into multiple coding units. To divide coding units within a picture, recursive tree structures such as a Quad Tree, Ternary Tree, or Binary Tree can be used. A coding unit divided into other coding units, with a single image or the largest coding unit as the root, can have as many child nodes as the number of divided coding units. A coding unit that is no longer divided according to certain limits becomes a leaf node. For example, assuming Quad Tree division is applied to a single coding unit, a single coding unit can be divided into up to four different coding units.

[0068] In the embodiments of the present disclosure below, the encoding unit may be used to mean a unit that performs encoding, or a unit that performs decoding.

[0069] A prediction unit may be divided into at least one shape, such as a square or rectangle, of the same size within a single encoding unit, or one of the prediction units divided within a single encoding unit may be divided such that any one prediction unit has a different shape and / or size from another prediction unit.

[0070] When performing intra-frame prediction, the transformation unit and the prediction unit may be set to be the same. In this case, the encoding unit may be divided into multiple transformation units, and intra-frame prediction may be performed for each transformation unit. The encoding unit may be divided in a horizontal or vertical direction. The number of transformation units generated by dividing the encoding unit may be two or four, depending on the size of the encoding unit. Alternatively, if the size of the transformation unit is small, multiple transformation units may be set as a single prediction unit.

[0071] The prediction unit (120, 125) may include an inter-frame prediction unit (120) that performs inter-frame prediction and an intra-frame prediction unit (125) that performs intra-frame prediction. It may determine whether to use inter-frame prediction or perform intra-frame prediction for a encoding unit, and determine specific information (e.g., reference sample line, intra-frame prediction mode, motion vector, reference picture, etc.) according to each prediction method. At this time, the processing unit in which the prediction is performed and the processing unit in which the prediction method and specific details are determined may be different. For example, the prediction method and prediction mode, etc., may be determined by the encoding unit, and the prediction may be performed by the prediction unit or the conversion unit. The residual value (residual block) between the generated prediction block and the original block may be input to the conversion unit (130). In addition, the prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value in the entropy encoding unit (165) and transmitted to the decoding device. When using a specific encoding mode, it is also possible to encode the original block as is and transmit it to the decoding unit without generating a prediction block through the prediction unit (120, 125).

[0072] The inter-frame prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous picture or the subsequent picture of the current picture, and in some cases, may predict a prediction unit based on information of a partially encoded area within the current picture. The inter-frame prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.

[0073] In the reference picture interpolation unit, reference picture information is received from memory (155), and pixel information of integer pixels or less can be generated from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information of integer pixels or less in 1 / 4 pixel units. In the case of chrominance signals, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information of integer pixels or less in 1 / 8 pixel units.

[0074] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm), can be used to calculate motion vectors. Based on the interpolated pixels, the motion vector can have motion vector values ​​in units of 1 / 2 or 1 / 4 pixels. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various motion prediction methods, such as the Skip method, Merge method, AMVP (Advanced Motion Vector Prediction) method, and Intra Block Copy method, can be used.

[0075] The in-screen prediction unit (125) can generate a prediction block based on reference pixel information, which is pixel information within the current picture. Reference pixel information can be derived from one selected from a plurality of reference pixel lines. The Nth reference pixel line among the plurality of reference pixel lines may include left pixels with an x-axis difference of N with the top-left pixel in the current block and top pixels with a y-axis difference of N with said top-left pixel. The number of reference pixel lines that the current block can select may be 1, 2, 3, or 4.

[0076] If a neighboring block of the current prediction unit is a block that has undergone inter-frame prediction, and the reference pixel is a pixel that has undergone inter-frame prediction, the reference pixel included in the block that has undergone inter-frame prediction can be replaced with the reference pixel information of a neighboring block that has undergone intra-frame prediction. That is, if the reference pixel is not available, the information of the unavailable reference pixel can be replaced with the information of at least one of the available reference pixels.

[0077] In intra-frame prediction, the prediction mode may include a directional prediction mode that uses reference pixel information according to the prediction direction, and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information may be different, and the intra-frame prediction mode information used to predict luminance information or the predicted luminance signal information may be utilized to predict chrominance information.

[0078] When performing intra-frame prediction, if the size of the prediction unit and the size of the transformation unit are the same, intra-frame prediction for the prediction unit can be performed based on the pixels to the left of the prediction unit, the pixels at the top left, and the pixels at the top.

[0079] The in-frame prediction method can generate a prediction block after applying a smoothing filter to a reference pixel according to the prediction mode. Depending on the selected reference pixel line, it may be determined whether to apply the smoothing filter.

[0080] To perform an intra-frame prediction method, the intra-frame prediction mode of the current prediction unit can be predicted from the intra-frame prediction mode of the prediction unit existing in the vicinity of the current prediction unit. When predicting the prediction mode of the current prediction unit using the mode information predicted from the surrounding prediction unit, if the intra-frame prediction mode of the current prediction unit and the surrounding prediction unit are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction unit are the same can be transmitted using predetermined flag information; if the prediction modes of the current prediction unit and the surrounding prediction unit are different, entropy coding can be performed to encode the prediction mode information of the current block.

[0081] Additionally, a residual block can be generated that includes residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit. The generated residual block can be input to the conversion unit (130).

[0082] In the transformation unit (130), the residual block containing residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on at least one of the size of the transformation unit, the shape of the transformation unit, the prediction mode of the prediction unit, or the in-frame prediction mode information of the prediction unit. Meanwhile, the transformation can be performed by separating the horizontal direction and the vertical direction.

[0083] After performing transformations for the horizontal and vertical directions, a second transformation can be performed. The second transformation may be in a form where the horizontal and vertical directions are not separated. Final transformation coefficients can be generated by performing a second transformation on the transformation coefficients obtained by the first transformation. Meanwhile, the number of final transformation coefficients output by the second transformation may be smaller than the number of transformation coefficients input for the second transformation. Specifically, the second transformation can be performed using a reduced transformation matrix with different numbers of columns and rows.

[0084] The quantization unit (135) can quantize the values ​​converted into the frequency domain in the conversion unit (130). The quantization coefficient may vary depending on the block or the importance of the image. The values ​​produced by the quantization unit (135) may be provided to the inverse quantization unit (140) and the reordering unit (160).

[0085] The reordering unit (160) can perform reordering of coefficient values ​​for quantized residual values.

[0086] The reordering unit (160) can convert two-dimensional block-shaped coefficients into one-dimensional vector forms through a coefficient scanning method. For example, the reordering unit (160) can convert the coefficients from DC to high-frequency ranges into one-dimensional vector forms by scanning using a Zig-Zag Scan method. Depending on the size of the conversion unit and the in-frame prediction mode, instead of Zig-Zag Scan, a vertical scan that scans two-dimensional block-shaped coefficients in the column direction, a horizontal scan that scans two-dimensional block-shaped coefficients in the row direction, or a diagonal scan that scans two-dimensional block-shaped coefficients in the diagonal direction may be used. That is, depending on the size of the conversion unit and the in-frame prediction mode, it can be determined whether to use a Zig-Zag Scan, a vertical scan, a horizontal scan, or a diagonal scan.

[0087] The entropy encoding unit (165) can perform entropy encoding based on the values ​​calculated by the reordering unit (160). Entropy encoding can use various encoding methods, such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).

[0088] The entropy encoding unit (165) can encode various information from the reordering unit (160) and the prediction unit (120, 125), such as residual value coefficient information of the encoding unit, block type information, prediction mode information, division unit information, prediction unit information and transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information.

[0089] The entropy encoding unit (165) can entropy-encode the coefficient value of the encoding unit input from the rearrangement unit (160).

[0090] In the inverse quantization unit (140) and inverse transformation unit (145), the values ​​quantized in the quantization unit (135) are inversely quantized, and the values ​​transformed in the transformation unit (130) are inversely transformed. The residual value generated in the inverse quantization unit (140) and inverse transformation unit (145) can be combined with the predicted unit predicted through the motion estimation unit, motion compensation unit, and in-frame prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.

[0091] The filter section (150) may include at least one of a deblocking filter, an offset correction section, and an ALF (Adaptive Loop Filter).

[0092] The deblocking filter can remove block distortion caused by boundaries between blocks in the restored picture. To determine whether to perform deblocking, the decision to apply the deblocking filter to the current block can be made based on the pixels contained in a certain number of columns or rows within the block. When applying the deblocking filter to a block, a Strong Filter or a Weak Filter can be applied depending on the required deblocking filtering strength. Additionally, when applying the deblocking filter, horizontal and vertical filtering can be processed in parallel.

[0093] The offset correction unit can correct the offset from the original image on a pixel-by-pixel basis for the image that has undergone deblocking. To perform offset correction for a specific picture, a method can be used in which pixels included in the image are divided into a certain number of regions, the region to be offset is determined, and the offset is applied to that region, or a method can be used in which the offset is applied by considering the edge information of each pixel.

[0094] Adaptive Loop Filtering (ALF) can be performed based on a comparison between the filtered restored image and the original image. After dividing the pixels included in the image into predetermined groups, a single filter to be applied to each group can be determined, allowing for differential filtering for each group. Information regarding whether to apply ALF can be transmitted per coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied may vary depending on each block. Additionally, an ALF filter of the same form (fixed form) may be applied regardless of the characteristics of the block to be applied.

[0095] The memory (155) can store a restoration block or picture calculated through the filter unit (150), and the stored restoration block or picture can be provided to the prediction unit (120, 125) when performing inter-frame prediction.

[0096] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.

[0097] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a reordering unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).

[0098] When a video bitstream is input to a video encoding device, the input bitstream can be decoded by the reverse procedure of the video encoding device.

[0099] The entropy decoding unit (210) can perform entropy decoding in the opposite procedure to that which the entropy encoding unit of the image encoding device performed entropy encoding. For example, various methods such as Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) may be applied in correspondence with the method performed by the image encoding device.

[0100] The entropy decoding unit (210) can decode information related to intra-frame prediction and inter-frame prediction performed by the encoding device.

[0101] The reordering unit (215) can perform reordering based on the method of reordering the entropy-decoded bitstream in the encoding unit in the entropy decoding unit (210). It can reorder by restoring the coefficients expressed in the form of a one-dimensional vector back into coefficients in the form of a two-dimensional block. The reordering unit (215) can perform reordering by receiving information related to the coefficient scanning performed in the encoding unit and scanning in reverse based on the scanning order performed in the encoding unit.

[0102] The inverse quantization unit (220) can perform inverse quantization based on the coefficient values ​​of the rearranged block and the quantization parameters provided by the encoding device.

[0103] The inverse transform unit (225) can perform an inverse transform of the transform performed by the transform unit on the quantization result performed by the image encoding device. That is, it can perform at least one of an inverse transform of the second transform (second inverse transform) or an inverse transform for DCT, DST, and KLT (i.e., first inverse transform). The inverse transform can be performed based on a transmission unit determined by the image encoding device. The inverse transform unit (225) of the image decoder can determine a transform matrix for the second inverse transform or a transform technique for the first inverse transform (e.g., DCT, DST, KLT) according to a plurality of information such as a prediction method, the size and shape of the current block, a prediction mode, and an intra-frame prediction direction. Alternatively, information for determining the transform matrix or transform technique may be explicitly encoded and signaled.

[0104] The prediction unit (230, 235) can generate a prediction block based on the prediction block generation information provided by the entropy decoding unit (210) and the previously decoded block or picture information provided by the memory (245).

[0105] As described above, when performing intra-frame prediction identical to the operation in the video encoding device, if the size of the prediction unit and the size of the transform unit are the same, intra-frame prediction for the prediction unit is performed based on the pixels to the left of the prediction unit, the pixels to the top left, and the pixels to the top; however, if the size of the prediction unit and the size of the transform unit are different when performing intra-frame prediction, intra-frame prediction can be performed using reference pixels based on the transform unit. Additionally, intra-frame prediction using NxN partitioning only for the minimum encoding unit may also be used.

[0106] The prediction unit (230, 235) may include a prediction unit determination unit, an inter-frame prediction unit, and an intra-frame prediction unit. The prediction unit determination unit receives various information, such as prediction unit information input from the entropy decoding unit (210), prediction mode information of the intra-frame prediction method, and motion prediction related information of the inter-frame prediction method, distinguishes the prediction unit in the current encoding unit, and determines whether the prediction unit performs inter-frame prediction or intra-frame prediction. The inter-frame prediction unit (230) may perform inter-frame prediction for the current prediction unit based on information included in at least one picture among the previous picture or subsequent picture of the current picture containing the current prediction unit, using information necessary for inter-frame prediction of the current prediction unit provided by the video encoding device. Alternatively, it may perform inter-frame prediction based on information of a partially restored area within the current picture containing the current prediction unit.

[0107] To perform inter-frame prediction, based on the encoding unit, it is possible to determine whether the motion prediction method of the prediction unit included in the corresponding encoding unit is Skip Mode, Merge Mode, AMVP Mode, or Intra-frame Block Copy Mode.

[0108] The intra-frame prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra-frame prediction, it can perform intra-frame prediction based on the intra-frame prediction mode information of the prediction unit provided by the video encoding device. The intra-frame prediction unit (235) may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part that performs filtering on the reference pixel of the current block, and can determine whether to apply the filter based on the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixel of the current block using the prediction mode of the prediction unit and the AIS filter information provided by the video encoding device. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0109] The reference pixel interpolation unit can generate a reference pixel of an integer value or less by interpolating the reference pixel when the prediction mode of the prediction unit is a prediction unit that performs intra-frame prediction based on the pixel value interpolated from the reference pixel. If the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is DC mode.

[0110] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.

[0111] Information regarding whether a deblocking filter has been applied to the corresponding block or picture can be received from the video encoding device, and if a deblocking filter has been applied, information regarding whether a strong filter or a weak filter has been applied. The deblocking filter of the video decoder receives information related to the deblocking filter provided by the video encoding device, and the video decoder can perform deblocking filtering on the corresponding block.

[0112] The offset correction unit can perform offset correction on the restored image based on the type of offset correction and offset value information applied to the image during encoding.

[0113] ALF can be applied to the encoding unit based on information on whether to apply ALF, ALF coefficient information, etc., provided by the encoding device. This ALF information can be provided included in a specific parameter set.

[0114] The memory (245) can store the restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to the output unit.

[0115] As described above, in the embodiments of the present disclosure below, the term "Coding Unit" is used as "encoding unit" for convenience of explanation, but it may be a unit that performs not only encoding but also decoding.

[0116] Additionally, the current block represents a block to be encoded / decoded, and depending on the encoding / decoding stage, it may represent a coding tree block (or coding tree unit), an encoding block (or encoding unit), a conversion block (or conversion unit), a prediction block (or prediction unit), or a block to which an in-loop filter is applied. In this specification, 'unit' represents a basic unit for performing a specific encoding / decoding process, and 'block' may represent a pixel array of a predetermined size. Unless otherwise distinguished, 'block' and 'unit' may be used with the same meaning. For example, in the embodiments described below, the encoding block (coding block) and the encoding unit (coding unit) may be understood as having the same meaning.

[0117] Furthermore, the picture containing the current block will be referred to as the current picture.

[0118] When encoding the current picture, duplicate data between pictures can be removed through inter-prediction. Inter-prediction can be performed on a block basis. Specifically, a prediction block of the current block can be generated from a reference picture using motion information of the current block. Here, the motion information may include at least one of a motion vector, a reference picture index, and a prediction direction.

[0119] Figure 3 is a diagram illustrating the process of performing inter-prediction in the encoder and decoder.

[0120] As shown in the example illustrated in FIG. 3, motion information for the current block can be obtained to perform inter-prediction (S310). Here, the motion information may include at least one of a motion vector, a reference picture index, or a weight applied to the prediction block. For the current block, motion information for at least one of the L0 direction or the L1 direction may be obtained.

[0121] In the encoder, motion information of the current block can be derived through motion estimation, and the derived motion information can be encoded and signaled to the decoder. Meanwhile, the encoding / decoding of motion information may be based on a motion information merging mode, a motion vector prediction mode, a template-based motion estimation method, or a two-way matching method, which will be described later.

[0122] In the decoder, movement information of the current block can be derived based on the information transmitted from the encoder.

[0123] Alternatively, motion information of the current block can be derived in the decoder in the same way as in the encoder. This method can be referred to as decoder-side motion estimation.

[0124] When motion information of the current block is derived, a prediction block for the current block can be obtained based on the derived motion information (S320). For example, a reference block spaced apart by a motion vector from the position of the current block in the reference picture can be set as the prediction block of the current block.

[0125] Below, we will explain in more detail the process of receiving inter-predictions.

[0126] The motion information of the current block can be generated through motion estimation.

[0127] Figure 4 shows an example where motion estimation is performed.

[0128] In Figure 4, it was assumed that the Picture Order Count (POC) of the current picture is T, and the POC of the reference picture is (T-1).

[0129] A search range for motion estimation can be set from the same location as the reference point of the current block within the reference picture. Here, the reference point may be the location of the top-left sample of the current block.

[0130] For example, in FIG. 4, a rectangle of sizes (w0+w01) and (h0+h1) centered on a reference point is exemplified as being set as a search range. In the above example, w0, w1, h0, and h1 may have mutually identical values. Alternatively, at least one of w0, w1, h0, and h1 may be set to have a different value. Or, the sizes of w0, w1, h0, and h1 may be determined so as not to exceed the Coding Tree Unit (CTU) boundary, slice boundary, tile boundary, or picture boundary.

[0131] Within the search range, reference blocks of the same size as the current block can be set, and the cost of each reference block relative to the current block can be measured. The cost can be calculated using the similarity between the two blocks.

[0132] For example, the cost can be calculated based on the sum of the absolute differences between the original samples in the current block and the original samples (or restored samples) in the reference block. The smaller the sum of the absolute values, the lower the cost can be.

[0133] Afterward, the cost of each of the reference blocks is compared, and the reference block with the optimal cost can be set as the prediction block of the current block.

[0134] In addition, the distance between the current block and the reference block can be set as a motion vector. Specifically, the x-coordinate difference and the y-coordinate difference between the current block and the reference block can be set as a motion vector.

[0135] Furthermore, the index of the picture containing the reference block identified through motion estimation is set as the reference picture index.

[0136] In addition, the prediction direction can be set based on whether the reference picture belongs to the L0 reference picture list or the L1 reference picture list.

[0137] Additionally, motion estimation can be performed for the L0 direction and the L1 direction, respectively. If prediction is performed for both the L0 direction and the L1 direction, motion information for the L0 direction and motion information for the L1 direction can be generated, respectively.

[0138] Figures 5 and 6 show examples of how a predicted block of the current block is generated based on motion information generated through motion estimation.

[0139] Figure 5 shows an example of generating a prediction block with unidirectional (i.e., L0 direction) prediction, and Figure 6 shows an example of generating a prediction block with bidirectional (i.e., L0 and L1 directions) prediction.

[0140] In the case of unidirectional prediction, a prediction block of the current block is generated using a single motion information. For example, the motion information may include an L0 motion vector, an L0 reference picture index, and prediction direction information covering the L0 direction.

[0141] In the case of bidirectional prediction, a prediction block is generated using two sets of motion information. For example, a reference block for the L0 direction, specified based on motion information for the L0 direction (L0 motion information), can be set as the L0 prediction block, and a reference block for the L1 direction, specified based on motion information for the L1 direction (L1 motion information), can be generated as the L1 prediction block. Subsequently, the prediction block of the current block can be generated by performing a weighted sum of the L0 prediction block and the L1 prediction block.

[0142] In the examples illustrated in FIGS. 4 to 6, the L0 reference picture is shown as existing in the direction before the current picture (i.e., having a smaller POC value than the current picture), and the L1 reference picture is shown as existing in the direction after the current picture (i.e., having a larger POC value than the current picture).

[0143] However, unlike the illustrated example, the L0 reference picture may exist in the direction after the current picture, or the L1 reference picture may exist in the direction before the current picture. For example, both the L0 reference picture and the L1 reference picture may exist in the direction before the current picture, or both may exist in the direction after the current picture. Alternatively, bidirectional prediction may be performed using the L0 reference picture existing in the direction after the current picture and the L1 reference picture existing in the direction before the current picture.

[0144] The motion information of the block for which inter-prediction has been performed can be stored in memory. In this case, the motion information can be stored on a sample basis. Specifically, the motion information of the block to which a specific sample belongs can be stored as the motion information of that specific sample. The stored motion information can be used to derive the motion information of neighboring blocks to be encoded / decoded in the future.

[0145] In the encoder, information encoding residual samples corresponding to the difference value between the sample of the current block (i.e., the original sample) and the prediction sample, and motion information necessary to generate the prediction block, can be signaled to the decoder. In the decoder, information regarding the signaled difference value is decoded to derive a difference sample, and a prediction sample within the prediction block generated using the motion information is added to the difference sample to generate a reconstructed sample.

[0146] At this time, one of a plurality of inter-prediction modes may be selected to effectively compress motion information signaled to the decoder. Here, the plurality of inter-prediction modes may include a motion information merging mode and a motion vector prediction mode.

[0147] The motion vector prediction mode is a mode that signals by encoding the difference between the motion vector and the motion vector prediction value. Here, the motion vector prediction value can be derived based on motion information of surrounding blocks or surrounding samples adjacent to the current block.

[0148] Figure 7 shows the location referenced to derive the motion vector prediction value.

[0149] For the sake of convenience of explanation, the current block is assumed to have a size of 4x4.

[0150] In the illustrated example, 'LB' represents a sample contained in the leftmost column and bottom row within the current block. 'RT' represents a sample contained in the rightmost column and top row within the current block. A0 through A4 represent samples adjacent to the left of the current block, and B0 through B5 represent samples adjacent to the top of the current block. For example, A1 represents a sample adjacent to the left of LB, and B1 represents a sample adjacent to the top of RT.

[0151] Col indicates the location of a sample adjacent to the bottom-right of the current block within the co-located picture. The co-located picture is a picture distinct from the current picture, and information to identify the co-located picture (e.g., co-located picture index) can be explicitly encoded and signaled in the bitstream. Alternatively, a reference picture having a predefined reference picture index can be set as the co-located picture.

[0152] The motion vector prediction value of the current block can be derived from at least one motion vector prediction candidate included in the Motion Vector Prediction List.

[0153] The number of motion vector prediction candidates that can be inserted into the motion vector prediction list (i.e., the size of the list) may be predefined in the encoder and decoder. For example, the maximum number of motion vector prediction candidates may be 2.

[0154] A motion vector stored at the location of a neighbor sample adjacent to the current block, or a scaled motion vector derived by scaling the said motion vector, can be inserted into the motion vector prediction list as a motion vector prediction candidate. At this time, the motion vector prediction candidate can be derived by scanning the neighbor samples adjacent to the current block according to a predefined order.

[0155] For example, it is possible to check whether a motion vector is stored at each location in the order from A0 to A4. Then, according to the above scan order, the first available motion vector found can be inserted into the motion vector prediction list as a motion vector prediction candidate.

[0156] As another example, checking whether a motion vector is stored at each location in the order from A0 to A4 allows the motion vector at the location with the same reference picture as the current block, found first, to be inserted into the motion vector prediction list as a motion vector prediction candidate. If no neighbor sample with the same reference picture as the current block exists, a motion vector prediction candidate can be derived based on the first available vector found. Specifically, the first available motion vector found can be scaled, and the scaled motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. In this case, scaling can be performed based on the difference in output order between the current picture and the reference picture (i.e., POC difference) and the difference in output order between the current picture and the neighbor sample's reference picture (i.e., POC difference).

[0157] Furthermore, it is possible to check whether a motion vector is stored at each location in the order from B0 to B5. Then, according to the above scan order, the first available motion vector found can be inserted into the motion vector prediction list as a motion vector prediction candidate.

[0158] As another example, checking whether a motion vector is stored at each location in the order from B0 to B5 allows the motion vector at the location with the same reference picture as the current block, found first, to be inserted into the motion vector prediction list as a motion vector prediction candidate. If no neighbor sample with the same reference picture as the current block exists, a motion vector prediction candidate can be derived based on the first available vector found. Specifically, the first available motion vector found can be scaled, and the scaled motion vector can be inserted into the motion vector prediction list as a motion vector prediction candidate. In this case, scaling can be performed based on the difference in output order between the current picture and the reference picture (i.e., POC difference) and the difference in output order between the current picture and the neighbor sample's reference picture (i.e., POC difference).

[0159] As in the example described above, motion vector prediction candidates can be derived from samples adjacent to the left of the current block, and motion vector prediction candidates can be derived from samples adjacent to the top of the current block.

[0160] In this case, a motion vector prediction candidate derived from the left sample may be inserted into the motion vector prediction list before a motion vector prediction candidate derived from the top sample. In this case, the index assigned to the motion vector prediction candidate derived from the left sample may have a smaller value than that of the motion vector prediction candidate derived from the top sample.

[0161] Conversely, motion vector prediction candidates derived from the top sample may be inserted into the motion vector prediction list before motion vector prediction candidates derived from the left sample.

[0162] Among the motion vector prediction candidates included in the above motion vector prediction list, the motion vector prediction candidate with the highest encoding efficiency can be set as the motion vector prediction value (Motion Vector Predictor, MVP) of the current block. Additionally, index information pointing to the motion vector prediction candidate set as the motion vector prediction value of the current block among multiple motion vector prediction candidates can be encoded and signaled to the decoder. If the number of motion vector prediction candidates is two, the index information may be a 1-bit flag (e.g., an MVP flag). Furthermore, the motion vector difference value (Motion Vector Difference, MVD), which is the difference between the motion vector of the current block and the motion vector prediction value, can be encoded and signaled to the decoder.

[0163] The decoder can construct a motion vector prediction list in the same way as the encoder. Additionally, it can decode index information from the bitstream and select one of multiple motion vector prediction candidates based on the decoded index information. The selected motion vector prediction candidate can be set as the motion vector prediction value of the current block.

[0164] In addition, the motion vector difference value can be decoded from the bitstream. Subsequently, the motion vector prediction value and the motion vector difference value are combined to derive the motion vector of the current block.

[0165] When bidirectional prediction is applied to the current block, motion vector prediction lists can be generated for both the L0 and L1 directions. That is, the motion vector prediction lists can consist of motion vectors of the same direction. Accordingly, the motion vector of the current block and the motion vector prediction candidates included in the motion vector prediction lists have the same direction.

[0166] When the motion vector prediction mode is selected, the reference picture index and prediction direction information can be explicitly encoded and signaled to the decoder. For example, if multiple reference pictures exist on a reference picture list and motion estimation is performed for each of the multiple reference pictures, a reference picture index for identifying the reference picture from which the motion information of the current block was derived among the multiple reference pictures can be explicitly encoded and signaled to the decoder.

[0167] In this case, if the reference picture list contains only one reference picture, the encoding / decoding of the reference picture index may be omitted.

[0168] The prediction direction information may be an index indicating one of L0 unidirectional prediction, L1 unidirectional prediction, or bidirectional prediction. Alternatively, an L0 flag indicating whether a prediction for the L0 direction is performed and an L1 flag indicating whether a prediction for the L1 direction is performed may be encoded and signaled, respectively.

[0169] The motion information merging mode is a mode that sets the motion information of the current block to be identical to the motion information of neighboring blocks. In the motion information merging mode, motion information can be encoded or decoded using a motion information merging list.

[0170] Motion information merging candidates can be derived based on motion information from neighboring blocks or neighbor samples adjacent to the current block. For example, after defining reference locations around the current block, it is possible to check whether motion information exists at the defined reference locations. If motion information exists at the defined reference locations, the motion information at those locations can be inserted into the motion information merging list as a motion information merging candidate.

[0171] In the example of FIG. 7, the previously defined reference positions may include at least one of A0, A1, B0, B1, B5, and Col. Furthermore, motion information merging candidates can be derived in the order of A1, B1, B0, A0, B5, and Col.

[0172] The motion information of the motion information merge candidate with the optimal cost among the motion information merge candidates included in the motion information merge list can be set as the motion information of the current block. Furthermore, index information (e.g., merge index) pointing to the selected motion information merge candidate among multiple motion information merge candidates can be encoded and transmitted to a decoder.

[0173] In the decoder, a motion information merge list can be configured in the same way as in the encoder. Then, motion information merge candidates can be selected based on the merge index decoded from the bitstream. The motion information of the selected motion information merge candidate can be set as the motion information of the current block.

[0174] Unlike the motion vector prediction list, the motion information merging list consists of a single list regardless of the prediction direction. That is, the motion information merging candidates included in the motion information merging list may have only L0 motion information or L1 motion information, or they may have bidirectional motion information (i.e., L0 motion information and L1 motion information).

[0175] Movement information of the current block can also be derived using a restoration sample area around the current block. Here, the restoration sample area used to derive the movement information of the current block may be referred to as a template.

[0176] Figure 8 is a diagram illustrating a template-based motion estimation method.

[0177] In FIG. 4, it was explained that the predicted block of the current block is determined based on the cost between the current block and the reference block within the search range. According to the present embodiment, unlike FIG. 4, motion estimation for the current block can be performed based on the cost between a template adjacent to the current block (hereinafter referred to as the current template) and a reference template having the same size and shape as the current template.

[0178] For example, the cost can be calculated based on the sum of the absolute differences between the restored samples in the current template and the restored samples in the reference block. The smaller the sum of the absolute values, the lower the cost can be.

[0179] When a reference template with the optimal cost and the current template within the search range is determined, a reference block adjacent to the reference template can be set as the predicted block of the current block.

[0180] Additionally, movement information of the current block can be set based on the distance between the current block and the reference block, the index of the picture to which the reference block belongs, and whether the reference picture is included in the L0 or L1 reference picture list.

[0181] Since the template is defined by the previously restored area surrounding the current block, the decoder can perform motion estimation itself in the same manner as the encoder. Accordingly, when deriving motion information using a template, there is no need to encode and signal the motion information, except for information indicating whether the template is being used.

[0182] The current template may include at least one of an area adjacent to the top of the current block or an area adjacent to the left. In this case, the area adjacent to the top may include at least one row, and the area adjacent to the left may include at least one column.

[0183] Figure 9 shows examples of template configurations.

[0184] The current template can be configured following one of the examples shown in Fig. 9.

[0185] Alternatively, unlike the example shown in FIG. 9, the template may be configured using only the area adjacent to the left of the current block, or only the area adjacent to the top of the current block.

[0186] The size and / or shape of the current template may be predefined in the encoder and decoder.

[0187] Alternatively, multiple template candidates of different sizes and / or shapes can be defined, and index information specifying one of the multiple template candidates can be encoded and signaled to a decoder.

[0188] Alternatively, one of a plurality of template candidates may be adaptively selected based on at least one of the size, shape, or location of the current block. For example, if the current block touches the top boundary of the CTU, the current template may be configured using only the area adjacent to the left of the current block.

[0189] Motion estimation based on a template can be performed for each of the reference pictures stored in the reference picture list. Alternatively, motion estimation can be performed for only some of the reference pictures. For example, motion estimation can be performed only for the reference picture with a reference picture index of 0, or only for reference pictures with a reference picture index smaller than a threshold value, or for reference pictures with a POC difference with the current picture smaller than a threshold value.

[0190] Alternatively, after explicitly encoding and signaling the reference picture index, motion estimation can be performed only on the reference picture pointed to by the reference picture index.

[0191] Alternatively, motion estimation can be performed on a reference picture of a neighbor block corresponding to the current template. For example, if the template consists of a left neighbor area and a top neighbor area, at least one reference picture can be selected using at least one of the reference picture index of the left neighbor block or the reference picture index of the top neighbor block. Subsequently, motion estimation can be performed on the selected at least one reference picture.

[0192] Information indicating whether template-based motion estimation has been applied can be encoded and signaled to a decoder. The information may be a 1-bit flag. For example, if the flag is true (1), it indicates that template-based motion estimation is applied to the L0 and L1 directions of the current block. On the other hand, if the flag is false (0), it indicates that template-based motion estimation is not applied. In this case, motion information of the current block can be derived based on a motion information merging mode or a motion vector prediction mode.

[0193] Conversely to the above, if it is determined that the motion information merging mode and the motion vector prediction mode are not applied to the current block, then a template-based motion estimation may be applied. For example, if a first flag indicating whether the motion information merging mode is applied and a second flag indicating whether the motion vector prediction mode is applied are both 0, then a template-based motion estimation may be performed.

[0194] For each of the L0 and L1 directions, information indicating whether template-based motion estimation has been applied can be signaled. That is, whether template-based motion estimation is applied to the L0 direction and whether it is applied to the L1 direction can be determined independently of each other. Accordingly, while template-based motion estimation is applied to either the L0 or L1 direction, another mode (e.g., motion information merging mode or motion vector prediction mode) may be applied to the other.

[0195] If template-based motion estimation is applied to both the L0 and L1 directions, the prediction block of the current block can be generated based on the weighted sum operation of the L0 prediction block and the L1 prediction block. Alternatively, even if template-based motion estimation is applied to one of the L0 and L1 directions, but another mode is applied to the other, the prediction block of the current block can be generated based on the weighted sum operation of the L0 prediction block and the L1 prediction block.

[0196] Alternatively, a template-based motion estimation method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether to apply the template-based motion estimation method may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate points to the template-based motion estimation method.

[0197] Based on the two-way matching method, movement information of the current block can also be generated.

[0198] Figure 10 is a diagram illustrating a motion estimation method based on a two-way matching method.

[0199] The two-way matching method can be performed only when the temporal order of the current picture (i.e., POC) exists between the temporal order of the L0 reference picture and the temporal order of the L1 reference picture.

[0200] When a two-way matching method is applied, a search range can be set for each of the L0 reference picture and the L1 reference picture. In this case, an L0 reference picture index for identifying the L0 reference picture and an L1 reference picture index for identifying the L1 reference picture can be encoded and signaled, respectively.

[0201] As another example, only the L0 reference picture index is encoded and signaled, and an L1 reference picture can be selected based on the distance between the current picture and the L0 reference picture (hereinafter referred to as the L0 POC difference). For example, among the L1 reference pictures included in the L1 reference picture list, an L1 reference picture can be selected in which the absolute value of the distance from the current picture (hereinafter referred to as the L1 POC difference) is equal to the absolute value of the distance between the current picture and the L0 reference picture. If there is no L1 reference picture having an L1 POC difference identical to the L0 POC difference, the L1 reference picture among the L1 reference pictures in which the L1 POC difference is most similar to the L0 POC difference can be selected.

[0202] At this time, among the L1 reference pictures, only L1 reference pictures that have a different temporal direction from the L0 reference picture can be used for two-way matching. For example, if the POC of the L0 reference picture is smaller than that of the current picture, one of the L1 reference pictures with a POC larger than that of the current picture can be selected.

[0203] Conversely to the above, only the L1 reference picture index is encoded and signaled, and the L0 reference picture is selected based on the distance between the current picture and the L1 reference picture.

[0204] Alternatively, a two-way matching method may be performed using the L0 reference picture closest to the current picture among the L0 reference pictures and the L1 reference picture closest to the current picture among the L1 reference pictures.

[0205] Alternatively, a two-way matching method may be performed using an L0 reference picture (e.g., index 0) assigned to a previously defined index in the L0 reference picture list and an L1 reference picture (e.g., index 0) assigned to a previously defined index in the L1 reference picture list.

[0206] Alternatively, LX (X is 0 or 1) reference picture may be selected based on an explicitly signaled reference picture index, and L|X-1| reference picture may be selected as the reference picture closest to the current picture among L|X-1| reference pictures, or as a reference picture having a predefined index within the L|X-1| reference picture list.

[0207] As another example, L0 and / or L1 reference pictures can be selected based on movement information of neighbor blocks of the current block. For example, L0 and / or L1 reference pictures to be used for bidirectional matching can be selected using the reference picture index of the left or top neighbor block of the current block.

[0208] The search range can be set within a predetermined range from the collocated blocks within the reference picture.

[0209] As another example, the search range can be set based on initial movement information. The initial movement information can be derived from the neighbor blocks of the current block. For example, the movement information of the current block's left neighbor block or top neighbor block can be set as the current block's initial movement information.

[0210] When the two-way matching method is applied, the L0 motion vector and the L1 motion vector are set in opposite directions. This indicates that the sign of the L0 motion vector and the L1 motion vector have opposite signs. Additionally, the magnitude of the LX motion vector can be proportional to the distance between the current picture and the LX reference picture (i.e., the POC difference).

[0211] Subsequently, motion estimation can be performed using the cost between a reference block (hereinafter referred to as the L0 reference block) within the search range of the L0 reference picture and a reference block (hereinafter referred to as the L1 reference block) within the search range of the L1 reference picture.

[0212] If an L0 reference block is selected with a vector (x, y) with respect to the current block, an L1 reference block can be selected at a location spaced (-Dx, -Dy) away from the current block. Here, D can be determined by the ratio of the distance between the current picture and the L0 reference picture to the distance between the L1 reference picture and the current picture.

[0213] For example, in the example illustrated in FIG. 10, the absolute value of the distance between the current picture (T) and the L0 reference picture (T-1) and the absolute value of the distance between the current picture (T) and the L1 reference picture (T+1) are mutually identical. Accordingly, in the illustrated example, the L0 motion vector (x0, y0) and the L1 motion vector (x1, y1) have the same magnitude but opposite distances. If the L1 reference picture with POC (T+2) is used, the L1 motion vector (x1, y1) will be set to (-2*x0, -2*y0).

[0214] When the L0 reference block and L1 reference block having the optimal cost are selected, the L0 reference block and L1 reference block can be set as the L0 prediction block and L1 prediction block of the current block, respectively. Subsequently, the final prediction block of the current block can be generated through a weighted sum operation of the L0 reference block and L1 reference block.

[0215] When a two-way matching method is applied, the decoder can perform motion estimation in the same way as the encoder. Accordingly, information indicating whether a two-way motion matching method is applied is explicitly encoded / decoded, while the encoding / decoding of motion information, such as motion vectors, can be omitted. As previously explained, at least one of the L0 reference picture index or the L1 reference picture index may be explicitly encoded / decoded.

[0216] As another example, information indicating whether a two-way matching method has been applied may be explicitly encoded / decoded; if the two-way matching method has been applied, the L0 motion vector or the L1 motion vector may be explicitly encoded and signaled. If the L0 motion vector is signaled, the L1 motion vector can be derived based on the POC difference between the current picture and the L0 reference picture and the POC difference between the current picture and the L1 reference picture. If the L1 motion vector is signaled, the L0 motion vector can be derived based on the POC difference between the current picture and the L0 reference picture and the POC difference between the current picture and the L1 reference picture. In this case, the encoder may explicitly encode the smaller of the L0 motion vector and the L1 motion vector.

[0217] Information indicating whether a two-way matching method is applied may be a 1-bit flag. For example, if the flag is true (e.g., 1), it may indicate that a two-way matching method is applied to the current block. If the flag is false (e.g., 0), it may indicate that a two-way matching method is not applied to the current block. In this case, a motion information merging mode or a motion vector prediction mode may be applied to the current block.

[0218] Conversely to the above, a two-way matching method may be applied only when it is determined that the motion information merging mode and the motion vector prediction mode are not applied to the current block. For example, if both the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are 0, the two-way matching method may be applied.

[0219] Alternatively, a two-way matching method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether to apply the two-way matching method may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate points to the two-way matching method.

[0220] In the two-way matching method, it was exemplified that the temporal order of the current picture must exist between the temporal order of the L0 reference picture and the temporal order of the L1 reference picture. A one-way matching method, to which the constraints of the above two-way matching method do not apply, may be applied to generate a predicted block of the current block. Specifically, in the one-way matching method, two reference pictures with a temporal order (i.e., POC) smaller than the current block or two reference pictures with a temporal order larger than the current block may be used. In this case, both of the two reference pictures may be derived from the L0 reference picture list or the L1 reference picture list. Alternatively, one of the two reference pictures may be derived from the L0 reference picture list and the other from the L1 reference picture list.

[0221] Figure 11 is a diagram illustrating a motion estimation method based on a unidirectional matching method.

[0222] A unidirectional matching method can be performed based on two reference pictures (i.e., Forward reference pictures) that have a POC smaller than the current picture or two reference pictures (i.e., Backward reference pictures) that have a POC larger than the current picture. In FIG. 11, motion estimation based on a unidirectional matching method is exemplified as being performed based on a first reference picture (T-1) and a second reference picture (T-2) that have a POC smaller than the current picture (T).

[0223] At this time, a first reference picture index for identifying the first reference picture and a second reference picture index for identifying the second reference picture can each be encoded and signaled. At this time, among the two reference pictures used in the unidirectional matching method, the reference picture with a smaller POC difference with the current picture can be set as the first reference picture. Accordingly, when the first reference picture is selected, only reference pictures among the reference pictures included in the reference picture list that have a POC difference with the current picture greater than that of the first reference picture can be set as the second reference picture. The second reference picture index can be set to point to the index of one of the reordered reference pictures after reordering the reference pictures that have the same temporal direction as the first reference picture and have a POC difference with the current picture greater than that of the first reference picture.

[0224] Conversely to the above, the reference picture with the larger POC difference with the current picture among the two reference pictures may be set as the first reference picture. In this case, the index of the second reference picture may be set to point to the index of one of the reordered reference pictures after reordering the reference pictures that have the same temporal direction as the first reference picture and have a smaller POC difference with the current picture than the first reference picture.

[0225] Alternatively, a unidirectional matching method may be performed using a reference picture assigned to a predefined index within the reference picture list and a reference picture having the same temporal direction. For example, a reference picture with an index of 0 within the reference picture list may be set as the first reference picture, and among the reference pictures with the same temporal direction as the first reference picture within the reference picture list, the reference picture with the smallest index may be selected as the second reference picture.

[0226] Both the first reference picture and the second reference picture can be selected from the L0 reference picture list or the L1 reference picture list. In FIG. 11, two L0 reference pictures are shown being used in a unidirectional matching method. Alternatively, the first reference picture may be selected from the L0 reference picture list and the second reference picture may be selected from the L1 reference picture list.

[0227] Information indicating whether the first reference picture and / or the second reference picture belongs to the L0 reference picture list or the L1 reference picture list may be additionally encoded / decoded.

[0228] Alternatively, unidirectional matching can be performed using one of the L0 reference picture list and the L1 reference picture list set as the default. Alternatively, two reference pictures can be selected from the L0 reference picture list and the L1 reference picture list that has a larger number of reference pictures.

[0229] Afterwards, a search range can be set within the first reference picture and the second reference picture.

[0230] The search range can be set within a predetermined range from the collocated blocks within the reference picture.

[0231] As another example, the search range can be set based on initial movement information. The initial movement information can be derived from the neighbor blocks of the current block. For example, the movement information of the current block's left neighbor block or top neighbor block can be set as the current block's initial movement information.

[0232] Subsequently, motion estimation can be performed using the cost between the first reference block within the search range of the first reference picture and the second reference block within the search range of the second reference picture.

[0233] At this time, under the unidirectional matching method, the magnitude of the motion vector should be set to increase in proportion to the distance between the current picture and the reference picture. Specifically, if a first reference block is selected with a vector (x, y) with respect to the current picture, the second reference block should be separated from the current block by (Dx, Dy). Here, D can be determined by the ratio of the distance between the current picture and the first reference picture to the distance between the current picture and the second reference picture.

[0234] For example, in the example of FIG. 11, the distance between the current picture and the first reference picture (i.e., POC difference) is 1, and the distance between the current picture and the second reference picture (i.e., POC difference) is 2. Accordingly, if the first motion vector for the first reference block in the first reference picture is (x0, y0), the second motion vector (x1, y1) for the second reference block in the second reference picture can be set to (2x0, 2y0).

[0235] When a first reference block and a second reference block having optimal costs are selected, the first reference block and the second reference block can be set as the first prediction block and the second prediction block of the current block, respectively. Subsequently, the final prediction block of the current block can be generated through a weighted sum operation of the first prediction block and the second prediction block.

[0236] When a unidirectional matching method is applied, the decoder can perform motion estimation in the same way as the encoder. Accordingly, information indicating whether a unidirectional motion matching method is applied is explicitly encoded / decoded, while the encoding / decoding of motion information, such as motion vectors, can be omitted. As previously explained, at least one of the first reference picture index or the second reference picture index may be explicitly encoded / decoded.

[0237] As another example, information indicating whether a unidirectional matching method has been applied may be explicitly encoded / decoded, and if a unidirectional matching method has been applied, a first motion vector or a second motion vector may be explicitly encoded and signaled. If the first motion vector is signaled, the second motion vector may be derived based on the POC difference between the current picture and the first reference picture and the POC difference between the current picture and the second reference picture. If the second motion vector is signaled, the first motion vector may be derived based on the POC difference between the current picture and the first reference picture and the POC difference between the current picture and the second reference picture. In this case, the encoder may explicitly encode the one with the smaller magnitude between the first motion vector and the second motion vector.

[0238] Information indicating whether a unidirectional matching method is applied may be a 1-bit flag. For example, if the flag is true (e.g., 1), it may indicate that a unidirectional matching method is applied to the current block. If the flag is false (e.g., 0), it may indicate that a unidirectional matching method is not applied to the current block. In this case, a motion information merging mode or a motion vector prediction mode may be applied to the current block.

[0239] Conversely to the above, a unidirectional matching method may be applied only when it is determined that the motion information merging mode and the motion vector prediction mode are not applied to the current block. For example, if both the first flag indicating whether the motion information merging mode is applied and the second flag indicating whether the motion vector prediction mode is applied are 0, a unidirectional matching method may be applied.

[0240] Alternatively, a unidirectional matching method may be inserted as a motion information merging candidate in the motion information merging mode or as a motion vector prediction candidate in the motion vector prediction mode. In this case, whether to apply the unidirectional matching method may be determined based on whether the selected motion information merging candidate or the selected motion vector prediction candidate points to the unidirectional matching method.

[0241] By adjusting the precision of the motion vector, the movement of an object between frames can also be detected. Specifically, the position of each pixel within a picture is specified as an integer. On the other hand, the movement of an object between frames may not be represented by an integer position.

[0242] Considering this, motion vectors can be explored in fractional pixel units by performing interpolation on the reference picture.

[0243] Figures 12 and 13 illustrate examples in which prediction blocks are generated according to the precision of the motion vectors.

[0244] FIG. 12 shows the position of the current block in the current picture, and FIG. 13 illustrates an example in which a predicted block is acquired according to a motion vector.

[0245] Specifically, FIG. 13 (a) shows an example where the motion vector precision is in integer pixel units, and FIG. 13 (b) and (c) show examples where the motion vector precision is in 1 / 2 pixel units and 1 / 4 pixel units, respectively.

[0246] Motion vector precision can also be set in units smaller than those described. For example, motion vector precision can be set in units of 1 / 8 pixel, 1 / 16 pixel, or 1 / 32 pixel.

[0247] When the motion vector of the current block is expressed in integer units, a reference block composed of integer position samples can be set as the prediction block of the current block, as in the example illustrated in FIG. 13 (a).

[0248] On the other hand, when the motion vector of the current block is expressed in fractional units, a reference block composed of fractional position samples can be set as the prediction block of the current block, as in the examples illustrated in FIG. 13 (b) and (c). In this case, the fractional position samples within the reference block can be generated by interpolating integer position samples. The interpolation filter can have a size of 4 taps or 8 taps.

[0249] As another example, to reduce complexity, fractional position samples can be generated through linear interpolation using only integer position samples adjacent to the fractional position.

[0250] Information indicating the motion vector precision of the current block can be encoded and signaled. For example, after assigning different indices to each of multiple motion vector precision candidates, the index of the motion vector precision candidate corresponding to the motion vector precision of the current block can be encoded and signaled.

[0251] At this time, the number and / or types of available motion vector candidates may be determined based on at least one of the size of the current block, the shape of the current block, the reference picture, or the motion compensation model. Here, the motion compensation model may include at least one of a translation model, a zooming model, or a rotation model. A motion compensation model in which at least one of a zooming model or a rotation model is combined with a translation model may be referred to as an affine model.

[0252] An index indicating one of the motion vector candidates available for the current block can be encoded. Depending on the number of motion vector candidates available for the current block, the maximum number of bits required to encode the index can be determined.

[0253] By adjusting the precision of the motion vector, the motion vector can be explored more precisely, and accordingly, the prediction accuracy for the current block can be improved.

[0254] Meanwhile, motion vectors expressed as fractional positions can be scaled up to integers and encoded.

[0255] Compensation for the movement of an object may be performed based on at least one of a translation model to compensate for linear movement of the object (e.g., movement in the horizontal and / or vertical directions), a zooming model to compensate for changes in the size of the object, and a rotation model to compensate for rotational movement of the object. Here, zooming may refer to enlargement or reduction in size.

[0256] Figure 14 shows an example in which motion compensation based on a translational model and a zooming model is performed for the current block.

[0257] For the convenience of explanation, the current block is assumed to have a size of 4x4, as shown in FIG. 12.

[0258] In FIG. 14, the variable α represents the scaling parameter. The size of the reference block can be derived by multiplying the size of the current block by the variable α.

[0259] A scaling parameter α less than 1 indicates that the reference block is smaller than the current block, and a scaling parameter α greater than 1 indicates that the reference block is larger than the current block.

[0260] Figures 14 (a) and (b) show examples where the scaling parameter α is less than 1, and Figure 14 (c) shows an example where the scaling parameter α is greater than 1.

[0261] Based on the motion vector of the current block, the top-left position of the reference block can be determined. Specifically, the top-left position of the reference block can be set to a position offset by the motion vector from the position corresponding to the top-left sample of the current block within the reference picture. Subsequently, a reference block can be set such that its width and height are each α times the width and height of the current block, respectively, according to a scaling parameter. Fractional position samples within the reference block can be generated by interpolating integer position samples.

[0262] The reference block derived by the motion vector and scaling parameter can be set as the prediction block of the current block.

[0263] Meanwhile, information regarding the size adjustment parameter α can be encoded and signaled. Specifically, a different index is assigned to each of the multiple size adjustment parameter candidates, and an index specifying the size adjustment parameter candidate applied to the current block can be encoded and signaled.

[0264] Alternatively, the size parameter of the current block may be derived based on the size parameter of a neighbor block. For example, the size parameter of a neighbor block at a predefined location can be set as the size parameter of the current block.

[0265] Alternatively, when multiple neighbor blocks are searched sequentially, the size adjustment parameter of the first available neighbor block found can be set as the size adjustment parameter of the current block.

[0266] Alternatively, a size control parameter of a neighboring block can be set as a size control parameter candidate. In this case, a list of size control parameter candidates containing multiple size control parameter candidates can be generated by sequentially searching multiple neighboring blocks. One of the multiple size control parameter candidates included in the list of multiple size control parameter candidates can be set as the size control parameter of the current block. In this case, an index indicating a candidate among the multiple size control parameter candidates that is identical to the size control parameter of the current block can be encoded and signaled.

[0267] Meanwhile, the neighbor blocks used to derive the size adjustment parameters of the current block may include at least one of the top neighbor block, left neighbor block, top-left neighbor block, top-right neighbor block, or bottom-left neighbor block.

[0268] Figure 15 shows an example in which motion compensation based on a translational model and a rotational model is performed for the current block.

[0269] For the convenience of explanation, the current block is assumed to have a size of 4x4, as shown in FIG. 12.

[0270] First, as shown in (a) of FIG. 15, the position of a temporary block within a reference picture can be determined based on the motion vector of the current block. Specifically, a block position can be determined by taking a position spaced apart by the motion vector from a position corresponding to the top-left sample of the current block within the reference picture as the top-left sample.

[0271] Afterwards, the temporary block can be rotated as in the example shown in FIG. 15 (b). The block at the rotated position is set as a reference block, and the reference block can be set as a prediction block of the current block.

[0272] Meanwhile, a rotation matrix may be used when rotating a temporary block specified by a motion vector. That is, the predicted sample for the current block can be set to a sample at a position obtained by applying a rotation matrix to the sample position within the temporary block.

[0273] Mathematical Equation 1 represents the rotation matrix.

[0274]

[0275] In the above mathematical formula 1, (pos_x, pos_y) represents the position of a sample within a temporary block. That is, (pos_x, pos_y) can be derived by adding a motion vector to the position of the target sample to be predicted within the current block.

[0276] (pos_x', pos_y') represents the position rotated from the position of the sample within the temporary block, and θ represents the rotation angle.

[0277] The sample value at position (pos_x', pos_y') within the reference picture can be set as the value of the predicted sample for the position of the sample to be predicted. If position (pos_x', pos_y') is a fractional position, the sample at that position can be generated by interpolating integer position samples.

[0278] Meanwhile, information representing the rotation angle θ can be encoded and signaled. For example, after assigning different indices to each of a plurality of rotation angle candidates, the index of the rotation angle candidate corresponding to the rotation angle of the current block can be encoded and signaled.

[0279] Alternatively, the rotation angle of the current block can be derived based on the rotation angle of a neighbor block. For example, the rotation angle of a neighbor block at a predefined position can be set as the rotation angle of the current block.

[0280] Alternatively, when multiple neighbor blocks are searched sequentially, the rotation angle of the first available neighbor block found can be set as the rotation angle of the current block.

[0281] Alternatively, the rotation angle of a neighboring block can be set as a rotation angle candidate. In this case, a rotation angle candidate list containing multiple rotation angle candidates can be generated by sequentially searching multiple neighboring blocks. One of the multiple rotation angle candidates included in the list of multiple rotation angle candidates can be set as the rotation angle of the current block. In this case, an index indicating the candidate among the multiple rotation angle candidates that is identical to the rotation angle of the current block can be encoded and signaled.

[0282] Meanwhile, the neighbor block used to induce the rotation angle of the current block may include at least one of the top neighbor block, left neighbor block, top-left neighbor block, top-right neighbor block, or bottom-left neighbor block.

[0283] Although not explicitly stated, motion compensation for the current block can also be performed by simultaneously applying translational, zooming, and rotational models.

[0284] Meanwhile, the motion vector precision for the current block or the number and / or types of motion vector precision candidates available for the current block may be determined differently depending on the motion compensation model.

[0285] For example, the number and / or types of motion vector precision candidates available for the current block may differ between the case where only a translation model is applied and the case where at least one of a zooming model or a rotation model is applied.

[0286] As a specific example, when a translation model is applied to the current block, candidates of at least 1 / 4 pixel unit may be available for the current block. On the other hand, when at least one of a zooming model or a rotation model is additionally applied along with the translation model to the current block, candidates of at least 1 / 16 pixel unit may be available for the current block.

[0287] Alternatively, if a translation model is applied to the current block, the motion vector precision of the current block may be set to 1 / 4 pixel units. On the other hand, if at least one of a zooming model or a rotation model is additionally applied to the current block along with the translation model, the motion vector precision of the current block may be set to 1 / 16 pixel units.

[0288] Meanwhile, available motion vector precision or available motion vector precision candidates for each motion compensation model may be stored in the encoder and decoder. Alternatively, information representing available motion vector precision or available motion vector precision candidates for each motion compensation model may be encoded and signaled through an upper header.

[0289] Motion compensation for an affine model, to which a zooming model and / or a rotation model are added to a translation model, can be performed using the motion vector of a control point. Here, the control point may correspond to a corner of the current block. For example, to perform motion compensation based on an affine model, at least one of the motion vector of the top-left corner, the motion vector of the top-right corner, or the motion vector of the bottom-left corner may be used.

[0290] Hereinafter, the motion vector of a control point will be referred to as the control point motion vector.

[0291] Figures 16 and 17 show an example of generating a prediction block for the current block using control point motion vectors.

[0292] For the convenience of explanation, the current block is assumed to have a size of 4x4, as shown in FIG. 12.

[0293] In FIG. 16 (a) and (b), a prediction block for the current block is exemplified by the motion vector of the first control point corresponding to the top-left corner of the current block (first control point motion vector, A) and the motion vector of the second control point corresponding to the top-right corner of the current block (second control point motion vector, B).

[0294] Beyond the illustrated examples, it is also possible to derive the predicted block of the current block by additionally utilizing the motion vector of the bottom-left corner or by using the motion vector of the bottom-left corner instead of the top-right corner.

[0295] Figure 18 shows an example of generating a prediction block for the current block using three control point motion vectors.

[0296] In FIG. 18 (a) and (b), a prediction block for the current block is exemplified by the motion vector of the first control point corresponding to the upper-left corner of the current block (first control point motion vector, A), the motion vector of the second control point corresponding to the upper-right corner of the current block (second control point motion vector, B), and the motion vector of the third control point corresponding to the lower-left corner of the current block (third control point motion vector, C).

[0297] As shown in the examples illustrated in FIGS. 16 to 18, translation, zooming, and rotational motion compensation for the current block can be performed using two or three control point motion vectors.

[0298] Information indicating the number of control point motion vectors can be encoded and signaled. The information can be signaled in blocks. For example, the information can indicate whether two control point motion vectors or three control point motion vectors are used in the current block.

[0299] Alternatively, the number of control point motion vectors can be adaptively determined based on at least one of the size or shape of the current block.

[0300] Alternatively, if the control point motion vectors of the current block are derived from neighboring blocks, the number of control point motion vectors for the current block can be set to be equal to the number of control point motion vectors of neighboring blocks.

[0301] Using control point motion vectors, the motion vector for each sample within the current block can be derived. Equation 2 represents a formula for deriving a motion vector for each sample using two control point motion vectors.

[0302]

[0303] In the above mathematical formula 2, (mv x , mv y ) represents the motion vector at the (x, y) position within the current block. (mv Ax , mv Ay ) represents the first control point motion vector (A), and (mv Bx , mv By ) represents the second control point motion vector (B). W represents the width of the current block.

[0304] When three control point motion vectors are used, a motion vector per sample can be derived by the following mathematical formula 3.

[0305]

[0306] In the above mathematical formula 3, (mv Cx , mv Cy ) represents the third control point motion vector (C).

[0307] When motion vectors are derived for each sample, motion compensation can be performed for each sample, as in the example illustrated in FIG. 17. Specifically, a reference sample indicated by the motion vector of the sample to be predicted can be set as a prediction sample for the sample to be predicted.

[0308] Meanwhile, if the motion vector of the sample to be predicted is expressed in fractional units, integer position samples can be interpolated to generate fractional position samples, and the generated fractional position samples can be set as prediction samples for the sample to be predicted.

[0309] At this time, the precision of the motion vector for each sample may differ. For example, the motion vector for the first prediction target sample may be derived in units of 1 / 2 pixels, while the motion vector for the second prediction target sample may be derived in units of 1 / 4 pixels.

[0310] In this case, fractional position samples can be generated according to the motion vector precision for each of the prediction target samples. Alternatively, the motion vector of the prediction target sample can be adjusted according to the reference motion vector precision, and then prediction samples for the prediction target sample can be derived based on the adjusted motion vector. For example, if the reference motion vector precision is 1 / 2, the motion vector for the second prediction target sample can be adjusted in 1 / 4 pixel increments.

[0311] The reference motion vector precision can be determined in block units. Alternatively, the precision of the control point motion vectors can be set to the reference motion vector precision. Alternatively, the reference motion vector precision may be predefined in the encoder and decoder.

[0312] As another example, to reduce complexity, motion vectors can be derived at the sub-block level.

[0313] Figure 19 shows an example in which motion vectors are derived in sub-block units.

[0314] The size and / or shape of the sub-block may be predefined in the encoder and decoder. For example, the sub-block may be a square block of size 2x2 or 4x4.

[0315] Alternatively, the size and / or shape of the sub-block may be adaptively determined based on the size and / or shape of the current block. For example, if the current block is square, the sub-block may also be square. Conversely, if the current block is non-square, the sub-block may also be non-square.

[0316] Alternatively, information regarding at least one of the partitioning method or partitioning form of the current block may be explicitly encoded and signaled. For example, information regarding at least one of the size of a sub-block, the shape of a sub-block, the location of a partition line dividing the current block, or the number of partition lines may be explicitly encoded and signaled. The information may be encoded and signaled on a block-by-block basis, or it may be encoded and signaled through an upper header.

[0317] In Fig. 19, it was assumed that the sub-block is a square block of size 2x2.

[0318] The motion vector of a sub-block can be derived using the coordinates of a predefined location within the sub-block. Here, the predefined location may be one of the location of the top-left sample, the top-right sample, the bottom-left sample, the bottom-right sample, or the center location within the sub-block.

[0319] By substituting the coordinates of a predefined position within the sub-block into (x, y) of Equation 2, the motion vector of the sub-block can be derived.

[0320] As in the example described above, motion vectors can be derived in sub-block units based on an affine motion model.

[0321] Meanwhile, motion vectors can also be derived in sub-block units using collocated pictures. As described above, deriving motion vectors in sub-block units using collocated pictures can be referred to as SbTMVP (Sub-block Temporal Motion Vector Prediction).

[0322] A collocated picture may be one of the reference pictures included in the reference picture list. For example, a picture with index 0 in the reference picture list may be selected as the collocated picture.

[0323] Alternatively, information indicating the index of a reference picture set as a collocated picture within the reference picture list may be explicitly encoded and signaled.

[0324] Figures 20 and 21 show an example where motion vectors are induced in sub-block units within the current block when SbTMVP is applied.

[0325] The size and / or shape of the sub-block may be predefined in the encoder and decoder.

[0326] Alternatively, the size and / or shape of the sub-block may be adaptively determined according to the size and / or shape of the current block. For example, if at least one of the width or height of the current block is greater than a threshold value, the size of the sub-block may be set to 8x8. Otherwise, the size of the sub-block may be set to 4x4.

[0327] Alternatively, information indicating the size and / or shape of the sub-block may be explicitly encoded and signaled.

[0328] In the example illustrated in FIG. 20, it is assumed that the current block size is 16x16 and the sub-block size is 4x4.

[0329] When SbTMVP is applied, the initial motion vector of the current block can be derived. The initial motion vector can be derived based on at least one of a motion vector prediction list or a motion information merge list. For example, an index indicating one of the motion vector prediction candidates included in the motion vector prediction list can be encoded and signaled. The initial motion vector can be derived by adding a motion vector difference value to the motion vector prediction candidate indicated by the index. Meanwhile, the motion vector difference value can also be explicitly encoded and signaled.

[0330] Alternatively, the encoding of the index may be omitted, and a motion vector prediction candidate with a predefined index within the motion vector prediction list may be set as the prediction value for the initial motion vector. Here, the motion vector prediction candidate with a predefined index may be a motion vector prediction candidate with an index of 0 or a motion vector prediction candidate with the largest index.

[0331] Alternatively, an index indicating one of the motion information merge candidates included in the motion information merge list may be encoded and signaled. The initial motion vector may be set to be identical to the motion vector of the motion information merge candidate indicated by the index.

[0332] Alternatively, the encoding of the index can be omitted, and an initial motion vector can be derived based on a motion information merging candidate having a predefined index within the motion information merging list. Here, the motion information merging candidate having a predefined index may be a motion information merging candidate with an index of 0 or a motion information merging candidate with the largest index.

[0333] Alternatively, an initial motion vector can be derived using the motion vector of a neighbor block at a predefined position. Here, the neighbor block at the predefined position may be a left neighbor block or an top neighbor block.

[0334] The motion vector of a neighbor block at a predefined position can be set as the predicted value of the initial motion vector, and the initial motion vector can be derived by adding a difference value to the predicted value.

[0335] Alternatively, the motion vector of a neighbor block at a predefined position can be set as the initial motion vector.

[0336] Alternatively, the initial motion vector can be derived using a template-based motion estimation method (i.e., a template matching method) or two-way matching.

[0337] The precision of the initial motion vector may be predefined in the encoder and decoder. For example, the precision of the initial motion vector may be fixed in integer pixel units.

[0338] Alternatively, information indicating the precision of the initial motion vector may be explicitly encoded and signaled. The information may be an index indicating one of a plurality of motion vector precision candidates.

[0339] When deriving an initial motion vector using motion vector prediction candidates, motion vector prediction candidates can be derived based on the motion vector precision of the initial motion vector. That is, after adjusting the motion vector prediction candidates to match the motion vector precision of the initial motion vector, the adjusted initial motion vector prediction candidates can be inserted into the motion vector prediction list.

[0340] When deriving initial motion vectors using motion information merging candidates, motion information merging candidates can be derived based on the motion vector precision of the initial motion vectors. That is, after adjusting the motion information merging candidates according to the motion vector precision of the initial motion vectors, the adjusted initial motion information merging candidates can be inserted into the motion information merging list.

[0341] Meanwhile, among the motion information merging candidates included in the motion information merging list, only those candidates whose reference picture is identical to the collocated picture of the current block can be used to derive the initial motion vector. That is, if the reference picture of a motion information merging candidate is different from the collocated picture of the current block, the initial motion vector may not be derived from that motion information merging candidate.

[0342] If there are multiple candidates among the motion information merging candidates for which the reference picture is identical to the collocated picture of the current block, an index indicating one of the multiple candidates can be encoded and signaled. Alternatively, if there are multiple candidates among the motion information merging candidates for which the reference picture is identical to the collocated picture of the current block, an initial motion vector can be derived from the candidate with the smallest index or the candidate with the largest index among the multiple candidates.

[0343] If a motion information merging candidate has both motion information in the L0 direction and motion information in the L1 direction, one of the motion information in the L0 direction and the motion information in the L1 direction is selected according to a preset priority, and an initial motion vector can be derived from the selected motion information.

[0344] The priority can be determined based on at least one of the magnitude of the motion vector of the motion merge candidate, the index of the reference picture of the motion merge candidate, or whether the reference picture of the motion merge candidate is the same as the collocated picture.

[0345] Alternatively, it may be set to always derive an initial motion vector based on motion information in the L0 direction.

[0346] When initial motion vectors are derived based on a template matching method, motion estimation can be performed according to the precision of the initial motion vectors. For example, if the precision of the initial motion vectors is in the integer pixel unit, motion estimation based on template matching can also be performed only at integer locations.

[0347] Similarly, when an initial motion vector is derived based on two-way matching, motion estimation can be performed according to the precision of the initial motion vector.

[0348] Meanwhile, as a result of the two-way matching, a motion vector for the L0 direction (L0 motion vector) and a motion vector for the L1 direction (L1 motion vector) are derived. In this case, according to a pre-set priority, one of the L0 motion vector and the L1 motion vector can be set as the initial motion vector.

[0349] Alternatively, it may be set to always derive an initial motion vector based on motion information in the L0 direction.

[0350] Alternatively, information indicating which of the L0 motion vector and L1 motion vector is set as the initial motion vector may be encoded and signaled.

[0351] Once an initial motion vector is derived, the position of a collocated block within a collocated block can be determined using the initial motion vector. For example, a block located at a position offset by the initial motion vector from a position corresponding to the current block within a reference picture can be set as a collocated block. In this case, the position of the collocated block can be determined based on a predefined position within the current block. Here, the predefined position may be the top-left position, top-right position, bottom-left position, bottom-right position, or center position.

[0352] Depending on the division method of the current block, the collocated block can be divided into multiple collocated sub-blocks. Additionally, the motion vector of each collocated sub-block within the collocated block can be set as the motion vector of each sub-block within the current block.

[0353] As another example, the positions of collocated sub-blocks corresponding to each of the sub-blocks within the current block in the collocated picture can be determined using initial motion vectors. In this case, the positions of the collocated sub-blocks can be derived based on predefined positions within the sub-blocks. Here, the predefined positions may be the top-left, top-right, bottom-left, bottom-right, or center positions.

[0354] Subsequently, the motion vector of the collocated sub-block corresponding to the sub-block can be set as the motion vector of the sub-block. Specifically, the motion vector stored at a position corresponding to a predefined position within the sub-block within the collocated sub-block can be set as the motion vector of the sub-block.

[0355] Meanwhile, if the motion information of the collocated sub-block is unavailable, a predefined motion vector can be set as the motion vector of the sub-block. Here, the predefined motion vector may be a zero vector (i.e., (0, 0)) or an initial motion vector.

[0356] Alternatively, if the motion information of the collocated sub-block corresponding to the sub-block is unavailable, the motion vector of the sub-block may be derived from another location within the collocated sub-block.

[0357] Specifically, when a position corresponding to a predefined position within a collocated sub-block is encoded by intra-prediction, there is no motion vector at that position. For example, if a predefined position is assumed to be a central position (e.g., c10 in FIG. 21), and no motion vector is stored at the central position, the motion vector of the sub-block cannot be derived.

[0358] In this case, the motion vector of the sub-block can be derived based on the motion vector stored at a location different from the center position. Specifically, the motion vector of the sub-block can be derived from the motion vector stored at a location adjacent to the center position (e.g., top adjacent position c6, left adjacent position c9, or top-left adjacent position c5).

[0359] Alternatively, if the center location is unavailable, samples within the collocated sub-block may be searched according to the scan order, and the first available motion vector found may be set as the motion vector of the sub-block. Here, the scan order may be a horizontal scan, a vertical scan, a diagonal scan, or a raster scan.

[0360] Alternatively, if the motion information of the collocated sub-block is unavailable, the motion vector of the sub-block can be set as the motion vector of the collocated block. For example, the motion vector stored at a position corresponding to a previously defined position within the current block within the collocated block can be set as the motion vector of the sub-block.

[0361] As in the example described above, motion vectors can be derived in sub-block units using an affine motion model or SbTMVP. When motion vectors are derived in sub-block units, motion compensation can be performed for each sub-block based on the motion vector of each sub-block.

[0362] By performing motion compensation for each of the sub-blocks, a prediction block for the current block can be obtained. That is, the prediction block may be composed of prediction samples for each of the sub-blocks.

[0363] When detecting movement between frames, the precision of the motion vector can be adjusted. Specifically, the position of each sample within a picture is defined as an integer. However, the position reflecting the movement can be a decimal position rather than an integer position.

[0364] Considering this, motion vectors can be explored more precisely through reference picture interpolation.

[0365] Figures 22 and 23 are diagrams illustrating examples in which a prediction block is derived according to the precision of the motion vector.

[0366] FIG. 22 shows the position of the current block in the current picture, and FIG. 23 shows the position of the reference block according to the motion vector precision.

[0367] As shown in the examples illustrated in FIGS. 22 and 23, the motion vector of the current block can be defined as the distance from a sample corresponding to the top-left position of the current block in the reference picture to a sample corresponding to the top-left position of the reference block in the reference picture.

[0368] FIG. 23 (a) illustrates the case where the motion vector precision of the current block is an integer Pel, FIG. 23 (b) illustrates the case where the motion vector precision of the current block is 1 / 2 Pel. Also, FIG. 23 (c) illustrates the case where the motion vector precision of the current block is 1 / 4 Pel.

[0369] In FIG. 23, the vector precision is expressed up to 1 / 4, but the motion vector can be expressed with even greater precision, such as 1 / 8, 1 / 16, or 1 / 32.

[0370] Meanwhile, information for indicating the motion vector precision of the current block may be encoded and signaled. For example, the information may be an index identifying one of the motion vector precision candidates. Specifically, a different index may be assigned to each of the motion vector precision candidates, and the information may indicate the index of the motion vector precision candidate applied to the current block.

[0371] By adjusting the precision of the motion vectors used for cross-frame prediction, more precise motion vector detection may be possible. If the reference block indicated by the motion vector exists at a real-valued location, the samples at the real-valued location can be generated using samples at integer locations and an interpolation filter. Additionally, motion vectors represented by real numbers can be scaled up to integers for encoding / decoding.

[0372] Thus, the motion vector (MV), motion vector predicted value (MVP), and motion vector difference value (MVD) can be encoded / decoded into integer values ​​through integerization. Specifically, the motion vector, motion vector predicted value, and / or motion vector difference value can be integerized based on the motion vector precision.

[0373] For example, if the motion vector precision is 1 / N, the motion vector difference value MVD can be converted to an integer by multiplying it by N. For example, if the motion vector difference value MVD is (4 / 16, 8 / 16), the motion vector difference value MVD can be converted to an integer by multiplying it by 16. That is, the converted motion vector difference value MVD can be expressed as (4, 8).

[0374] Based on motion vector precision, the actual MVD can be derived from the integerized MVD. For example, if the motion vector precision is 1 / N, the actual MVD can be derived by dividing the integerized MVD by N. For example, if the integerized MVD is (4, 8) and the motion vector precision is 1 / 8, the actual MVD can be (4 / 8, 8 / 8). Or, if the integerized MVD is (4, 8) and the motion vector precision is 1 / 4, the actual MVD can be (4 / 4, 8 / 4).

[0375] Depending on the motion vector precision, the range of representation of the integerized MVD may differ. For example, assume that the motion vector difference value MVD is (4 / 16, 8 / 16) (i.e., (1 / 4, 2 / 4)). When the motion vector precision is 1 / 16, the integerized MVD is derived as (4, 8). On the other hand, when the motion vector precision is 1 / 4, the integerized MVD is derived as (1, 2).

[0376] Comparing the two cases above, if the motion vector precision is adjusted from 1 / 16 to 1 / 4, the value of the integerized MVD can be reduced from (4, 8) to (1, 2).

[0377] Consequently, depending on the motion vector precision, the number of bits required to encode / decode the integerized motion vector difference value MVD may vary. Accordingly, a motion vector precision that minimizes the number of bins can be selected when encoding / decoding the motion vector difference value MVD. Then, based on the selected motion vector precision, the motion vector difference value MVD can be converted to an integer, and the integerized motion vector difference value MVD can be encoded / decoded. In addition, information regarding the motion vector precision can be additionally encoded / decoded.

[0378] In the decoder, the actual MVD can be restored from the decoded MVD based on motion vector precision. Then, the motion vector MV can be derived by combining the restored MVD and the motion vector prediction value MVP.

[0379] As described above, adjusting the value of the motion vector difference value MVD, which is encoded / decoded based on motion vector precision, is called the AMVR (Adaptive Motion Vector Resolution) method.

[0380] FIGS. 24 and FIGS. 25 are diagrams illustrating the process of encoding and decoding motion vector difference values ​​when the AMVR method is applied, respectively.

[0381] For the sake of convenience of explanation, it is assumed that the motion vector and the motion vector difference value are expressed in units of 1 / 16 before integerization is performed, and 1 / 16 is referred to as the original motion vector precision.

[0382] The motion vector difference value MVD can be derived by differencing the motion vector prediction value MVP from the motion vector MV (S2410).

[0383] The motion vector difference value MVD may consist of a horizontal component (i.e., the x-axis component) and a vertical component (i.e., the y-axis component).

[0384] When the motion vector difference value is 0, that is, when both the horizontal and vertical components are 0, the value of the motion vector difference value MVD to be encoded becomes 0 regardless of the motion vector precision. Therefore, when the motion vector difference value MVD is 0, the encoding of AMVR-related information can be omitted (S2420).

[0385] On the other hand, if the motion vector difference value is not zero, that is, if at least one of the horizontal component and the vertical component is not zero, the motion vector precision can be determined (S2430). Meanwhile, the motion vector precision can be encoded as AMVR-related information.

[0386] Information related to AMVR may include at least one of a flag (e.g., amvr_flag) indicating whether the AMVR method is applied to the current block and an index (e.g., amvr_prec_idx) indicating one of a plurality of motion precision candidates if the AMVR method is applied.

[0387] If the AMVR method is not applied to the current block, the motion vector precision can be set to a default value. In this case, amvr_flag can be encoded as a value of 0. Meanwhile, the default value can be 1, 1 / 2, 1 / 4, 1 / 8, or 1 / 16.

[0388] When the AMVR method is applied to the current block, an index indicating one of multiple motion vector precision candidates, i.e., amvr_prec_idx, may be additionally decoded. In this case, amvr_flag is encoded with a value of 1, and amvr_prec_idx may be encoded with a value from 0 to (n-1). Here, n represents the number of motion vector precision candidates. For example, multiple motion vector precision candidates may include at least one of 4, 2, 1, 1 / 2, 1 / 4, 1 / 8, or 1 / 16. Meanwhile, the default value may not be set to the multiple motion vector precision candidates indicated by the index. That is, if the motion vector precision of the current block is the default value, it is encoded and signaled as 0, which is the value of amvr_flag, and the encoding of amvr_prec_idx may be omitted.

[0389] In the encoder, the optimal motion vector precision can be determined by performing Rate Distortion Optimization (RDO) for each combination of amvr_flag and amvr_prec_idx. That is, by performing RDO for the following cases, the combination with the optimal cost can be selected.

[0390] 1) When amvr_flag is 0

[0391] 2) When amvr_flag is 1 and amvr_prec_idx is 0

[0392] 3) When amvr_flag is 1 and amvr_prec_idx is 1

[0393] 4) When amvr_flag is 1 and amvr_prec_idx is 2

[0394] Depending on the motion vector precision of the current block, a variable for scaling the motion vector difference value, i.e., a scaling parameter, can be set. For example, Table 1 shows the values ​​of the variable amvrshift according to the motion vector precision.

[0395] amvr_flagamvr_prec_idxamvrshift0 (1 / 4)-210 (1 / 2)311 (1-pel)412 (4-pel)6

[0396] If the finest motion vector precision applicable to the current block is 1 / 16, the motion vector precision can be expressed as shown in the following mathematical formula 4.

[0397]

[0398] As shown in Table 1, when the value of amvr_flag is 0, the variable amvrshift is set to 2. This indicates that the motion vector precision is 1 / 4 according to Equation 4.

[0399] When the value of amvr_flag is 1, the variable amvrshift can be determined according to the value of amvr_prec_idx. For example, when amvr_prec_idx is 1, the variable amvrshift is set to 4. This indicates that the motion vector precision is 1 according to Equation 4.

[0400] In the encoder, the motion vector difference value MVD can be scaled down and encoded using the variable amvrshift, which is based on the motion vector precision. As an example, Equation 5 shows an example of a scale-down operation being performed on the motion vector difference value MVD.

[0401]

[0402] In the above mathematical formula 5, MVD_x represents the horizontal component of the motion vector difference value, and MVD_y represents the vertical component of the motion vector difference value. MVD'_x and MVD'_y represent the results of performing a scale-down operation.

[0403] The encoder can encode motion vector difference values ​​with changed precision and AMVR information (S2440).

[0404] In the decoder, the motion vector difference value MVD can be decoded (S2510).

[0405] If the motion vector difference value is 0, the decoding of AMVR-related information is omitted, and the motion vector MV of the current block can be set to be the same as the motion vector prediction value (S2520).

[0406] On the other hand, if the motion vector difference value is not zero, that is, if at least one of the horizontal component and the vertical component is not zero, information related to AMVR can be additionally decoded (S2530).

[0407] Based on AMVR information, a variable amvrshift for scaling motion vector difference values ​​can be derived. For example, as shown in the example in Table 1, a variable amvrshfit can be derived based on amvr_flag and / or amvr_prec_idx.

[0408] Afterwards, the decoded MVD can be scaled up using the variable amvrshift to obtain the motion vector difference value MVD restored to its original precision (S2540). Equation 6 shows an example of a scale-up operation applied to the decoded MVD.

[0409]

[0410] In Equation 6, MVD' represents the decoded motion vector difference value. MVD represents the motion vector difference value restored to its original precision, i.e., 1 / 16, through a scale-up operation.

[0411] Afterwards, the motion vector MV can be obtained by combining the motion vector difference value MVD restored to the original precision and the motion vector prediction value MVP.

[0412] As in the example above, when a motion vector prediction mode is applied, the decoder can derive the motion vector MV by combining the motion vector prediction value MVP and the motion vector difference value MVD.

[0413]

[0414] As described above, the movement information of the current block can be derived from at least one of the spatial neighbor blocks or temporal neighbor blocks of the current block.

[0415] As another example, movement information of the current block can be derived using blocks located at positions not adjacent to the current block (i.e., non-adjacent blocks). For instance, movement information merging candidates derived based on the movement information of non-adjacent blocks not adjacent to the current block can be added to the movement information merging list of the current block.

[0416] The location of a non-adjacent block can be determined based on at least one of the width or height of the current block. For example, the horizontal distance between a non-adjacent block and the current block may be N times the width of the current block. Here, the horizontal distance represents the difference in x-axis coordinates between the two blocks. Additionally, the vertical distance between a non-adjacent block and the current block may be M times the height of the current block. Here, the vertical distance may represent the difference in y-axis coordinates between the two blocks.

[0417] Meanwhile, N and M may each be integers including 0. Multiple non-adjacent blocks may be searched by increasing the values ​​of N and / or M by 1. For example, between the first non-adjacent block and the second non-adjacent block, at least one of the x-axis coordinate or the y-axis coordinate may be different.

[0418] For example, the second non-adjacent block may be spaced horizontally from the first non-adjacent block by the width of the current block.

[0419] Alternatively, the second non-adjacent block may be spaced vertically from the first non-adjacent block by the height of the current block.

[0420] Alternatively, the second non-adjacent block may be spaced apart from the first non-adjacent block by the width of the current block in the horizontal direction and by the height of the current block in the vertical direction.

[0421] Non-adjacent blocks may exist in the currently restored area within the picture.

[0422] Depending on the location of the restored area within the current picture, non-adjacent blocks may exist to the left or right of the current block in the horizontal direction. Additionally, non-adjacent blocks may exist above or below the current block in the vertical direction.

[0423] Even if motion information merge candidates derived from spatial neighbor blocks, spatial non-adjacent blocks, or temporal neighbor blocks are inserted into the motion information merge list, if the number of motion information merge candidates contained in the motion information merge list is less than a threshold value, additional motion information merge candidates can be generated using the motion information merge candidates already inserted into the motion information merge list. Here, the threshold value may be the maximum number of motion information merge candidates that the motion information merge list can contain, or a value obtained by subtracting an offset value from the said maximum number. For convenience of explanation, the threshold value is assumed to be the maximum number of motion information merge candidates, and the value is assumed to be 6.

[0424] Figure 26 shows an example of generating a new motion information merge candidate using motion information merge candidates that are already inserted into the motion information merge list.

[0425] In FIG. 26 (a), it is exemplified that there are only 4 motion information merging candidates in the motion information merging list.

[0426] Since the number of motion information merge candidates included in the motion information merge list is smaller than the threshold value, new motion information merge candidates can be derived based on the motion information merge candidates already inserted in the motion information merge list.

[0427] Specifically, a new motion information merge candidate can be generated using the top two motion information merge candidates (i.e., the two motion information merge candidates with the smallest indices) within the motion information merge list. Subsequently, the motion information merge list can be updated by inserting the newly generated motion information merge candidate into the motion information merge list.

[0428] In the example illustrated in Fig. 26 (b), a motion information merge candidate with index 4 is newly added to the motion information merge list.

[0429] Meanwhile, additional motion information merging candidates derived based on two motion information merging candidates already inserted in the motion information merging list can be called pair-wise motion information merging candidates.

[0430] When two motion information merging candidates have a common prediction direction, the motion vector of an additional motion information merging candidate for that direction can be derived by averaging the motion vectors of the two motion information merging candidates. For example, in the example illustrated in FIG. 26 (b), the L0 motion vector of an additional motion information merging candidate (i.e., a motion information merging candidate with an index of 4) can be derived by averaging the motion vector of a motion information merging candidate with an index of 0 and the L0 motion vector of a motion information merging candidate with an index of 1.

[0431] Furthermore, the reference picture index of the additional motion information merging candidate can be set to be the same as one of the two motion information merging candidates. Specifically, the reference picture index of the motion information merging candidate with the smaller index among the two motion information merging candidates can be set as the reference picture index of the additional motion information merging candidate. That is, in the example illustrated in FIG. 26 (b), the L0 reference picture index of the additional motion information merging candidate can be set to be the same as the reference picture index of the motion information merging candidate with an index of 0.

[0432] If, among the two motion information merging candidates, the first motion information merging candidate has motion information for a specific direction, while the second motion information merging candidate does not have motion information for a specific direction, the motion information for a specific direction of the additional motion information merging candidate can be set to be the same as that of the first motion information merging candidate.

[0433] For example, in the example illustrated in FIG. 26 (a), the motion information merging candidate with index 0 has bidirectional motion information, whereas the motion information merging candidate with index 1 has L0 direction motion information. Since the motion information merging candidate with index 1 does not have L1 direction motion information, the L1 direction motion information of the motion information merging candidate with index 0 can be set as the L1 direction motion information of the additional motion information merging candidate.

[0434] That is, among the two motion information merging candidates, the first motion information merging candidate has LX motion information and L(1-X) motion information, while the second motion information merging candidate has only LX motion information, the LX motion vector of the additional motion information merging candidate is derived by averaging the LX motion vector of the first motion information merging candidate and the LX motion vector of the second motion information merging candidate, whereas the L(1-X) motion vector of the additional motion information merging candidate can be set to be the same as the L(1-X) motion vector of the first motion information merging candidate. Here, X can be 0 or 1.

[0435] As another example, among two motion information merging candidates, if the first motion information merging candidate has bidirectional motion information while the second motion information merging candidate has unidirectional motion information, a third motion information merging candidate can be additionally used to derive an additional motion information merging candidate. In this case, the third motion information merging candidate may have motion information in a direction that the second motion information merging candidate does not have.

[0436] For example, it is assumed that the first motion information merging candidate has both LX motion information and L(1-X) motion information, while the second motion information merging candidate has LX motion information. In this case, a third motion information merging candidate having L(1-X) motion information can be additionally searched.

[0437] In this case, the LX motion vector of the additional motion information merging candidate is derived by averaging the LX motion vector of the first motion information merging candidate and the LX motion vector of the second motion information merging candidate, and the L(1-X) motion vector of the additional motion information merging candidate can be derived by averaging the L(1-X) motion vector of the first motion information merging candidate and the L(1-X) motion vector of the third motion information merging candidate.

[0438] Figure 27 shows an example in which additional motion information merging candidates are derived based on three motion information merging candidates.

[0439] In the example illustrated in FIG. 27(a), the motion information merging candidate with index 0 is exemplified as having motion information in the L0 direction and L1 direction, while the motion information merging candidate with index 1 is exemplified as having motion information in the L0 direction. Accordingly, additional motion information merging candidates having motion information in the L1 direction can be searched. The additional search may involve searching the remaining motion information merging candidates, excluding the previously selected motion information merging candidates, in index order.

[0440] Referring to the example illustrated in FIG. 27 (a), since the motion information merging candidate with index 2 contains motion information in the L1 direction, the motion information merging candidate with index 2 can be additionally selected.

[0441] Accordingly, the L0 motion vector of an additional motion information merging candidate (i.e., a motion information merging candidate with an index of 4) can be derived by averaging the L0 motion vector of a motion information merging candidate with an index of 0 and the L0 motion vector of a motion information merging candidate with an index of 1. Meanwhile, the L0 reference picture index of the additional motion information merging candidate can be set to be the same as the reference picture index of the motion information merging candidate with the smaller index among the two motion information merging candidates (i.e., a motion information merging candidate with an index of 0).

[0442] Likewise, the L1 motion vector of an additional motion information merge candidate (i.e., a motion information merge candidate with an index of 4) can be derived by averaging the L1 motion vector of a motion information merge candidate with an index of 0 and the L1 motion vector of a motion information merge candidate with an index of 2. Meanwhile, the L1 reference picture index of an additional motion information merge candidate can be set to be the same as the reference picture index of the motion information merge candidate with the smaller index among the two motion information merge candidates (i.e., a motion information merge candidate with an index of 0).

[0443] Unlike the example described, additional motion information merge candidates can be derived by selecting two motion information merge candidates that have bidirectional motion information from among the motion information merge candidates already inserted in the motion information merge list. That is, motion information merge candidates that have unidirectional motion information can be set as unavailable for use in generating additional motion information merge candidates.

[0444] By sequentially searching for motion information merging candidates in order of decreasing index, two motion information merging candidates having bidirectional motion information can be selected to derive additional motion information merging candidates. In the example illustrated in FIG. 27 (a), since the motion information merging candidate with index 0 and the additional motion information merging candidate with index 2 have bidirectional motion information, additional motion information merging candidates can be derived based on the above two motion information merging candidates.

[0445] On the other hand, in the example illustrated in FIG. 27 (a), the motion information merge candidate with index 1 and the motion information merge candidate with index 3, which have unidirectional motion information, may not be available for generating additional motion information merge candidates.

[0446] As another example, if the number of motion information merging candidates having bidirectional motion information is less than 2, additional motion information merging candidates can be derived using motion information merging candidates having unidirectional motion information according to the above-described embodiment.

[0447] As another example, if among two motion information merging candidates, the first motion information merging candidate has bidirectional motion information while the second motion information merging candidate has unidirectional motion information, the additional motion information merging candidate can also be configured to have unidirectional motion information.

[0448] For example, if the first motion information merging candidate has L0 motion information and L1 motion information, while the second motion information merging candidate has L0 motion information, the L0 motion vector of the additional motion information merging candidate can be derived based on the motion vectors for the L0 direction that the first and second motion information merging candidates have in common. On the other hand, since the second motion information merging candidate does not have motion information for the L1 direction, the motion information for the L1 direction of the additional motion information merging candidate can be set as unavailable.

[0449] Alternatively, additional motion information merge candidates can be derived by additionally selecting motion information merge candidates that have the same direction as the top motion information merge candidate. Here, the top motion information merge candidate may represent the motion information merge candidate with the smallest index.

[0450] For example, if a motion information merge candidate with index 0 has LX motion information but does not have L(1-X) motion information, additional motion information merge candidates having LX motion information can be searched. In this case, since the motion information merge candidate with index 0 has unidirectional motion information, the additionally searched motion information merge candidates can also be restricted to having unidirectional motion information.

[0451] On the other hand, if a motion information merge candidate with index 0 has bidirectional motion information, additional motion information merge candidates having bidirectional motion information can be searched.

[0452] Additional motion information merge candidates can be derived using the top motion information merge candidate and additionally searched motion information merge candidate. That is, if the motion information merge candidate with index 0 has only LX motion information, the additional motion information merge candidate can also have only LX motion information.

[0453] On the other hand, if a motion information merge candidate with index 0 has bidirectional motion information, additional motion information merge candidates may also have bidirectional motion information.

[0454] The number of newly generated additional movement information merge candidates can be 1.

[0455] As another example, the maximum number of newly generated additional motion information merge candidates can be set to N. Here, N can be a natural number such as 1, 2, 3, 4, 5, or 6.

[0456] The maximum number of additional motion information merge candidates can be determined based on at least one of the number of motion information merge candidates already inserted in the motion information merge list or the number of empty entries in the motion information merge list. Here, the number of empty entries may represent the value obtained by subtracting the number of motion information merge candidates already inserted in the motion information merge list from the maximum number of motion information merge candidates that the motion information merge list can contain.

[0457] For example, if M motion information merge candidates are inserted into the motion information merge list, then at most ( M C2) Additional movement information merge candidates can be generated.

[0458] Alternatively, the maximum number of additional motion information merge candidates can be set to the smaller value between the maximum number of combinations of motion information merge candidates already inserted in the motion information merge list and the number of empty entries.

[0459] Figure 28 shows an example in which additional motion information merging candidates are generated up to the maximum number of combinations of motion information merging candidates already inserted in the motion information merging list.

[0460] For the sake of convenience of explanation, it is assumed that the maximum number of motion information merge candidates that can be included in the motion information merge list is 10.

[0461] In the example illustrated in FIG. 28 (a), since four motion information merging candidates are inserted into the motion information merging list, up to six additional motion information merging candidates can be newly generated.

[0462] Specifically, through the following combinations, six additional motion information merging candidates can be generated.

[0463] Movement information merge candidate with index 0 and movement information merge candidate with index 1

[0464] Movement information merge candidate with index 0 and movement information merge candidate with index 2

[0465] Movement information merge candidate with index 1 and movement information merge candidate with index 2

[0466] Movement information merge candidate with index 0 and movement information merge candidate with index 3

[0467] Movement information merge candidate with index 1 and movement information merge candidate with index 3

[0468] Movement information merge candidate with index 2 and movement information merge candidate with index 3

[0469] Following the above order, additional motion information merge candidates can be generated sequentially, and the generated additional motion information merge candidates can be inserted into the motion information merge list.

[0470] Meanwhile, additional movement information merge candidates can be set as unavailable for generating other additional movement information merge candidates.

[0471] As another example, even though additional motion information merge candidates have been generated using all combinations of motion information merge candidates already inserted in the motion information merge list, if the number of motion information merge candidates contained in the motion information merge list is less than a threshold value, it may be configured to generate additional pairwise mean motion information candidates using the additional motion information merge candidates.

[0472] In other words, after adding additional motion information merge candidates, the number of motion information merge candidates included in the motion information merge list is compared with a threshold value to determine whether to use the additional motion information merge candidates to generate additional additional motion information merge candidates. Specifically, if the number of motion information merge candidates is less than the threshold value, the additional motion information merge candidates can be used to generate additional additional motion information merge candidates; conversely, if the number of motion information merge candidates is equal to or less than the threshold value, the additional motion information merge candidates can be set as unavailable for generating additional additional motion information merge candidates.

[0473] In the example described above, it was exemplified that an additional motion information merge candidate is generated based on two motion information merge candidates already inserted into the motion information merge list.

[0474] Unlike the example described, additional motion information merge candidates can be derived using a number of motion information merge candidates greater than two. Depending on the number of motion information merge candidates used to generate the additional motion information merge candidates, the additional motion information merge candidates can be referred to as M-th order motion information merge candidates. Here, M represents the number of motion information merge candidates used to derive the additional motion information merge candidates.

[0475] The value of M may be predefined in the encoder and decoder. For example, M may be a natural number such as 2, 3, or 4.

[0476] Alternatively, M can be determined based on the number of motion information merge candidates already included in the motion information merge list.

[0477] For example, if the number of motion information merge candidates already included in the motion information merge list is equal to or greater than the threshold value, the threshold value can be set to M. Conversely, if the number of motion information merge candidates already included in the motion information merge list is less than the threshold value, the number of motion information merge candidates already included in the motion information merge list can be set to M.

[0478] Figure 29 shows an example in which a candidate for 4th movement information merging is derived.

[0479] Four motion information can be used to derive motion information for additional motion information merging candidates (i.e., motion information merging candidates with index 4).

[0480] For example, in the example illustrated in FIG. 29, the L0 motion vector of the additional motion information merging candidate can be derived by averaging the L0 motion vectors of four motion information merging candidates from index 0 to index 3. That is, the L0 motion vector of the additional motion information merging candidate can be derived based on the following mathematical formula 7.

[0481]

[0482] Meanwhile, in the example illustrated in FIG. 29, since the motion information merging candidate with index 1 does not have L1 motion information, there are no four motion informations for the L1 direction within the motion information merging list. In this case, the L1 motion information of the additional motion information merging candidate can be set as unavailable.

[0483] In this way, if the number of available motion information for a specific direction is less than a threshold value (e.g., 4), the additional motion information for that specific direction may be set as unavailable.

[0484] Alternatively, even if the number of available motion information for a specific direction is less than the threshold value, motion information for a specific direction of additional motion information candidates can be derived based on the available motion information for a specific direction.

[0485] For example, in the example illustrated in FIG. 29, three of the four motion information merging candidates (i.e., motion information merging candidates with indices 0, 2, and 3) have motion information in the L1 direction. Accordingly, the L1 motion vector of an additional motion information merging candidate can be derived by averaging the L1 motion vectors of the three motion information merging candidates. That is, the L1 motion vector of an additional motion information merging candidate can be derived based on the following Equation 8.

[0486]

[0487] Alternatively, if the number of available motion information for a specific direction is less than a threshold value, a motion vector of additional motion information can be derived by using a predefined number of motion information for that specific direction.

[0488] For example, in the example illustrated in FIG. 29, among the three motion information merging candidates having motion information in the L1 direction, the L1 motion vector of additional motion information can be derived using the two motion information merging candidates with the smallest indices. That is, the L1 motion vector of the additional motion information merging candidate can be derived based on the following mathematical formula 9.

[0489]

[0490] Alternatively, by assuming the motion vector of a motion information candidate that does not have motion information for a specific direction as a zero vector, the motion vector for a specific direction of an additional motion information merging candidate can be derived.

[0491] For example, in the example illustrated in FIG. 29, the L0 motion vector of a motion information merging candidate with an index of 1 where L0 motion information does not exist can be assumed to be (0, 0), and the L1 motion vector of an additional motion information merging candidate can be derived. That is, the L1 motion vector of an additional motion information merging candidate can be derived based on the following mathematical formula 10.

[0492]

[0493] In deriving multiple additional motion information candidates, the order may be set differently.

[0494] For example, even though an A-th degree motion information merge candidate is inserted into the motion information merge list, if the number of motion information merge candidates included in the motion information merge list is less than a threshold value, an (A+n)-th degree motion information merge candidate can be additionally induced. For example, after inducing a 2-th degree motion information merge candidate based on 2 motion information merge candidates, a 3-th degree or 4-th degree motion information merge candidate can be additionally induced based on 3 or 4 motion information merge candidates.

[0495] In the example described above, the reference picture index of an additional motion information merge candidate is exemplified as being derived from the motion information merge candidate with the smallest index among the multiple motion information merge candidates used to derive the said additional motion information merge candidate.

[0496] As another example, after deriving template costs for multiple reference pictures, the reference picture index of the additional motion information merge candidate can be set to point to the reference picture with the smallest cost.

[0497] Here, multiple reference pictures may be all reference pictures included in the reference picture list or reference pictures indicated by the reference picture index of multiple motion information merge candidates used to derive additional motion information merge candidates.

[0498] For example, as shown in FIG. 26, if additional motion information is derived based on two motion information merging candidates, a template cost can be calculated for each of the first L0 reference picture indicated by the L0 reference picture index refIdx_B1_L0 of the first motion information merging candidate (i.e., the motion information merging candidate with index 0) among the two motion information merging candidates, and the second L0 reference picture indicated by the L0 reference picture index refIdx_A1_L0 of the second motion information merging candidate (i.e., the motion information merging candidate with index 1) among the two motion information merging candidates.

[0499] Template costs can be calculated based on the motion vectors of additional motion information merging candidates. That is, template costs for each of the first L0 reference picture and the second L0 reference picture can be calculated based on the average of the L0 motion vectors of the two motion information merging candidates.

[0500] The template cost can be derived based on the difference between the current template and the reference template. Here, the reference template may be an area located away from the current template within the reference picture by the motion vector of the additional motion information merge candidate.

[0501] When the template cost for each of the reference pictures is calculated, the reference picture index of the additional motion information merge candidate can be set to point to the reference picture with the smallest cost. For example, if the template cost of the first L0 reference picture is smaller than the template cost of the second L0 reference picture, the L0 reference picture index of the additional motion information merge candidate can be set to point to the first L0 reference picture.

[0502] Even when more than two motion information merge candidates are used, a reference picture index can be set based on the template cost.

[0503] When the motion information merging mode is used, information indicating one of the multiple motion information merging candidates included in the motion information merging list may be encoded and signaled. The said information may be an index indicating one of the multiple motion information merging candidates.

[0504] Meanwhile, the maximum number of motion information merging candidates that the motion information merging list can include may be predefined in the encoder and decoder. Alternatively, information indicating the maximum number may be encoded and signaled.

[0505] In the following embodiments, it is assumed that the maximum number of motion information merging candidates that can be included in the motion information merging list is 5.

[0506] Figure 30 illustrates an example of the configuration of a motion information merging list.

[0507] In the example illustrated in FIG. 30, each motion information merging candidate may have motion information for at least one of the L0 direction or the L1 direction. The motion information may include a motion vector and a reference picture index. Although not illustrated, the motion information may further include at least one of a prediction direction flag indicating whether motion information for the L0 direction or the L1 direction exists, a flag indicating whether brightness compensation is used, or weight information for bidirectional prediction.

[0508] If a motion information merging candidate has motion information for the L0 direction and L1 direction, respectively, the motion information merging candidate can be understood as having bidirectional motion information. On the other hand, if a motion information merging candidate has motion information for either the L0 direction or the L1 direction, the motion information merging candidate can be understood as having unidirectional motion information.

[0509] In the example illustrated in FIG. 30, a unique index is assigned to each of the motion information merging candidates. In the encoder, the index of a selected motion information merging candidate among the motion information merging candidates can be encoded and signaled to the decoder.

[0510] A unique index may be assigned to combinations of motion information merging candidates. In the encoder, the index of a selected combination among the combinations of motion information merging candidates can be encoded and signaled to the decoder. Accordingly, the decoder can select multiple motion information merging candidates based on a single index.

[0511] Figure 31 shows an example in which a unique index is assigned to a combination of multiple motion information merging candidates.

[0512] In the example illustrated in FIG. 31, a unique index is assigned to a combination of a motion information merger candidate belonging to the first motion information merger list and a motion information merger candidate belonging to the second motion information merger list. That is, index N can represent a combination of a motion information merger candidate with index N in the first motion information merger list and a motion information merger candidate with index N in the second motion information merger list.

[0513] In the example illustrated in FIG. 31, a unique index is assigned to a combination of two motion vector merge candidates. Unlike the illustrated example, a unique index may be assigned to a combination of more than two motion vector merge candidates.

[0514] For example, the number of motion information merging candidates included in a single combination may be a power of 2. For example, a single combination may contain 2, 4, or 8 motion information merging candidates.

[0515] When one of a plurality of combinations is selected, a prediction block of the current block can be obtained based on a plurality of motion information merging candidates belonging to the selected combination. For example, if the selected combination includes two motion information merging candidates, a prediction block of the current block can be obtained based on each of the two motion information merging candidates. For example, a first prediction block for the current block can be obtained based on the first motion information merging candidate among the two motion information merging candidates, and a second prediction block for the current block can be obtained based on the second motion information merging candidate among the two motion information merging candidates.

[0516] Afterwards, the final prediction block of the current block can be obtained by weighting the multiple prediction blocks (i.e., the first prediction block and the second prediction block).

[0517] Meanwhile, the weights applied to the first prediction block and the second prediction block may be the same or different. In the encoder, information regarding the weights applied to the first prediction block and the second prediction block can be encoded and signaled.

[0518] Alternatively, it can be configured so that a higher weight is assigned to the prediction block derived from the higher priority among the two motion information merge candidates.

[0519] For example, if the priority of the first motion information merging candidate is higher than that of the second motion information merging candidate, a larger weight can be assigned to the first prediction block than to the second prediction block. Meanwhile, the ratio of weights assigned to the two prediction blocks can be 3:1 or 7:1.

[0520] Meanwhile, if a motion information merging candidate has bidirectional motion information, the prediction block derived from the said motion information merging candidate can be derived by weighted summing the L0 prediction block and the L1 prediction block.

[0521] At this time, the weights assigned to the L0 prediction block and the L1 prediction block can be set to be the same.

[0522] Alternatively, weights assigned to the L0 prediction block and L1 prediction block can be set based on the weight information held by the motion information merging candidate.

[0523] The motion information of a motion information merging candidate can be corrected, and a prediction block can be derived based on the corrected motion information. For example, if the motion information merging candidate has bidirectional motion information, the L0 motion vector and the L1 motion vector can be corrected based on a bidirectional matching method. Subsequently, an L0 prediction block is derived based on the corrected L0 motion vector, and an L1 prediction block is derived based on the corrected L1 motion vector. Then, the L0 prediction block and the L1 prediction block are weighted to obtain a prediction block based on the motion information merging candidate.

[0524] This results in the same outcome as correcting the L0 prediction block obtained by the initial L0 motion vector and the L1 prediction block obtained by the initial L1 motion vector using a two-way matching method.

[0525] If a flag indicating whether to use brightness compensation for motion information merging candidates indicates that brightness compensation is performed, a prediction block may be obtained by applying brightness compensation to an initial prediction block derived based on the motion information of the motion information merging candidates. Meanwhile, a flag indicating whether to use brightness compensation may exist in both the L0 direction and the L1 direction. Based on the flag in the L0 direction, it may be determined whether to perform brightness compensation on the initial prediction block in the L0 direction, and based on the flag in the L1 direction, it may be determined whether to perform brightness compensation on the initial prediction block in the L1 direction.

[0526] Meanwhile, motion information merging candidates to be inserted into the first motion information merging list and the second motion information merging list can be derived from the candidate blocks of the current block.

[0527] Candidate blocks may include at least one of the temporal neighbor blocks, spatial neighbor blocks, and non-adjacent blocks of the current block.

[0528] FIG. 32 is a diagram illustrating candidate blocks for inducing motion information merging candidates.

[0529] In the example illustrated in FIG. 32, A0 to A4 represent neighboring blocks adjacent to the left of the current block, and B0 to B5 represent neighboring blocks adjacent to the top of the current block.

[0530] Furthermore, a0 to a4 represent blocks adjacent to the left of A0 to A4, and b0 to b5 represent blocks adjacent to the top of B0 to B5. Meanwhile, b6 represents a block adjacent to the left of neighbor block B5, which is adjacent in the upper-left diagonal direction from the current block, and b7 represents a block adjacent in the upper-left diagonal direction from B5. Meanwhile, in this embodiment, 'block' may represent a coding block, which is an actual encoding / decoding unit.

[0531] Alternatively, a 'block' may be an area of ​​arbitrary size. For example, in the example illustrated in FIG. 32, A1 and A2 may be included in a single CB. That is, A1 may correspond to a first area within a neighboring coding block, and A2 may correspond to a second area within a neighboring coding block. In this case, the width and height of each area may be determined based on the width and height of the current block.

[0532] Alternatively, the width and height of each region may be predefined in the encoder / decoder.

[0533] In the illustrated example, a0 to a4 represent non-adjacent blocks located to the left of the current block and not adjacent to the current block, and 0 to b7 represent non-adjacent blocks located at the top of the current block and not adjacent to the current block.

[0534] Candidates for motion information merging can be derived by sequentially searching neighboring blocks of the current block according to a predefined order. At this time, a predefined number of candidates for motion information merging can be inserted into the first motion information merging list, and the remaining candidates for motion information merging can be inserted into the second motion information merging list. That is, candidates for motion information merging can be added to the first motion information merging list until the number of candidates for motion information merging contained in the first motion information merging list reaches a threshold value, and when the number of candidates for motion information merging contained in the first motion information merging list reaches a threshold value, candidates for motion information merging can be inserted into the second motion information list.

[0535] Here, the predefined order may be to scan neighbor blocks in an upward direction starting from the neighbor block adjacent to the bottom-left diagonal of the current block (i.e., from A0 to A4), and then scan neighbor blocks in a right direction starting from the neighbor block adjacent to the top-left diagonal of the current block (i.e., from B5 to B0).

[0536] Alternatively, motion information merging candidates may be inserted alternately into the first motion information merging list and the second motion information list. For example, when neighboring blocks are searched according to a predefined order, the first available motion information found can be inserted into the first motion information merging list as a motion information merging candidate. Conversely, when neighboring blocks are searched according to a predefined order, the second available motion information found can be inserted into the second motion information merging list as a motion information merging candidate.

[0537] Consequently, in the first motion information merging list, motion information merging candidates corresponding to odd-numbered search orders can be inserted, and in the second motion information merging list, motion information merging candidates corresponding to even-numbered search orders can be inserted.

[0538] Alternatively, motion information merging candidates derived from neighbor blocks of a predefined location may be inserted into the first motion information merging list, and motion information merging candidates derived from blocks adjacent to neighbor blocks of a predefined location may be inserted into the second motion information merging list.

[0539] For example, if a motion information merge candidate derived from neighbor block A1 is inserted into the first motion information merge list, a motion information merge candidate derived from neighbor block A0 or neighbor block A2 adjacent to neighbor block A1 can be inserted into the second motion information merge list.

[0540] Meanwhile, if motion information of a block adjacent to a neighbor block at a predefined location is not available, motion information merging candidates can be derived based on motion information of a non-adjacent block that is adjacent to the neighbor block at the predefined location but not adjacent to the current block. For example, if motion information of neighbor block A0 or neighbor block A2 adjacent to neighbor block A1 is not available, motion information merging candidates can be derived from the motion information of a non-adjacent block a1 that is adjacent to neighbor block A1 but not adjacent to the current block.

[0541] As another example, motion information merging candidates derived from neighboring blocks adjacent to the current block may be inserted into the first motion information merging list, and motion information merging candidates derived from non-adjacent blocks of the current block may be inserted into the second motion information merging list. That is, motion information merging candidates derived from at least one of A0 to A4 and B0 to B5 may be inserted into the first motion information merging list, and motion information merging candidates derived from at least one of a0 to a4 and b0 to b7 may be inserted into the second motion information merging list.

[0542] As another example, a single motion information merge list (i.e., an initial motion information merge list) can be derived for the current block, and then the single motion information merge list can be separated to generate a first motion information merge list and a second motion information merge list. For example, the spatial neighbor blocks, spatially non-adjacent blocks, and temporal neighbor blocks of the current block can be sequentially searched to generate the initial motion information merge list of the current block.

[0543] Afterwards, motion information merge candidates with indices of 0 or even within the initial motion information merge list can be reorganized into the first motion information merge list, and motion information merge candidates with indices of odd within the initial motion information merge list can be reorganized into the second motion information merge list.

[0544] Alternatively, a predetermined number of motion information merge candidates with small indices within the initial motion information merge list can be reorganized into a first motion information merge list, and the remaining motion information merge candidates can be reorganized into a second motion information merge list.

[0545] Meanwhile, motion information merging candidates can be reordered according to the template cost. The template cost can be derived based on the reference template indicated by the motion information of the current template and the motion information merging candidate. For example, the template cost can be SAD, SSD, SATD, or MR-SAD between the current template and the reference template.

[0546] The rearrangement of motion information merge candidates can be performed for each of the first motion information merge list and the second motion information merge list. That is, the rearrangement for the first motion information merge list is performed only on the motion information merge candidates included in the first motion information merge list, and the rearrangement for the second motion information merge list can be performed on the motion information merge candidates included in the second motion information merge list.

[0547] Meanwhile, reordering may involve sorting the movement information merge candidates in order of lowest template cost. As a result of reordering, the smallest index (i.e., 0) may be assigned to the movement information merge candidate with the lowest template cost.

[0548] Alternatively, after performing a reorder on the initial motion information merge list, the first motion information merge candidate and the second motion information merge list may be constructed based on the reordered initial motion information merge list.

[0549] For example, motion information merge candidates can be reordered in order of lowest template cost, and then the reordered motion information merge candidates can be inserted alternately into the first motion information merge list and the second motion information merge list.

[0550] Alternatively, after rearranging the motion information merge candidates in order of lowest template cost, a predetermined number of motion information merge candidates can be inserted into the first motion information merge list, and the remaining motion information merge candidates can be inserted into the second motion information merge list.

[0551] The index assigned to a motion information merging candidate can be adaptively determined by considering whether the motion information merging candidate has bidirectional motion information or unidirectional motion information.

[0552] For example, a combination of motion information merging candidates may consist of two motion information merging candidates, at least one of which has unidirectional motion information, or two motion information merging candidates, both of which have bidirectional motion information. For convenience of explanation, a combination of two motion information merging candidates having at least one unidirectional motion information will be referred to as a unidirectional motion information combination, and a combination of two motion information merging candidates having both of which have bidirectional motion information will be referred to as a bidirectional motion information combination.

[0553] The index assigned to the unidirectional motion information combination can be set to have a smaller value than the index assigned to the bidirectional motion information combination.

[0554] That is, when configuring the first motion information merge candidate and the second motion information merge candidate, the index assigned to the motion information merge candidate having unidirectional motion information can be set to have a smaller value than the index assigned to the motion information merge candidate having bidirectional motion information.

[0555] Alternatively, conversely, the index assigned to the bidirectional motion information combination can be set to have a smaller value than the index assigned to the unidirectional motion information combination.

[0556] That is, when configuring the first motion information merge candidate and the second motion information merge candidate, the index assigned to the motion information merge candidate having bidirectional motion information can be set to have a smaller value than the index assigned to the motion information merge candidate having unidirectional motion information.

[0557] Meanwhile, the index assigned to the motion information merge candidates can be adjusted by adjusting the insertion order of the motion information merge candidates or by reordering the motion information merge candidates inserted into the motion information merge list.

[0558] Meanwhile, information indicating whether to perform a prediction using multiple motion information merge lists for the current block may be encoded and signaled. The information may be a 1-bit flag.

[0559] If the above information indicates that multiple motion information merge lists are used, a prediction for the current block can be performed by additionally using a second motion information merge list in addition to the first motion information merge list. In this case, a combination of motion information merge candidates included in the first motion information merge list and motion information merge candidates included in the second motion information merge list can be selected by the index information.

[0560] On the other hand, if the above information indicates that multiple motion information merge lists are not used, a prediction for the current block can be performed based on a single motion information merge list. In this case, one of the motion information merge candidates included in the motion information merge list can be selected by the index information.

[0561] The maximum number of motion information merge candidates that can be included in the first motion information merge list and the motion information merge candidates that can be included in the second motion information merge list can be set to be the same.

[0562] Alternatively, the maximum number of motion information merge candidates that can be included in the first motion information merge list and the maximum number of motion information merge candidates that can be included in the second motion information merge list may be set differently. Here, the maximum number may be a natural number such as 1, 2, 3, 4, 5, or 6.

[0563] Meanwhile, if the number of motion information merge candidates included in the first motion information merge list and the number of motion information merge candidates included in the second motion information merge list are different, either of the motion information merge candidates may be used multiple times.

[0564] For example, it is assumed that the first motion information merging list contains N motion information merging candidates, and the second motion information merging list contains M motion information merging candidates. Additionally, it is assumed that M is a number smaller than N.

[0565] In this case, combinations of M motion information merge candidates can be generated by using the M motion information merge candidates included in the first motion information merge list and the M motion information merge candidates included in the second motion information merge list.

[0566] Subsequently, the remaining (NM) motion information merge candidates included in the first motion information merge list can be combined with (NM) motion information merge candidates with small index or template costs in the second motion information merge list to additionally generate combinations of (NM) motion information merge candidates.

[0567] Alternatively, the remaining (N-) motion information merge candidates included in the first motion information merge list can be combined with the motion information merge candidate with the smallest index or template cost in the second motion information merge list to generate additional combinations of (NM) motion information merge candidates. That is, the motion information merge candidate with the smallest index or template cost in the second motion information merge list can be used by default to generate additional combinations.

[0568] Consequently, at least one motion information merging candidate within the first motion information merging list and the second motion information merging list can be used as a composition of multiple combinations.

[0569] In the example described above, multiple motion information merging candidates are selected by a single index. That is, the index is exemplified as identifying one of multiple combinations.

[0570] Unlike the example described, multiple motion information merging candidates may also be selected based on multiple indices. For example, a first index pointing to one of the motion information merging candidates included in a first motion information merging list may be encoded / decoded, and a second index pointing to one of the motion information merging candidates included in a second motion information merging list may be encoded / decoded.

[0571] When each of the two selected motion information merging candidates has bidirectional motion information, the prediction block derived from each of the two motion information merging candidates can be obtained by weighted summing the L0 prediction block and the L1 prediction block. For example, among the two motion information merging candidates, the first prediction block based on the first motion information merging candidate can be obtained by weighted summing the first L0 prediction block derived based on the L0 motion information of the first motion information merging candidate and the first L1 prediction block derived based on the L1 motion information of the first motion information merging candidate.

[0572] Likewise, among the two motion information merging candidates, the second prediction block based on the second motion information merging candidate can be obtained by weighting the second L0 prediction block derived based on the L0 motion information of the second motion information merging candidate and the second L1 prediction block derived based on the L1 motion information of the second motion information merging candidate.

[0573] Specifically, the final predicted block of the current block can be obtained through the following mathematical formulas 11 to 13.

[0574]

[0575]

[0576]

[0577] In Equation 11, ref0_list0 represents a first L0 prediction block obtained based on the L0 motion information of a first motion information merging candidate, and ref0_list1 represents a first L1 prediction block obtained based on the L1 motion information of a first motion information merging candidate. α0 and β0 represent weights assigned to the first L0 prediction block and the first L1 prediction block, respectively. Pred0 represents the first prediction block.

[0578] In Equation 12, ref1_list0 represents a second L0 prediction block obtained based on the L0 motion information of a second motion information merging candidate, and ref1_list1 represents a second L1 prediction block obtained based on the L1 motion information of a second motion information merging candidate. α1 and β1 represent weights assigned to the second L0 prediction block and the second L1 prediction block, respectively. Pred1 represents the second prediction block.

[0579] In Equation 13, pred represents the final prediction block of the current block. α and β represent the weights assigned to the first prediction block and the second prediction block, respectively.

[0580] Meanwhile, in Equations 11 to 13, normalization is performed by dividing the value obtained by assigning weights to two prediction blocks by the sum of the weights. Specifically, in Equation 11, normalization is performed by dividing the weighted sum result of the first L0 prediction block and the first L1 prediction block by the sum of the weights (α0 + β0), and in Equation 8, normalization is performed by dividing the weighted sum result of the second L0 prediction block and the second L1 prediction block by the sum of the weights (α1 + β1).

[0581] However, as the number of normalization cycles increases, losses accumulate, which may lead to a decrease in encoding / decoding efficiency. To resolve the above problem, Equations 11 to 13 may be modified as shown in Equations 14 to 16 below.

[0582]

[0583]

[0584]

[0585] According to mathematical formulas 10 to 12, normalization is not performed when acquiring the first prediction block and the second prediction block, and normalization is performed only when acquiring the final prediction block. That is, the number of normalization operations is reduced to one, thereby minimizing the loss caused by normalization.

[0586] In at least one of the above mathematical formulas 11 to 16, weights may be set so that the value of the denominator for normalization is expressed as a power of 2 (e.g., 2, 4, 8, 16, 32, or 64).

[0587]

[0588] Applying the embodiments described with reference to the decoding process or the encoding process to the encoding process or the decoding process is included within the scope of the present disclosure. Changing the embodiments described in a predetermined order to a different order from that described is also included within the scope of the present disclosure.

[0589] Although the above disclosure is described based on a series of steps or flowcharts, this does not limit the chronological order of the invention and may be performed simultaneously or in a different order as necessary. Furthermore, each component constituting the block diagram in the above disclosure (e.g., unit, module, etc.) may be implemented as a hardware device or software, or a plurality of components may be combined to be implemented as a single hardware device or software. As an example, the hardware device may include at least one of a processor for performing operations, a memory for storing data, a transmitter for transmitting data, and a receiver for receiving data.

[0590] The aforementioned disclosure may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either individually or in combination.

[0591] Additionally, according to the present disclosure, a computer-readable recording medium may be provided for storing a bitstream generated by the encoding method described above. The bitstream may be transmitted by an encoding device, and a decoding device may receive the bitstream and decode an image.

[0592] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present disclosure, and vice versa.

[0593] The present disclosure may be applied to a computing or electronic device capable of encoding / decoding a video signal.

Claims

1. A step of constructing a merge list of movement information for the current block; A step of inducing motion information for the current block based on the above motion information merge list; and The method includes the step of obtaining a predicted block of the current block based on the movement information of the current block, The above motion information merging list is derived by adding a first additional motion information merging candidate to the initial motion information merging list, and The above-mentioned first additional motion information merging candidate is derived based on the motion information merging candidates included in the above-mentioned initial motion information merging list, and A video decoding method characterized in that the reference picture index for a predetermined direction of the first additional motion information merging candidate is set to indicate the one with the smallest template cost among the reference pictures for a predetermined direction of the motion information merging candidates.

2. In Paragraph 1, The motion vector for the predetermined direction of the first additional motion information merging candidate is derived by averaging the motion vectors for the predetermined direction of the plurality of motion information merging candidates, and An image decoding method characterized in that the above-mentioned predetermined direction is the L0 direction or the L1 direction.

3. In Paragraph 2, A video decoding method characterized in that, if at least one of the plurality of motion information merging candidates does not have motion information for the predetermined direction, the at least one motion information merging candidate is not used to derive the motion vector for the predetermined direction of the first additional motion information merging candidate.

4. In Paragraph 2, A video decoding method characterized by deriving a motion vector for the predetermined direction of the first additional motion information merging candidate, wherein if at least one of the plurality of motion information merging candidates does not have motion information for the predetermined direction, the motion vector for the predetermined direction of the at least one motion information merging candidate is considered to be 0.

5. In Paragraph 2, A video decoding method characterized in that the template cost of each of the above reference pictures is derived based on a current template composed of a restored area around the current block and a reference template at a position spaced apart from the current template within the reference picture by the motion vector of the first additional motion information merging candidate.

6. In Paragraph 1, A video decoding method characterized in that the number of the plurality of motion information merging candidates is greater than 2.

7. In Paragraph 1, The above motion information merging list further includes a second additional motion information merging candidate, and A video decoding method characterized in that the number of motion information merging candidates used to derive the first additional motion information merging candidate and the number of motion information merging candidates used to derive the second additional motion information merging candidate are different.

8. In Paragraph 7, A video decoding method characterized in that the first additional motion information merging candidate is unusable to induce the second additional motion information merging candidate.

9. In Paragraph 1, A video decoding method characterized in that the above motion information merging list is divided into a first motion information merging list and a second motion information merging list.

10. In Paragraph 9, The above movement information is derived from one selected combination among a plurality of combinations, and A video decoding method characterized in that the above combination consists of a first motion information merging candidate belonging to the first motion information merging list and a second motion information merging candidate belonging to the second motion information merging list.

11. In Paragraph 10, A video decoding method characterized in that the above prediction block is obtained by weighted summing a first prediction block derived based on the first motion information of the first motion information merging candidate and a second prediction block derived based on the second motion information of the second motion information merging candidate.

12. In Paragraph 9, A video decoding method characterized in that each of the first motion information merging list and the second motion information merging list is rearranged by a template cost.

13. In Paragraph 9, The first motion information merging list above is composed of motion information merging candidates whose indices within the motion information merging list are even, and A video decoding method characterized in that the second motion information merging list is composed of motion information merging candidates in which the index within the motion information merging list is odd.

14. Step of constructing a merge list of movement information for the current block; A step of inducing motion information for the current block based on the above motion information merge list; and The method includes the step of obtaining a predicted block of the current block based on the movement information of the current block, The above motion information merging list is derived by adding a first additional motion information merging candidate to the initial motion information merging list, and The above-mentioned first additional motion information merging candidate is derived based on the motion information merging candidates included in the above-mentioned initial motion information merging list, and A video encoding method characterized in that the reference picture index for a predetermined direction of the first additional motion information merging candidate is set to indicate the one with the smallest template cost among the reference pictures for a predetermined direction of the motion information merging candidates.

15. A processor for acquiring compressed video data; and It includes a transmission unit that transmits the above-mentioned compressed video data, The above compressed video data is, Step of constructing a merge list of movement information for the current block; A step of inducing motion information for the current block based on the above motion information merge list; and Based on the movement information of the current block, the predicted block of the current block is obtained through the step of obtaining, and The above motion information merging list is derived by adding a first additional motion information merging candidate to the initial motion information merging list, and The above-mentioned first additional motion information merging candidate is derived based on the motion information merging candidates included in the above-mentioned initial motion information merging list, and A device for transmitting compressed video data, characterized in that the reference picture index for a predetermined direction of the first additional motion information merging candidate is set to indicate the one with the smallest template cost among the reference pictures for a predetermined direction of the motion information merging candidates.

Citation Information

Patent Citations

  • Vehicle door latch apparatus

    KR1020250024319A

  • Educational teaching aids for dental care and oral health and educational method using the same

    KR1020250115838A

  • Display device and method for manufacturing the same

    KR1020260007455A

  • Decoder side motion information derivation

    US20240305786A1

  • KR20240137041A